Program, information processing method, information processing device, and model generation method
The program generates accurate 3D shape data from X-ray images using learning models, addressing the inefficiency of existing methods by reducing computational requirements and enabling cost-effective, radiation-free medical imaging.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2026-03-05
AI Technical Summary
Existing methods for reconstructing a three-dimensional image from an X-ray image require significant computational resources, such as computation time and memory usage, making them inefficient for widespread medical applications.
A program that utilizes a series of learning models, including a first learning model to generate depth images and a second learning model to correct and synthesize 3D shape data from a single X-ray image, reducing the need for extensive computational resources.
Enables the generation of accurate 3D shape data from X-ray images without requiring large computational resources, facilitating efficient medical imaging without the need for costly or radiation-exposing CT or MRI scans.
Smart Images

Figure JP2025030308_05032026_PF_FP_ABST
Abstract
Description
Program, information processing method, information processing device, and model generation method
[0001] The present disclosure relates to a program, an information processing method, an information processing device, and a model generation method.
[0002] For example, in artificial joint replacement surgery or fracture treatment, understanding the three-dimensional shape of bones is useful for performing surgery and treatment, and medical institutions take CT (Computed Tomography) images or MRI (Magnetic Resonance Imaging) images in addition to X-ray images. However, CT and MRI examinations are more expensive than X-ray examinations, and CT examinations are affected by radiation exposure, while MRI examinations take a long time to perform, so they are performed less frequently than X-ray examinations. In particular, in Europe and the United States, CT and MRI examinations are often not performed due to the cost and radiation exposure involved.
[0003] Therefore, Patent Document 1 and Non-Patent Document 1 disclose techniques for restoring a three-dimensional image from an X-ray image of an object. By restoring a three-dimensional image of an object from an X-ray image as in the techniques disclosed in Patent Document 1 and Non-Patent Document 1, it becomes possible to grasp the three-dimensional shape of the object without taking CT images or MRI images.
[0004] International Publication No. 2022 / 229816
[0005] Ryoya Shiode et al., “2D-3D reconstruction of distal forearm bone from actual X-ray images of the wrist using convolutional neural networks”, Scientific Reports, 2021 Jul 27, 11(1):15249
[0006] However, the techniques for reconstructing a three-dimensional image of an object from an X-ray image as disclosed in Patent Document 1 and Non-Patent Document 1 require a large amount of computational resources, such as computation time and memory usage, to process the three-dimensional image.
[0007] The present disclosure aims to provide a program or the like that is capable of acquiring the three-dimensional shape of an object from an X-ray image of the object without requiring a large amount of computational resources.
[0008] A program according to one aspect of the present disclosure acquires an X-ray image of a target area, and inputs the acquired X-ray image into a learning model that outputs depth images of the front and back of the target area in the X-ray image when the X-ray image of the target area is input, thereby causing a computer to execute a process of acquiring depth images of the front and back of the target area in the input X-ray image.
[0009] In one aspect of the present disclosure, the three-dimensional shape of an object can be obtained from an X-ray image of the object without requiring a large amount of computational resources.
[0010] 1 is a block diagram showing an example of the configuration of an information processing device. FIG. 1 is an explanatory diagram showing an example of the configuration of a first learning model. FIG. 2 is an explanatory diagram showing an example of the configuration of a second learning model. FIG. 3 is an explanatory diagram showing an example of the configuration of a third learning model. FIG. 4 is a flowchart showing an example of a processing procedure for generating training data for the first learning model. FIG. 5 is an explanatory diagram of a processing procedure for generating depth images for training. FIG. 6 is a flowchart showing an example of a processing procedure for generating training data for the first learning model. FIG. 7 is a flowchart showing an example of a processing procedure for generating training data for the second learning model. FIG. 8 is a flowchart showing an example of a processing procedure for generating training data for the third learning model. FIG. 9 is a flowchart showing an example of a processing procedure for generating 3D shape data of a target region. FIG. 10 is an explanatory diagram showing an example of a screen. FIG. 11 is an explanatory diagram showing an effect of using the information processing device of the present disclosure. FIG. 12 is an explanatory diagram showing a modified example of the first learning model. FIG. 13 is an explanatory diagram showing a modified example of the second learning model. FIG. 14 is an explanatory diagram showing a modified example of the third learning model. FIG. 15 is an explanatory diagram showing another modified example of the first learning model. FIG. 16 is an explanatory diagram showing another modified example of the second learning model. FIG. 17 is an explanatory diagram showing another modified example of the third learning model. FIG. 18 is a flowchart showing an example of a processing procedure for generating 3D shape data of a target region according to embodiment 3. 10 is an explanatory diagram of a process for generating 3D shape data. FIG. ... thickness image. FIG. 10 is an explanatory diagram showing an example of the configuration of a fourth learning model. FIG. 10 is a flowchart showing an example of a process for generating training data for the fourth learning model. FIG. 10 is a flowchart showing an example of a process for generating 3D shape data of a target part in embodiment 4. FIG. 10 is an explanatory diagram showing an example of a recess. FIG. 10 is an explanatory diagram showing a modified example of the first learning model. FIG. 10 is an explanatory diagram showing a modified example of the second learning model. FIG. 10 is an explanatory diagram showing a modified example of the third learning model.
[0011] Hereinafter, a program, an information processing method, an information processing device, and a model generation method according to the present disclosure will be described with reference to the drawings illustrating embodiments thereof.
[0012] (Embodiment 1) An information processing device that generates 3D (dimensional) shape data of a target region based on an X-ray image of the target region captured by an X-ray device will be described. In this embodiment, a configuration will be described in which 3D shape data of the pelvis and femur is generated from a single frontal X-ray image of the hip joint, which includes the pelvis and femur within an imaging range and is captured by irradiating X-rays from the front of the subject (patient). However, the target region is not limited to the pelvis and femur, and may be other bone regions, such as the lumbar vertebrae and thoracic vertebrae. Furthermore, the target region is not limited to a bone region, but may be a muscle region or an organ. For example, a configuration may be used in which 3D shape data of muscle regions, such as the gluteus maximus, gluteus medius, hamstrings, and quadriceps, is generated from a frontal X-ray image of the hip joint. Furthermore, in this embodiment, 3D shape data of the target region is generated from a single X-ray image of the target region captured from one direction. However, 3D shape data of the target region may also be generated from multiple X-ray images of the target region captured from multiple directions. For example, 3D shape data of the target region may be generated from each of a plurality of X-ray images, and the 3D shape data of the target region may be generated by averaging the generated plurality of 3D shape data.
[0013] FIG. 1 is a block diagram showing an example configuration of an information processing device. The information processing device 10 is a device capable of various information processing and information transmission / reception, such as a personal computer, server computer, workstation, or tablet PC (personal computer). The information processing device 10 is installed and used in a medical institution, testing institution, research institution, or the like. The information processing device 10 is not limited to a single computer, but may be a multi-computer configured including multiple computers, or may be a virtual machine virtually constructed by software within a single device. When the information processing device 10 is configured as a server computer, the information processing device 10 may be a local server installed in a medical institution, or a cloud server connected to the information processing device via a network such as the Internet. The following description will be given assuming that the information processing device 10 is a single computer.
[0014] The information processing device 10 of this embodiment generates 3D shape data of a target region (e.g., a pelvis and a femur) based on an X-ray image of the target region (e.g., a frontal X-ray image of a hip joint including a pelvis and a femur). Specifically, the information processing device 10 prepares a first learning model M1 and a second learning model M2 that have been trained in advance by performing machine learning (deep learning) to learn predetermined training data. The first learning model M1 is a model trained to receive an X-ray image as input and output depth images of the front and rear surfaces of the target region in the X-ray image. The second learning model M2 is a model trained to receive 3D shape data of the target region as input and output corrected 3D shape data (corrected three-dimensional shape data) in which missing parts, etc. in the 3D shape data have been corrected (complemented). For example, the information processing device 10 inputs a frontal X-ray image of a hip joint into the first learning model M1, thereby acquiring depth images of the front and rear surfaces of the pelvis and a femur in the frontal X-ray image of the hip joint from the first learning model M1. The front surface of the target area refers to the surface on the front side of the subject in the X-ray image (the surface on the near side in the X-ray irradiation direction), and the rear surface refers to the surface on the back side of the subject (the surface on the far side in the X-ray irradiation direction). The front and rear surfaces of the target area may also refer to the rear and front sides of the subject. A depth image is an image (2D image) on a 2D coordinate plane perpendicular to the imaging direction (X-ray irradiation direction) of the target area, and the pixel value of each pixel in the image indicates the distance from the X-ray tube (imaging center) to the imaging surface in the imaging direction. The information processing device 10 generates 3D shape data of the pelvis and femur by synthesizing the front and rear depth images of the pelvis and femur, respectively. The information processing device 10 then inputs the generated 3D shape data of the pelvis and femur into the second learning model M2, thereby obtaining corrected 3D shape data with defects and the like corrected from the second learning model M2. This makes it possible to generate 3D shape data of the target area captured in the X-ray image from the X-ray image.
[0015] The information processing device 10 of this embodiment also prepares a third learning model M3 that is trained to input an X-ray image and output silhouette images in which the shapes (silhouettes) of the front and back surfaces of a target region in the X-ray image are extracted. The information processing device 10 inputs a frontal X-ray image of a hip joint into the third learning model M3 and acquires silhouette images of the front and back surfaces of the pelvis and femur in the frontal X-ray image of the hip joint from the third learning model M3. The information processing device 10 then performs masking on the depth images of the front and back surfaces of the pelvis and femur acquired from the first learning model M1 based on the silhouette images of the front and back surfaces of the pelvis and femur, thereby removing (e.g., replacing with black pixels) areas other than the pelvis and femur in the depth images. The information processing device 10 then synthesizes the depth images of the front and back surfaces of the pelvis and femur from which the background areas have been removed to generate 3D shape data of the pelvis and femur. By removing background areas from the depth images of the front and back of the target area and then combining them, noise-free 3D shape data can be generated. In this case, the information processing device 10 also inputs the generated 3D shape data of the pelvis and femur into the second learning model M2 to obtain corrected 3D shape data of the pelvis and femur.
[0016] The information processing device 10 includes a control unit 11, a storage unit 12, a communication unit 13, an input unit 14, a display unit 15, a reading unit 16, etc., and these units are connected via a bus. The control unit 11 includes one or more processors (arithmetic processing devices), such as a CPU (Central Processing Unit), an MPU (Micro-Processing Unit), a GPU (Graphics Processing Unit), an AI chip (AI semiconductor), etc. The control unit 11 executes the processes to be performed by the information processing device 10 by appropriately executing a program P stored in the storage unit 12. Note that when the control unit 11 includes multiple processors, each process may be executed by the same processor, or each process may be executed by a different processor.
[0017] The storage unit 12 includes RAM (Random Access Memory), flash memory, a hard disk, an SSD (Solid State Drive), etc. The storage unit 12 stores the program P (program product, computer program) executed by the control unit 11 and various data necessary for executing the program P. The storage unit 12 also temporarily stores data generated when the control unit 11 executes the program P. The storage unit 12 also stores a first learning model M1 to a third learning model M3. The learning models M1 to M3 are expected to be used as program modules constituting artificial intelligence software. The learning models M1 to M3 perform predetermined calculations on input values and output the calculation results. The storage unit 12 stores information defining the learning models M1 to M3, such as information on the layers included in the learning models M1 to M3, information on the nodes constituting each layer, and weights (coupling coefficients) between nodes. The storage unit 12 also stores a medical image DB 12a and a training DB 12b. The medical image DB 12a stores corresponding frontal hip joint X-ray images and CT images prepared for learning the learning models M1 to M3. The training DB 12b stores training data used in the learning process of the learning models M1 to M3, and the training data is stored in the training DB 12b when the information processing device 10 performs a training data generation process described below. The storage unit 12 may be composed of multiple storage devices, and part of the storage unit 12 may be another storage device connected to the information processing device 10 or another storage device with which the information processing device 10 can communicate.
[0018] The communication unit 13 is a communication module for performing processes related to wired or wireless communication, and transmits and receives information to and from other devices via the network N. The network N may be a dedicated network, the Internet, a public telephone network, or a local area network (LAN) established within the facility where the information processing device 10 is installed. The input unit 14 accepts user input and sends control signals corresponding to the operation content to the control unit 11. The display unit 15 is a liquid crystal display, an organic EL display, or the like, and displays various information in accordance with instructions from the control unit 11. A part of the input unit 14 and the display unit 15 may be an integrated touch panel. Note that the input unit 14 and the display unit 15 are not essential, and the information processing device 10 may be configured to accept operations via a connected computer and output information to be displayed to an external display device.
[0019] The reading unit 16 reads information stored in a portable storage medium 10a, such as a CD (Compact Disc), a DVD (Digital Versatile Disc), a USB (Universal Serial Bus) memory, an SD card, or a CompactFlash (registered trademark). The program P and various data stored in the storage unit 12 may be read by the control unit 11 from the portable storage medium 10a via the reading unit 16 and stored in the storage unit 12. The program P and various data may be written to the storage unit 12 when the information processing device 10 is manufactured or when it is first used, or may be downloaded by the control unit 11 from another device via the communication unit 13 and stored in the storage unit 12.
[0020] In this embodiment, the program P may be located and executed on a single computer, or at one site, or may be deployed to run on multiple computers distributed across multiple sites and interconnected via a network N.
[0021] FIG. 2 is an explanatory diagram showing an example configuration of the first learning model M1. The first learning model M1 in FIG. 2 is a trained model that uses a frontal hip joint X-ray image as input data, performs calculations to predict depth images of the front and rear surfaces of the pelvis and femur in the frontal hip joint X-ray image, and outputs the calculation results. Note that the first learning model M1 of this embodiment predicts and outputs depth images of the front and rear surfaces of the left and right halves of the pelvis and the left and right femur in the input frontal hip joint X-ray image. The first learning model M1 is a neural network having an encoder and a decoder. As long as it has a configuration including an encoder and a decoder, it may be configured using any algorithm, or may be configured by combining multiple algorithms.
[0022] The encoder of the first learning model M1 has a convolutional layer and a pooling layer. The convolutional layer extracts image features from pixel information of the input X-ray image by filtering or the like to generate a feature map, and the pooling layer compresses (downsamples) the generated feature map. The decoder of the first learning model M1 has a deconvolutional layer (transposed convolutional layer) and enlarges (upsamples) the feature map generated by the encoder to the original image size. The deconvolutional layer detects the positions of the front and back surfaces of each target region in the image (here, the left and right halves of the pelvis, and the left and right femurs) based on the features extracted by the encoder, and generates front and back depth images for each detected target region.
[0023] The first learning model M1 is generated by preparing training data that associates training X-ray images (frontal X-ray images of the hip joint) with depth images of the front and back surfaces of each training target region (left and right halves of the pelvis, left femur, and right femur), and then performing machine learning (deep learning) using this training data. The training depth images are images generated for the front and back surfaces of each target region from CT images obtained by using a CT imaging device to capture an area including the pelvis and femur of the same subject as the training X-ray images. The depth images used indicate the anterior-posterior position (coordinate values, depth) of each point on the front and back surfaces of each target region. Such training data is stored, for example, in a training DB 12b provided in the memory unit 12 and used during the learning process.
[0024] The first learning model M1 learns to output correct depth images (depth images of the front and back of each target region) included in the training data when an X-ray image included in the training data is input. During the learning process, the first learning model M1 inputs the input X-ray image to an encoder, and generates a feature map by extracting image features while reducing the image size using the encoder. The first learning model M1 inputs the feature map generated by the encoder to a decoder, and generates depth images of the front and back of the target region in the X-ray image using the decoder. The first learning model M1 then compares each of the generated depth images of the front and back of the target region with the correct depth images indicated by the training data and optimizes parameters used in the calculation process so that the two images approximate each other. For example, the first learning model M1 optimizes parameters such as weights (coupling coefficients) between nodes in the decoder and encoder, and coefficients of functions used at each node (weights of filters (kernels) in the convolution layer) using backpropagation, steepest descent, or the like. This results in a first learning model M1 that, when a frontal X-ray image of a hip joint is input, outputs depth images of the front and back surfaces of the target area (pelvis and femur) in the X-ray image.
[0025] FIG. 3 is an explanatory diagram showing an example configuration of the second learning model M2. The second learning model M2 in FIG. 3 is a trained model that uses 3D shape data of a target region (a femur in the example of FIG. 3 ) as input data, performs calculations to correct (complement) missing portions in the 3D shape data, and outputs the calculation results (corrected 3D shape data). The second learning model M2 can also be configured as a neural network having an encoder and a decoder. As long as it has an encoder and a decoder, it can be configured using any algorithm, or it can be configured by combining multiple algorithms. Furthermore, the second learning model M2 can be configured, for example, as a statistical shape model (SSM) using a dimensionality reduction algorithm such as principal component analysis (PCA).
[0026] The second learning model M2 is generated by preparing training data that associates training 3D shape data with training corrected 3D shape data and performing machine learning (deep learning) using this training data. The training 3D shape data can be 3D shape data of a target region (e.g., a femur) generated by synthesizing depth images of the front and back surfaces of the target region generated from CT images of the target region. That is, 3D shape data can be used in which the anterior-posterior coordinate values of each point on the front and back surfaces of the target region are extracted from the CT images, and depth images based on the extracted coordinate values are synthesized with the contour lines of the front and back surfaces. The training corrected 3D shape data can be 3D shape data indicating the shape of the target region generated from the same CT image as the training 3D shape data. That is, the coordinate values (3D coordinate values) of each point on the surface of the target region are extracted from the CT images, and 3D shape data generated based on the extracted coordinate values can be used. Such training data is stored in, for example, a training DB 12b provided in the storage unit 12 and is used during the learning process.
[0027] The second learning model M2 learns to output corrected 3D shape data included in the training data when 3D shape data included in the training data is input. During the learning process, the second learning model M2 performs calculations based on the input 3D shape data and generates corrected 3D shape data in which missing portions of the input 3D shape data have been corrected as output data. The second learning model M2 then compares the generated corrected 3D shape data with the correct 3D shape data indicated by the training data and optimizes parameters used in the calculation process so that the two data approximate each other. The second learning model M2 optimizes parameters such as weights (coupling coefficients) between nodes in the decoder and encoder and coefficients of functions used at each node using backpropagation, steepest descent, or the like. This results in a second learning model M2 that, when 3D shape data of a target region is input, outputs corrected 3D shape data in which missing portions of the 3D shape data have been corrected.
[0028] When the second learning model M2 is configured as a statistical shape model using principal component analysis, the second learning model M2 performs principal component analysis on the 3D shape data included in the training data and the corrected 3D shape data as a single piece of shape data, and calculates and sets a vector that represents the statistical tendency of the corrected 3D shape data relative to the 3D shape data for each principal component. As a result, when 3D shape data is input, the second learning model M2 can output corrected 3D shape data by performing calculations based on the 3D shape data and the set vector.
[0029] FIG. 4 is an explanatory diagram showing an example configuration of the third learning model M3. The third learning model M3 is a model that recognizes a predetermined object contained in an input image and can classify objects in the image pixel by pixel, for example, by semantic segmentation. The third learning model M3 in FIG. 4 is trained to use a frontal hip joint X-ray image as input data, perform calculations to recognize the front and back surfaces of the pelvis and femur in the frontal hip joint X-ray image, and output the calculation results. Note that the third learning model M3 of this embodiment recognizes the front and back surfaces of each part, i.e., the left and right halves of the pelvis and the left and right femur, in the input frontal hip joint X-ray image. For the front and back surfaces of each part, it classifies each pixel of the input X-ray image into a region corresponding to the part and a background region, and outputs a classified X-ray image (hereinafter referred to as a silhouette image showing the shape of each part) in which each pixel is associated with a label for each region. In the example of Figure 4, the third learning model M3 outputs a silhouette image for the front and back of each body part, in which pixels classified into the area of each body part are shown in white and pixels classified into the background area are shown in black.
[0030] The third learning model M3 is a neural network having an encoder and a decoder, and may be configured using any algorithm or a combination of multiple algorithms as long as it has an encoder and a decoder. For example, the third learning model M3 may be configured using an image segmentation algorithm such as U-Net, FCN (Fully Convolutional Network), or SegNet, or an object detection algorithm such as YOLO (You Only Look Once), SSD (Single Shot Multi-Box Detector), or ViT (Vision Transformer).
[0031] The third learning model M3 is generated by machine learning (deep learning) using training data that associates training X-ray images (frontal X-ray images of the hip joint) with silhouette images of the front and back of each training target region (left and right halves of the pelvis, left femur, and right femur). The training silhouette images are images generated for the front and back of each target region from CT images obtained by using a CT scanner to capture an area including the pelvis and femur of the same subject as in the training X-ray images. The silhouette images used are those in which each pixel in a two-dimensional image of the subject viewed from the front is labeled with data indicating the front or back region of each target region or the background region (pixels in each region are labeled with 0 (white) and pixels in the background region are labeled with 1 (black)). Such training data is also stored, for example, in a training DB 12b provided in the memory unit 12 and used during the learning process.
[0032] The third learning model M3 learns to output correct silhouette images (front and back silhouette images of each target region) included in the training data when an X-ray image included in the training data is input. During the learning process, the third learning model M3 performs calculations based on the input X-ray images and generates, as output data, silhouette images representing the recognition results of the front and back regions of the target region in the input X-ray images. Specifically, the third learning model M3 generates a silhouette image for each recognized region in which each pixel in the X-ray image is labeled with a value (0 or 1) indicating the region or background region of each recognized region. The third learning model M3 then compares each of the generated front and back silhouette images of the target region with the correct silhouette images indicated in the training data and optimizes parameters used in the calculation process so that the two images approximate each other. The third learning model M3 optimizes parameters such as weights (coupling coefficients) between nodes in the decoder and encoder of the third learning model M3 and coefficients of functions used in each node using backpropagation, steepest descent, etc. This results in a third learning model M3 that, when a frontal X-ray image of a hip joint is input, outputs silhouette images of the front and back of the target area (pelvis and femur) in the X-ray image.
[0033] The third learning model M3 is not limited to the example in Fig. 4. For example, the third learning model M3 may be configured to include the configuration of the first learning model M1, predict depth images of the anterior and posterior surfaces of the pelvis and femur in an input frontal X-ray image of the hip joint, recognize the anterior and posterior surfaces of the pelvis and femur in the predicted depth image based on the predicted depth image, and generate silhouette images of the anterior and posterior surfaces of the pelvis and femur.
[0034] The information processing device 10 prepares the above-mentioned learning models M1 to M3 in advance and uses them when generating 3D shape data of the target area from X-ray images of the target area. The learning of the learning models M1 to M3 may be performed on another learning device. The trained learning models M1 to M3 generated by training on the other learning device are downloaded from the learning device to the information processing device 10 via the network N or the portable storage medium 10a, for example, and stored in the storage unit 12.
[0035] The following describes the process of generating training depth images used in learning the first learning model M1. Figure 5 is a flowchart showing an example of the process procedure for generating training data for the first learning model M1, and Figure 6 is an explanatory diagram of the process of generating training depth images. The following process is executed by the control unit 11 of the information processing device 10 according to the program P stored in the memory unit 12, but may also be executed by another information processing device or learning device. In the following process, it is assumed that the X-ray images and CT images used to generate the training data are pairs of frontal X-ray images and CT images of the hip joint, which are images of the subject's pelvis and areas including the left and right femurs, and are stored in correspondence with each other in the medical image DB 12a.
[0036] The control unit 11 of the information processing device 10 reads one pair of a hip joint frontal X-ray image and a CT image stored in the medical image DB 12a (S11). First, the control unit 11 performs a brightness value calibration process on the read CT image (S12) to correct each brightness value (CT value) in the CT image. The brightness value (CT value) of each pixel measured by CT varies due to individual differences in the X-ray CT device, the installation environment, and imaging conditions. Therefore, calibration is required to correct the deviation in the measured value. The CT image calibration process is performed using calibration data acquired, for example, when the X-ray CT device is installed, when components such as the tube are replaced, at the start of imaging, or periodically. The calibration data is generated based on the radiation density (CT value expressed in Hounsfield units (HU)) obtained by imaging a phantom made of a material with known characteristics using the X-ray CT device and the tissue density of the material. Specifically, a conversion formula for converting the density of radiation obtained by passing through the phantom into the tissue density of the substance is used in the calibration data.
[0037] The calibration process can be performed using the method described in the paper "Automated segmentation of an intensity calibration phantom in clinical CT images using a convolutional neural network" by the inventors of the present application, Keisuke Uemura et al., published online in the International Journal of Computer-Assisted Radiology and Surgery (IJCARS) on March 17, 2021. This paper discloses a technique for automatically extracting the imaging area of each material in the phantom using an X-ray CT scanner along with a subject, using a convolutional neural network (CNN). Using the technology disclosed in this paper, calibration data can be generated to convert the CT value (radiation density) of each material's imaging area extracted from the CT image into the tissue density of each material. By using this generated calibration data to perform a calibration process on the CT value of the subject's imaging area, accurate tissue density can be obtained. It should be noted that the process of step S12 does not necessarily have to be performed.
[0038] Next, the control unit 11 performs a process on the calibrated CT image using a segmentation deep neural network (DNN) to classify each pixel in the CT image into one of multiple regions (musculoskeletal regions): bone region, muscle region, and other region (S13). The process of classifying each pixel in the CT image into a musculoskeletal region can be performed using, for example, the method described in the paper "Automated Muscle Segmentation from Clinical CT Using Bayesian U-Net for Personalized Musculoskeletal Modeling" by the inventors of the present application, Yoshito Otake et al., published on pages 1030-1040 of IEEE Transactions on Medical Imaging, Vol. 39, No. 4, April 2020. This paper discloses a musculoskeletal segmentation model that inputs a CT image, classifies each pixel in the input CT image into one of bone region, muscle region, and other region, and outputs a classified CT image (musculoskeletal labeled image) in which each pixel is associated with a label for each region. The musculoskeletal segmentation model disclosed in this paper is constructed using a Bayesian U-net. As a result, as shown in (1) in Figure 6, each pixel in the CT image is classified into one of three regions, and a musculoskeletal labeled image can be obtained in which each region is associated with a label. In Figure 6, each pixel in the musculoskeletal labeled image is schematically represented by a color (shade) corresponding to the classified region and muscle type, etc.
[0039] The control unit 11 extracts data of the region of interest (target region) from the CT image based on the musculoskeletal label image generated from the CT image (S14). Figure 6 shows an example in which the left femur is the region of interest. In this case, the control unit 11 extracts data of the bone region from the CT image based on the musculoskeletal label image, as shown in (2) of Figure 6, and then extracts data of the region of interest (left femur) from the extracted data of the bone region (CT image), as shown in (3) of Figure 6. The process of extracting the region of interest (left femur) from the bone region can be performed, for example, by pattern matching using a template. In this case, a template representing the shape of the left femur is stored in the storage unit 12 in advance. The control unit 11 determines whether or not there is a region matching the template in the CT image of the bone region, and extracts the region matching the template from the bone region, thereby extracting data of the femur from the bone region. The process of extracting a region of interest from a bone region can also be performed using a learning model that has been machine-learned to output the region of interest (left femur) in the bone region when a CT image of the bone region is input. In this case, the control unit 11 inputs the CT image of the bone region into the learning model and can identify and extract the region of interest in the bone region based on output information from the learning model. The control unit 11 may also extract data of the region of interest directly from the CT image without extracting data of the bone region from the CT image.
[0040] Next, the control unit 11 aligns the region of interest (here, the left femur) between the X-ray image (frontal hip joint X-ray image) acquired in step S11 and the CT image of the region of interest extracted in step S14 (S15). The control unit 11 detects the brightness gradient (edge) of the X-ray image based on the pixel values of each pixel, and identifies the region of interest in the X-ray image based on the detected brightness gradient. The control unit 11 may identify the region of interest in the X-ray image by pattern matching using a pre-prepared template or by using a pre-trained learning model. The control unit 11 then identifies the imaging direction of the region of interest (left femur) in the CT image that matches the region of interest (left femur) in the X-ray image, and generates a CT image of the region of interest viewed from the identified direction. This allows a CT image of the region of interest aligned with the region of interest in the X-ray image to be acquired, as shown in (4) in FIG. 6 .
[0041] The registration of the subject in the X-ray image with the subject in the CT image can be performed using, for example, the method described in the paper titled "Can Anatomic Measurements of Stem Anteversion Angle Be Considered as the Functional Anteversion Angle?" by the inventors of the present application, Keisuke Uemura et al., which was published on pages 595-600 of The Journal of Arthroplasty 33 (2018). This paper discloses a technique for performing segmentation using a hierarchical statistical shape model on CT images of the pelvis and femur to identify the pelvis and femur in the CT image, and then aligning (associating) the pelvis and femur in the CT image with the pelvis and femur in the X-ray image. Furthermore, the registration of X-ray images and CT images can be performed using the method described in the paper "3D-2D registration in mobile radiographs: algorithm development and preliminary clinical evaluation" by the inventors of the present application, Yoshito Otake et al., which was published on pages 2075-2090 of Physics in Medicine and Biology 60 (2015). This paper discloses a technique for generating a CT image in which the subject is viewed from the same direction as the X-ray image by translating and rotating the CT image.
[0042] The control unit 11 then generates depth images of the region of interest (left femur) by extracting coordinate values of each pixel in the region of interest in the same direction as the X-ray image from the CT image of the region of interest (left femur) that has been aligned with the region of interest (left femur) in the X-ray image (S16). Here, the region of interest is the anterior and posterior surfaces of the left femur, so the control unit 11 generates depth images of the anterior and posterior surfaces of the left femur, as shown in (5) in FIG. 6 . Since the frontal hip joint X-ray image includes the pelvis and femur, the region of interest for which depth images are generated from the corresponding CT image can be the pelvis and femur. Therefore, in this embodiment, the control unit 11 generates depth images of the anterior and posterior surfaces of each of the regions of interest: the left half of the pelvis, the right half of the pelvis, the left femur, and the right femur.
[0043] The control unit 11 determines whether the process of generating depth images based on CT images has been completed for all regions of interest to be extracted from the CT images (S17). If it determines that the process has not been completed (S17: NO), the control unit 11 returns to step S14 and performs steps S14 to S16 for the unprocessed regions of interest. If it determines that the process of generating depth images based on CT images for all regions of interest has been completed (S17: YES), the control unit 11 associates the frontal hip joint X-ray image acquired in step S11 with the front and back depth images of each region of interest generated for each region of interest in step S16, and stores them as training data in the training DB 12b (S18). This allows the depth images of each region of interest captured in the X-ray images to be used as training data. Although Figure 6 only shows depth images of the front and back surfaces of the left femur, similar processing can be performed on the left half of the pelvis, the right half of the pelvis, and the right femur to generate depth images of the front and back surfaces of the left half of the pelvis, the right half of the pelvis, and the right femur.
[0044] The control unit 11 determines whether there are any unprocessed X-ray images and CT images stored in the medical image DB 12a, for which the above-described training data generation process has not yet been performed (S19). If it is determined that there are unprocessed images (S19: YES), the control unit 11 returns to step S11 and executes steps S11 to S18 for the X-ray images and CT images for which training data generation has not yet been performed. If it is determined that there are no unprocessed images (S19: NO), the control unit 11 terminates the series of processes. Through the above-described process, training data used to train the first learning model M1 can be generated based on the X-ray images and CT images stored in the medical image DB 12a and stored in the training DB 12b. In the above-described process, the X-ray images and CT images used to generate the training data are stored in the medical image DB 12a. However, the control unit 11 may also be configured to acquire X-ray images and CT images stored in another device, for example, via the network N or a portable storage medium 10a. For example, the control unit 11 may be configured to acquire X-ray images and CT images from electronic medical record data stored in an electronic medical record server. In addition, in the above-described process, the region of interest is extracted from the CT image and then aligned with the X-ray image, but the region of interest may be extracted from the CT image after alignment with the X-ray image.
[0045] Next, we will explain the process of learning the training data generated by the above-mentioned process to generate the first learning model M1. Figure 7 is a flowchart showing an example of the process procedure for generating the first learning model M1. The following process is executed by the control unit 11 of the information processing device 10 according to the program P stored in the memory unit 12, but may also be performed by another learning device. Furthermore, the process for generating the training data shown in Figure 5 and the process for generating the first learning model M1 shown in Figure 7 may each be performed by a different device.
[0046] The control unit 11 of the information processing device 10 acquires one piece of training data from the training DB 12b (S21). Specifically, the control unit 11 reads one set of a hip joint X-ray front image and multiple depth images (depth images of the front and back surfaces of each of multiple body parts) stored in the training DB 12b. The control unit 11 performs a training process for the first learning model M1 using the read training data (S22). Here, the control unit 11 inputs the hip joint X-ray front image included in the training data into the first learning model M1 and acquires output data (depth images of the front and back surfaces of the target body part) from the first learning model M1. The control unit 11 compares the output data with the correct depth images (depth images of the front and back surfaces of each of the multiple body parts) indicated by the training data and optimizes parameters such as weights between nodes in the encoder and decoder of the first learning model M1 using, for example, backpropagation algorithm, so that the output data and the correct depth images are similar to each other.
[0047] The control unit 11 determines whether there is unprocessed training data that has not been subjected to the learning process among the training data stored in the training DB 12b (S23). If it is determined that there is unprocessed training data (S23: YES), the control unit 11 returns to step S21 and executes the processes of steps S21 to S22 for the training data that has not been subjected to the learning process. If it is determined that there is no unprocessed training data (S23: NO), the control unit 11 ends the series of processes. Through the above-described learning process, by inputting a frontal X-ray image of a hip joint, a first learning model M1 is generated that outputs depth images of the front and rear surfaces of the left half of the pelvis, the right half of the pelvis, the left femur, and the right femur captured in the X-ray image.
[0048] By repeatedly performing the learning process using the training data described above, the first learning model M1 can be further optimized. The first learning model M1 of this embodiment is trained using training data that enables spatial correspondence between the 3D data of each part obtained by CT and the X-ray image with high accuracy, using a segmentation technique that accurately classifies musculoskeletal regions from CT images and a technique that accurately aligns target parts in X-ray images with target parts in CT images. Therefore, the first learning model M1 can be realized to generate highly accurate depth images without requiring a large number of cases (training data).
[0049] Next, the process of generating training data used to learn the second learning model M2 will be described. FIG. 8 is a flowchart showing an example of the process procedure for generating training data for the second learning model M2. The following process is also executed by the control unit 11 of the information processing device 10 according to the program P stored in the memory unit 12, but may also be executed by another information processing device or learning device. Furthermore, the CT images used to generate the training data for the second learning model M2 may be the same as or different from the CT images used to generate the training data for the first learning model M1.
[0050] The control unit 11 of the information processing device 10 reads one CT image stored in the medical image DB 12a (S31). The control unit 11 performs processes similar to steps S12 and S13 in FIG. 5 (S32-S33), calibrating the brightness values of the read CT image and classifying each pixel in the CT image into multiple musculoskeletal regions. The control unit 11 then extracts data on the region of interest from the CT image based on the musculoskeletal labeled image generated from the CT image (S34). For example, the control unit 11 extracts data on the anterior and posterior surfaces of the left femur.
[0051] The control unit 11 generates a depth image of the front surface of the region of interest by extracting coordinate values (e.g., coordinate values in the anterior-posterior direction of the subject) of each pixel of the front surface of the region of interest from the extracted CT image of the region of interest (S35). The control unit 11 also generates a depth image of the posterior surface of the region of interest by extracting coordinate values of each pixel of the posterior surface of the region of interest (S36). The control unit 11 generates 3D shape data of the region of interest by connecting the contours of each region of interest in the generated depth images of the front and posterior surfaces of the region of interest (S37). The 3D shape data can be generated using software such as 3D CAD (Computer Aided Design), and can be polygon data generated using 3D CAD. The 3D shape data can also be voxel data generated using a voxel method.
[0052] Next, the control unit 11 generates 3D shape data of the extracted region of interest from the CT image based on the 3D coordinate values of each pixel in the CT image (S38). This 3D shape data may be polygon data generated using 3D CAD or the like, or voxel data generated by a voxel method. The control unit 11 associates the 3D shape data generated in step S37 with the 3D shape data generated in step S38 (corrected 3D shape data) and stores them as training data in the training DB 12b (S39).
[0053] The control unit 11 determines whether the 3D shape data generation process has been completed for all regions of interest that can be extracted from the CT image acquired in step S31 (S40). If it determines that the process has not been completed (S40: NO), the control unit 11 returns to step S34 and performs the processes of steps S34 to S39 for the unprocessed regions of interest. For example, in the case of a CT image whose imaging range is an area including the pelvis and femur, the left half of the pelvis, the right half of the pelvis, the left femur, and the right femur are set as the regions of interest, and the control unit 11 generates 3D shape data for each region of interest. If it determines that the 3D shape data generation process has been completed for all regions of interest (S40: YES), the control unit 11 determines whether there are any unprocessed images among the CT images stored in the medical image DB 12a for which the above-mentioned training data generation process has not been performed (S41). If it is determined that there are unprocessed images (S41: YES), the control unit 11 returns to the process of step S31 and executes the processes of steps S31 to S40 for the CT images for which training data generation has not been processed. If it is determined that there are no unprocessed images (S41: NO), the control unit 11 ends the series of processes.
[0054] By the above-described processing, training data used for learning the second learning model M2 can be generated based on the CT images stored in the medical image DB 12a and stored in the training DB 12b. The CT images used to generate the training data for the second learning model M2 may be stored in the medical image DB 12a in advance, or may be obtained from another device via the network N or the portable storage medium 10a.
[0055] The training data generated by the above-described process is used to train the second learning model M2. Training the second learning model M2 can be achieved by a process similar to that shown in FIG. 7 . In the training process for the second learning model M2, the control unit 11 inputs the 3D shape data included in the training data into the second learning model M2 and obtains output data (corrected 3D shape data) from the second learning model M2. The control unit 11 compares the output data with the corrected 3D shape data indicated by the training data and optimizes the parameters of the second learning model M2, for example, using backpropagation algorithms, so that the output data and the corrected 3D shape data are similar to each other. If the second learning model M2 is a statistical shape model using principal component analysis, the control unit 11 performs principal component analysis on the 3D shape data and the corrected 3D shape data included in the training data as a single piece of shape data, calculates a vector representing the statistical tendency of the corrected 3D shape data relative to the 3D shape data for each principal component, and sets the vector as a parameter of the second learning model M2. As a result, when 3D shape data is input, the second learning model M2 can output corrected 3D shape data by performing calculations based on the 3D shape data and the set vector.
[0056] Next, the process of generating training data used for learning the third learning model M3 will be described. Figure 9 is a flowchart showing an example of the process procedure for generating training data for the third learning model M3. The process of generating training data for the third learning model M3 is the same as the process of generating training data for the first learning model M1 up to a certain point, but the process in Figure 9 is the process in Figure 5 with step S51 added instead of step S16 and step S52 added instead of step S18. Explanations of steps that are the same as those in Figure 5 will be omitted.
[0057] After performing steps S11 to S15 in FIG. 5 , the control unit 11 of the information processing device 10 extracts coordinate values (coordinate values on a 2D coordinate plane perpendicular to the aligned imaging direction (X-ray irradiation direction)) of each pixel of the region of interest (left femur) from the CT image of the region of interest (left femur) that has been aligned with the region of interest (left femur) in the X-ray image, thereby generating a silhouette image of the region of interest viewed from the same direction as the imaging direction of the X-ray image (S51). Here, the control unit 11 generates respective silhouette images of the anterior and posterior surfaces of the region of interest (left femur). The control unit 11 determines whether the process of generating silhouette images based on CT images has been completed for all regions of interest to be extracted from the CT image (S17). If it is determined that the process has not been completed (S17: NO), the control unit 11 returns to step S14 and performs the processes of steps S14 to S15 and S51 for the unprocessed regions of interest.
[0058] If the control unit 11 determines that the process of generating silhouette images based on CT images for all regions of interest has been completed (S17: YES), the control unit 11 associates the hip joint frontal X-ray image acquired in step S11 with the front and rear silhouette images of each region of interest generated for each region of interest in step S51, and stores the associated training data in the training DB 12b (S52). This results in training data including silhouette images of the front and rear of the left half of the pelvis, the right half of the pelvis, the left femur, and the right femur, generated from CT images of the region including the pelvis and femur. The control unit 11 then performs step S19. If the control unit 11 determines that there are unprocessed images (S19: YES), the control unit 11 returns to step S11. If the control unit 11 determines that there are no unprocessed images (S19: NO), the control unit 11 terminates the process. Through the above-described process, training data used to train the third learning model M3 can be generated based on the X-ray images and CT images stored in the medical image DB 12a and stored in the training DB 12b.
[0059] The training data generated by the above-described process is used to train the third learning model M3. Training the third learning model M3 can be achieved by a process similar to that shown in FIG. 7 . In the training process for the third learning model M3, the control unit 11 inputs the hip joint X-ray front image included in the training data into the third learning model M3 and obtains output data (silhouette images of the front and back of the target body part) from the third learning model M3. The control unit 11 compares the output data with the correct silhouette images (silhouette images of the front and back of multiple body parts) indicated by the training data, and optimizes parameters such as weights between nodes in the encoder and decoder of the third learning model M3, using, for example, backpropagation, so that the output data approximates the correct silhouette images.
[0060] As a result, by inputting a frontal X-ray image of the hip joint, a third learning model M3 is generated that outputs silhouette images of the front and back surfaces of the left half of the pelvis, the right half of the pelvis, the left femur, and the right femur captured in the X-ray image. The third learning model M3 also uses a segmentation technique that accurately classifies musculoskeletal regions from CT images and a technique that accurately aligns target areas in the X-ray image with target areas in the CT image, and uses training data that allows for highly accurate spatial correspondence between the 3D data of each part obtained by CT and the X-ray image. Therefore, the third learning model M3 can be realized, which is capable of generating highly accurate silhouette images without requiring a large number of cases (training data).
[0061] Next, we will explain the process of generating (estimating) 3D shape data of the left half of the pelvis, the right half of the pelvis, the left femur, and the right femur from a frontal X-ray image of the hip joint including the pelvis and femur of a subject using the first learning model M1 to the third learning model M3 generated as described above. Figure 10 is a flowchart showing an example of the processing procedure for generating 3D shape data of the target region, and Figure 11 is an explanatory diagram showing an example screen.
[0062] The control unit 11 of the information processing device 10 acquires a frontal hip joint X-ray image of a region including the subject's pelvis and left and right femurs, captured with an X-ray device (S61). The control unit 11 acquires the frontal hip joint X-ray image of the subject for which 3D shape data of the pelvis or femurs is to be generated, for example, from electronic medical record data stored in an electronic medical record server. The frontal hip joint X-ray image of the subject may be read from the portable storage medium 10a by the reading unit 16.
[0063] Based on the hip joint frontal X-ray image, the control unit 11 generates front and rear silhouette images of the region of interest in the X-ray image (S62). Here, the control unit 11 inputs the hip joint frontal X-ray image to the third learning model M3 and obtains front and rear silhouette images of the region of interest captured in the X-ray image as output data from the third learning model M3. The regions of interest here are the left half of the pelvis, the right half of the pelvis, the left femur, and the right femur, and the control unit 11 generates front and rear silhouette images of each of the left half of the pelvis, the right half of the pelvis, the left femur, and the right femur.
[0064] The control unit 11 generates depth images of the front and back surfaces of the region of interest in the X-ray image based on the front X-ray image of the hip joint (S63). Here, the control unit 11 inputs the front X-ray image of the hip joint to the first learning model M1 and obtains depth images of the front and back surfaces of the region of interest captured in the X-ray image as output data from the first learning model M1. Here, the control unit 11 also generates depth images of the front and back surfaces of each of the left half of the pelvis, the right half of the pelvis, the left femur, and the right femur.
[0065] The control unit 11 performs masking processing on the depth images of the front and back surfaces of each region of interest generated in step S63 based on the silhouette images of the front and back surfaces of each region of interest generated in step S62 (S64). The silhouette images are images in which background areas other than the region of interest (specifically, the front or back surface of the region of interest) are masked. By performing masking processing using these silhouette images, each pixel in the region other than the region of interest in the depth image of the region of interest is masked to 0 or 1. Thus, a highly accurate depth image can be obtained from which noise contained in regions other than the region of interest has been removed. Note that the third learning model M3 may include the configuration of the first learning model M1 and be configured to predict depth images of the front and back surfaces of the pelvis and femur captured in the input frontal hip joint X-ray image and generate silhouette images of the front and back surfaces of the pelvis and femur from the predicted depth images. In this case, the control unit 11 may be configured to perform a process to mask the background area other than the area of interest on the silhouette image (the silhouette image in which the area of interest in the depth image is detected) obtained as output data from the third learning model M3, thereby generating a mask image in which the background area is masked, and then perform masking processing using the generated mask image.
[0066] The control unit 11 generates 3D shape data for the region of interest by connecting the contours of the front and back depth images of the region of interest that were masked in step S64 (S65). Step S65 is the same process as step S37 in FIG. 8 . Here, the control unit 11 also generates 3D shape data for each region: the left half of the pelvis, the right half of the pelvis, the left femur, and the right femur. The control unit 11 corrects missing portions of the 3D shape data generated in step S65 (S66). Here, the control unit 11 inputs the generated 3D shape data for each region into the second learning model M2 and obtains corrected 3D shape data, in which missing portions in the 3D shape data for each region have been corrected, as output data from the second learning model M2.
[0067] The control unit 11 stores the corrected 3D shape data of each body part obtained by the correction process of step S66, for example, in the electronic medical record data of the electronic medical record server (S67). The control unit 11 then generates a screen displaying the examination results, outputs the screen to, for example, the display unit 15 (S68), and displays the screen on the display unit 15, thereby ending the process. For example, the control unit 11 generates an examination result screen such as that shown in FIG. 11. The screen shown in FIG. 11 displays the subject's identification information (e.g., patient ID, patient name, etc.), the frontal hip joint X-ray image, and the date and time of its capture. Furthermore, the screen shown in FIG. 11 displays 3D shape data of each body part (left half of the pelvis, right half of the pelvis, left femur, right femur) captured in the frontal hip joint X-ray image as an estimation result of the 3D shape data of each body part based on the frontal hip joint X-ray image (3D shape reconstruction data). Although FIG. 11 shows only the 3D shape data of the left femur as an estimation result, the 3D shape data of the left half of the pelvis, the right half of the pelvis, and the right femur may also be displayed.
[0068] The above-described processing can estimate 3D shape data for each part of the hip joint, including the left half of the pelvis, the right half of the pelvis, the left femur, and the right femur, captured in a frontal hip joint X-ray image taken using an X-ray device commonly used in medical institutions. In this embodiment, the 3D shape data for each part generated from the frontal hip joint X-ray image can be presented to a physician. Therefore, physicians can understand the condition of each part using the 3D shape data, which can be used when formulating treatment and surgical plans and simulating artificial joint replacement surgery. The 3D shape data for each part is displayed so that the observation direction of each part can be changed by operating the input unit 14, allowing physicians to observe each part from any direction. The technology disclosed herein generates a front depth image and a rear depth image of the target part from an X-ray image, and the 3D shape data for the target part can be generated by stitching together the front depth image and the rear depth image. Since the 3D shape of the target area cannot be grasped using only depth images of the front or back of the target area, the technology disclosed herein has significant advantages over using depth images of one side (either the front or back).
[0069] In this embodiment, a depth image of the target area is generated (estimated) from an X-ray image of the target area, and 3D shape data of the target area is generated based on the depth image. This makes it possible to generate 3D shape data of the target area from an X-ray image captured from one direction. Furthermore, since the depth image is a 2D image on a 2D coordinate plane orthogonal to the imaging direction, calculations can be performed with fewer computational resources than when calculations are performed based on a 3D image (3D data). Therefore, it is possible to generate 3D shape data of the target area in the X-ray image with high accuracy, even for X-ray images captured over a wide area, without requiring a large amount of computational resources.
[0070] FIG. 12 is an explanatory diagram illustrating the effect of using the information processing device 10 of the present disclosure. Each graph in FIG. 12 , with error on the horizontal axis and frequency on the vertical axis, shows the frequency distribution of the difference (error) between 3D shape data of the femur and pelvis generated from frontal hip joint X-ray images using the information processing device 10 of this embodiment and 3D shape data generated from CT images. The top four graphs show errors for the femur, and the bottom four graphs show errors for the pelvis. Also, from left to right, the graphs show the frequency distribution of errors based on frontal hip joint X-ray images and CT images of patients diagnosed with osteoarthritis of the hip joint at Grade 1 (early stage), Grade 2 (early stage), Grade 3 (advanced stage), and Grade 4 (end-stage) according to the Kellgren-Lawrence (KL) classification. From each graph in FIG. 12 , it can be seen that, regardless of the severity of the patient's condition, the error between the 3D shape data of the femur and pelvis generated from frontal hip joint X-ray images and the 3D shape data generated from CT images was approximately 1 to 2 mm, indicating that the 3D shape data could be estimated with high accuracy.
[0071] The above-described process has been described with reference to an example in which 3D shape data is generated for the pelvis and femur captured in a frontal hip joint X-ray image. However, 3D shape data may be generated only for any target region captured in the X-ray image. For example, a user, such as a doctor, may specify a target region for which 3D shape data is to be generated via the input unit 14, and 3D shape data for the specified target region may be generated. In this case, after processing step S63 in the process of FIG. 10 , the control unit 11 extracts front and rear silhouette images of the specified target region from the front and rear silhouette images of each target region generated in step S62, extracts front and rear depth images of the specified target region from the front and rear depth images of each target region generated in step S63, and performs masking on the front and rear depth images of the target region based on the front and rear silhouette images of the target region. The control unit 11 then performs steps S65 and S66 for only the target region, thereby generating 3D shape data for the target region.
[0072] In this embodiment, the first learning model M1 automatically extracts features of the imaging state of the target area in X-ray images to generate depth images, and then stitches together the depth images of the front and back of the target area to generate 3D shape data of the target area. Therefore, the 3D shape data of the target area can be confirmed without using a device capable of capturing 3D data, such as an X-ray CT scanner or MRI scanner. Therefore, 3D shape data of the target area can be estimated from X-ray images taken during health checkups or at small clinics, making it easy to confirm the 3D shape of the target area. Furthermore, even in Europe and the United States, where CT scans and MRI scans are often not performed, the 3D shape of the target area can be confirmed from X-ray images. Furthermore, compared to CT scans and MRI scans, X-ray examinations (X-ray photography) can be performed in a shorter time and at lower cost, thereby reducing medical costs without increasing the burden on the subject.
[0073] As described above, in this embodiment, it is possible to achieve highly accurate spatial alignment between a target portion in an X-ray image (e.g., the left femur) and a target portion in a CT image. Therefore, by training the first learning model M1 and the third learning model M3 using training data generated from the highly accurately aligned X-ray images and CT images, it is possible to achieve the first learning model M1 and the third learning model M3 with high estimation accuracy without training a large amount of training data.
[0074] In this embodiment, a configuration for estimating 3D shape data of the pelvis and femur from a frontal X-ray image of the hip joint has been described. However, the target regions for 3D shape estimation are not limited to the pelvis and femur, and may be knee joints, ankle joints, wrist joints, elbow joints, shoulder joints, spine (cervical vertebrae, thoracic vertebrae, lumbar vertebrae, sacrum, coccyx), clavicles, ribs, hand bones, foot bones, or specific regions thereof. Similar processing is used to generate training data and learning models M1 to M3 for other regions, enabling estimation of 3D shape data using the learning models M1 to M3.
[0075] In this embodiment, a highly accurate depth image can be generated by performing masking processing on a depth image of the target area generated using the first learning model M1 based on a silhouette image of the target area generated using the third learning model M3. However, in this embodiment, the third learning model M3 is not necessarily required. 3D shape data of the target area may be generated directly using the depth image of the target area generated from an X-ray image using the first learning model M1. Even in this case, the generated 3D shape data can be corrected using the second learning model M2 to generate corrected 3D shape data in which missing areas are corrected.
[0076] In this embodiment, the first learning model M1 may be configured to simultaneously estimate the depth images of the front and rear surfaces of a target region in an input X-ray image so that the contours of the target region in the depth images match (more specifically, so that the contour error is within a predetermined range). That is, in the learning process, the first learning model M1 estimates the depth images of the front and rear surfaces of each target region so that the contours of the target region in the depth images of the front and rear surfaces of each target region in the input X-ray image match, and performs learning based on the comparison results between the estimated depth images and the correct depth images. Furthermore, the first learning model M1 may be configured to simultaneously estimate the depth images of each target region so that there is no inconsistency in the positional relationship between the target regions in the depth images of the multiple target regions in the input X-ray image (more specifically, so that the positional relationship between the target regions is a predetermined positional relationship). That is, in the learning process, the first learning model M1 estimates the depth image of each target part in the input X-ray image so that the positional relationship of the target part in the depth image of each target part in the input X-ray image is a predetermined positional relationship, and performs learning based on the comparison result between the estimated depth image and the correct depth image. As a result, the depth images of the front and back surfaces of each target part in the X-ray image that are consistent with each other can be obtained.
[0077] Furthermore, the third learning model M3 may be configured to simultaneously estimate the front and rear silhouette images of the target part in the input X-ray image so that the error in the contour of the target part in the silhouette image falls within a predetermined range. Furthermore, the third learning model M3 may be configured to simultaneously estimate the silhouette images of each target part in the input X-ray image so that the positional relationship between each target part is a predetermined positional relationship. This allows for the acquisition of front and rear silhouette images of each target part in the X-ray image that are consistent in terms of the positions of the front and rear.
[0078] In this embodiment, a frontal X-ray image of a hip joint taken in accordance with international guidelines for using an X-ray device, in which the distance from the X-ray tube (center of projection) to the imaging surface (FFD: Focus-Film Distance) is approximately 120 cm, was used. However, the FFD of the X-ray image is not limited to 120 cm, and X-ray images taken at other distances may also be used. The first learning model M1 and the third learning model M3 may also be configured to input the FFD of the X-ray image along with the X-ray image. The first learning model M1 and the third learning model M3 may also be configured to set the FFD of the input X-ray image. Note that if the X-ray image is in the Digital Imaging and Communications in Medicine (DICOM) format, the FFD is associated with the tag information of the X-ray image and can be obtained from the tag information of the X-ray image. This configuration enables the generation of depth images and silhouette images taking into account the FFD of the X-ray image input to the first learning model M1 and the third learning model M3. Furthermore, by generating (preparing) the first learning model M1 and the third learning model M3 for each FFD, it is possible to realize the first learning model M1 and the third learning model M3 that are optimal for the FFD.
[0079] In this embodiment, the training data generation process, the learning process of the learning models M1 to M3 using the training data, and the estimation process of the 3D shape data of the target region using the learning models M1 to M3 are not limited to being performed locally by the information processing device 10. For example, a separate information processing device may be provided to perform each of the above processes. Alternatively, a server may be provided to perform the training data generation process and the learning process of the learning models M1 to M3. In this case, the information processing device 10 is configured to transmit X-ray images and CT images used for the training data to the server, and the server generates training data from the X-ray images and CT images, generates the learning models M1 to M3 through a learning process using the generated training data, and transmits them to the information processing device 10. Thus, the information processing device 10 can realize the estimation process of the 3D shape data of the target region using the learning models M1 to M3 acquired from the server. Alternatively, a server may be provided to perform the estimation process of the 3D shape data of the target region using the learning models M1 to M3. In this case, the information processing device 10 is configured to transmit an X-ray image of the subject to a server, and the server performs an estimation process for 3D shape data of the target region using the learning models M1 to M3, and transmits the generated 3D shape data of the target region to the information processing device 10. Even with such a configuration, the same process as in the above-described embodiment is possible, and the same effects can be obtained.
[0080] The information processing device 10 of this embodiment has been described as being configured to estimate 3D shape data of a bone region from an X-ray image. However, this configuration is not limited thereto. It may also be configured to estimate 3D shape data of muscle regions or organs captured in the X-ray image. For example, it may be configured to estimate 3D shape data of muscle regions such as the gluteus maximus, gluteus medius, and hamstrings (biceps femoris, semitendinosus, and semimembranosus) from a frontal X-ray image of a hip joint, or to estimate 3D shape data of organs such as the uterus, bladder, and rectum from a frontal X-ray image of a hip joint. It may also be configured to estimate bone regions such as the ribs, scapula, and clavicle, muscle regions such as the pectoralis major, pectoralis minor, and serratus anterior, and organs such as the heart and lungs from a frontal X-ray image of a chest. Furthermore, in addition to bone regions, muscle regions, and organs, recesses such as holes and depressions in bone regions, muscle regions, and organs may also be used as target regions. That is, it may be configured to generate 3D shape data of recesses in bone regions, muscle regions, and organs as target regions. When a recess is the target area, 3D shape data of the recess can be generated by performing the above-mentioned processing, with the front side of the X-ray image of the subject (the surface closest to you in the direction of X-ray irradiation) as the front side and the back side of the subject (the surface further back in the direction of X-ray irradiation) as the back side of the contour of the recess.
[0081] FIG. 13A is an explanatory diagram showing a modified example of the first learning model M1, FIG. 13B is an explanatory diagram showing a modified example of the second learning model M2, and FIG. 13C is an explanatory diagram showing a modified example of the third learning model M3. For example, when estimating 3D shape data of muscle regions such as the gluteus maximus, gluteus medius, and hamstrings from a frontal hip joint X-ray image, the first learning model M1a to the third learning model M3a shown in FIGS. 13A to 13C are prepared. The first learning model M1a is a model that uses a frontal hip joint X-ray image as input data, performs a calculation to predict depth images of the anterior and posterior surfaces of muscle regions such as the gluteus maximus, gluteus medius, and hamstrings in the frontal hip joint X-ray image, and outputs the calculation results. The second learning model M2a is a model that uses 3D shape data of a target region (e.g., the gluteus maximus) as input data, performs a calculation to correct missing parts in the 3D shape data, and outputs the calculation results (corrected 3D shape data). The third learning model M3a is a model that uses a frontal X-ray image of a hip joint as input data, performs calculations to recognize the front and back surfaces of muscle regions such as the gluteus maximus, gluteus medius, and hamstrings in the frontal X-ray image of the hip joint, and outputs the calculation results (silhouette images of the front and back surfaces of each muscle region). The first learning model M1a to the third learning model M3a have the same configuration as the first learning model M1 to the third learning model M3 described in the first embodiment, and can be realized by the same learning process, except that the training data used for learning is different.
[0082] FIG. 14A is an explanatory diagram showing another variation of the first learning model M1, FIG. 14B is an explanatory diagram showing another variation of the second learning model M2, and FIG. 14C is an explanatory diagram showing another variation of the third learning model M3. For example, when estimating 3D shape data of organs such as the heart and lungs from a frontal chest X-ray image, the first learning model M1b to the third learning model M3b shown in FIGS. 14A to 14C are prepared. The first learning model M1b is a model that uses a frontal chest X-ray image as input data, performs calculations to predict depth images of the front and back surfaces of organs such as the heart and lungs in the frontal chest X-ray image, and outputs the calculation results. The second learning model M2b is a model that uses 3D shape data of a target region (e.g., the heart) as input data, performs calculations to correct missing areas in the 3D shape data, and outputs the calculation results (corrected 3D shape data). The third learning model M3b is a model that uses a frontal chest X-ray image as input data, performs calculations to recognize the front and back surfaces of each of the organs, such as the heart and lungs, in the frontal chest X-ray image, and outputs the calculation results (silhouette images of the front and back surfaces of each organ). The first learning model M1b to the third learning model M3b also have the same configuration as the first learning model M1 to the third learning model M3, and can be realized by the same learning process, except that the training data used for learning is different.
[0083] As described above, the first to third learning models can be generated by generating training data for the first to third learning models based on X-ray images and CT images of a region for which 3D shape data is to be generated, and then performing a learning process using the generated training data. By using such first to third learning models, 3D shape data of any region can be generated (estimated) with high accuracy from an X-ray image of the region.
[0084] (Embodiment 2) In the above-described embodiment 1, an X-ray image (simple X-ray image) and a CT image of the same subject are taken of the same imaging target, and the information processing device 10 performs a process of aligning a bone region (target region) in the X-ray image with a bone region (target region) in the CT image. The information processing device 10 of this embodiment has the same configuration as the information processing device 10 of embodiment 1, and therefore a description of the configuration will be omitted.
[0085] FIG. 15 is a flowchart showing an example of the alignment process procedure, and FIGS. 16A to 17B are explanatory diagrams of the alignment process. The alignment process shown in FIG. 15 can be the process of step S15 in FIGS. 5 and 9. Therefore, in this embodiment, the control unit 11 of the information processing device 10 executes the process of FIG. 15 after the process of step S14 in FIGS. 5 and 9, and then executes the process of step S16 in FIG. 5 or step S51 in FIG. 9. The following describes an example in which the pelvis is the region of interest and the pelvis in a frontal hip joint X-ray image is aligned with the pelvis in a CT image. However, the bone region used for alignment is not limited to the pelvis. Any bone region captured in the X-ray image and the CT image can be used for alignment. For example, in a frontal hip joint X-ray image, in addition to the pelvis, the femur, the proximal femur, etc. may be used for alignment.
[0086] 5 and 9 , the control unit 11 identifies the region of interest (here, the pelvis) in the X-ray image (here, the frontal X-ray image of the hip joint) acquired in step S11 (S71). The process of identifying the region of interest in the X-ray image can be performed, for example, by pattern matching using a template that indicates the shape of the region of interest. Alternatively, for example, the process can be performed using a learning model that has been machine-learned to output the region of interest in an X-ray image when the X-ray image is input. As a result, for example, the pelvis region indicated by the solid line in the X-ray image shown in FIG. 16A is identified.
[0087] Next, the control unit 11 generates a DRR (Digital Reconstructed Radiograph) image, which is a projection image of a three-dimensional region of the region of interest based on the region of interest in the CT image extracted in step S14 (S72). The DRR image is an X-ray image obtained by projection simulation from a three-dimensional region (region of interest) of a specific region in the CT image. Specifically, as shown in FIG. 16B , the control unit 11 places the region of interest in the CT image (a three-dimensional CT image of the pelvis) in the three-dimensional virtual space of the X-ray imaging system under predetermined projection conditions (position and angle relative to a virtual X-ray source), and generates a DRR image (projection image) by projecting the region of interest from the virtual X-ray source onto a two-dimensional X-ray imaging surface. Here, because the imaging conditions (e.g., the subject's posture, joint flexion angle, etc.) are different between the X-ray image and the CT image, the contour of the region of interest in the DRR image generated from the CT image does not match the contour of the region of interest in the X-ray image. 17A , the contour P1 of the region of interest in the X-ray image is indicated by a solid line, and the contour P2 of the region of interest in the DRR image is indicated by a dashed line. In this embodiment, the control unit 11 updates the projection conditions of the region of interest in the CT image to identify the projection conditions that maximize the correlation value between the contour P1 of the region of interest in the X-ray image and the contour P2 of the region of interest in the DRR image. By using these projection conditions, a DRR image can be obtained in which the region of interest in the DRR image is accurately aligned with the region of interest in the X-ray image, as shown in FIG. 17B .
[0088] Therefore, the control unit 11 calculates a correlation value between the contour of the region of interest in the X-ray image identified in step S71 and the contour of the region of interest in the DRR image generated in step S72 (S73), and determines whether the calculated correlation value is maximum (S74). If the control unit 11 determines that the correlation value is not maximum (S74: NO), it updates the projection conditions used to generate the DRR image from the CT image of the region of interest (S75) and repeats the processes of steps S72 to S74 under the updated projection conditions. The control unit 11 repeats the processes of steps S72 to S75 until it determines that the calculated correlation value is maximum (S74: YES), that is, if a DRR image with the maximum correlation value can be generated, it identifies the projection conditions used at that time (S76).
[0089] In this embodiment, the control unit 11 can perform the processes of steps S72 to S75 using, for example, the method described in a paper titled "3D-2D Registration in Mobile Radiographs: Algorithm Development and Preliminary Clinical Evaluation" by the present inventor, Yoshito Otake et al. This paper defines the achievement of registration between the contour of a target region (a bone region, in this case, the pelvis) in a DRR image generated from a 3D region of the target region in a CT image and the contour of the target region in an actual X-ray image as "maximizing the correlation between the gray-scale gradient intensity images of the X-ray image and the DRR image." The paper also discloses a method for determining a configuration (projection conditions, specifically, the 3D position and angle of the target region relative to the imaging system) that maximizes the correlation using a covariance matrix adaptation evolution strategy (CMA-ES). Therefore, the control unit 11 can identify a DRR image that maximizes the correlation between the contour of the target region in the X-ray image and the contour of the DRR image using the covariance matrix adaptation evolution strategy, thereby identifying the projection conditions for the identified DRR image. Under these projection conditions, a DRR image of the region of interest can be generated in which a contour P2 of the region of interest in the DRR image is aligned with a contour P1 of the region of interest in the X-ray image with high accuracy, as shown in Fig. 17B. Note that in step S74, the control unit 11 may be configured to determine whether the calculated correlation value is equal to or greater than a predetermined value, and if it is determined that the calculated correlation value is equal to or greater than the predetermined value, proceed to the processing of step S76.
[0090] The control unit 11 then proceeds to step S16 of FIG. 5 or step S51 of FIG. 9 . In step S16 of FIG. 5 , the control unit 11 generates a depth image of the region of interest by extracting coordinate values of each pixel in the region of interest in the projection direction according to the projection conditions specified in step S76 from the CT image of the region of interest. This generates a depth image of the region of interest under the same projection conditions (imaging direction) as the imaging conditions (position and angle relative to the X-ray source) of the region of interest in the X-ray image, thereby obtaining a depth image of the region of interest that is accurately aligned with the region of interest in the X-ray image. Furthermore, in step S51 of FIG. 9 , the control unit 11 extracts coordinate values (2D coordinate values) of each pixel of the region of interest in a 2D coordinate plane orthogonal to the projection direction according to the projection conditions specified in step S76 from the CT image of the region of interest, thereby generating a silhouette image of the region of interest. This generates a silhouette image of the area of interest viewed from the same imaging direction as the imaging conditions of the area of interest in the X-ray image (position and angle relative to the X-ray source), resulting in a silhouette image of the area of interest that is accurately aligned with the area of interest in the X-ray image.
[0091] By the above-mentioned processing, training data is generated that associates an X-ray image of the area of interest (the area of interest being a bone region) with a depth image of the area of interest that is aligned with high precision to the area of interest in the X-ray image, and training data is also generated that associates an X-ray image of the area of interest with a silhouette image of the area of interest that is aligned with high precision to the area of interest in the X-ray image, and the training data is stored in training DB12b.
[0092] Since bone regions are hard tissues, they are not likely to deform when X-ray images and CT images are taken. Therefore, by performing a registration process by optimizing the conditions (imaging direction relative to the region of interest) for generating depth images and silhouette images from CT images as in this embodiment, the contour of the region of interest (bone region) in the X-ray image can be accurately aligned with the contour of the region of interest in the DRR image generated from the CT image. Furthermore, the information processing device 10 of this embodiment is capable of executing the process shown in FIG. 7 . By performing the learning process of the first learning model M1 and the third learning model M3 using the training data generated as described above, the first learning model M1 can predict a depth image with high accuracy from an X-ray image, and the third learning model M3 can predict a silhouette image with high accuracy from an X-ray image. Furthermore, the information processing device 10 of this embodiment is capable of executing the processing shown in Figure 10, and by using the first learning model M1 and the third learning model M3 generated as described above, the depth images of the front and back of the area of interest can be estimated with high accuracy using the depth images and silhouette images predicted with high accuracy from the X-ray images, and as a result, it is possible to estimate the 3D shape data of the area of interest with high accuracy.
[0093] In the above-described process shown in FIG. 5 , the registration of the region of interest between the X-ray image and the CT image is performed based on the region of interest extracted from the CT image in step S14. However, this is not a limitation. For example, before extracting the region of interest from the CT image, a bone region including, for example, the pelvis and femur may be extracted, and the region of interest may be extracted from the extracted bone region. The registration of the bone region between the X-ray image and the CT image may be performed based on the bone region extracted from the CT image. In this case, after performing step S13 in FIG. 5 , the control unit 11 of the information processing device 10 extracts the bone region from the CT image, performs the registration process shown in FIG. 15 based on the extracted bone region, and then performs steps S14 and S16 based on the CT image after the registration process. Specifically, the control unit 11 extracts data of the region of interest from the CT image (the CT image of the bone region) after the alignment process (S14), and generates a depth image from the CT image of the extracted region of interest in the imaging direction according to the projection conditions specified in step S76 (S16). In the process shown in FIG. 15 , steps S71 to S76 are performed using the bone region extracted from the CT image as the region of interest. The control unit 11 then proceeds to step S17. This process also generates training data that associates an X-ray image with a depth image of the region of interest that is aligned with high accuracy to the region of interest in the X-ray image, and stores the training data in the training DB 12b. By executing the learning process of the first learning model M1 using the training data generated in this manner, the first learning model M1 can be realized, which can predict a high-accuracy depth image from an X-ray image. Furthermore, by using the first learning model M1 generated in this manner, it is possible to predict with high accuracy the depth image of the target area from the X-ray image, and the highly accurately predicted depth image makes it possible to estimate with high accuracy the 3D shape data of the area of interest.
[0094] The process shown in FIG. 9 is not limited to aligning the region of interest between the X-ray image and the CT image based on the region of interest extracted from the CT image in step S14. The process shown in FIG. 9 may also be configured to align the bone region between the X-ray image and the CT image based on the bone region extracted from the CT image. In this case, after processing step S13 in FIG. 9 , the control unit 11 of the information processing device 10 extracts the bone region from the CT image, performs the alignment process shown in FIG. 15 based on the extracted bone region, and performs steps S14 and S51 based on the CT image after the alignment process (the CT image of the bone region). That is, the control unit 11 extracts data of the region of interest from the CT image after the alignment process (S14), and generates a silhouette image from the CT image of the extracted region of interest in the imaging direction according to the projection conditions identified in step S76 (S51). Even in this process, training data is generated in which an X-ray image and a silhouette image of the region of interest that is aligned with high accuracy with respect to the region of interest in the X-ray image are associated, and the training data is stored in the training DB 12 b. By executing the learning process of the third learning model M3 using the training data generated in this manner, it is possible to realize the third learning model M3 that can predict a silhouette image with high accuracy from an X-ray image.
[0095] The information processing device 10 of this embodiment can perform the same processes as those of the first embodiment described above except for the above-described alignment process, and can achieve the same effects as those of the first embodiment. Furthermore, this embodiment can perform highly accurate spatial alignment between a region of interest in an X-ray image and a region of interest in a CT image. Therefore, by using such X-ray images and depth images and silhouette images generated from CT images aligned with the X-ray images with high accuracy as training data, a first learning model M1 that predicts a depth image from an X-ray image with high accuracy and a third learning model M3 that predicts a silhouette image from an X-ray image with high accuracy can be realized without requiring a large number of cases (training data). Furthermore, the modifications described in the first embodiment can also be applied to this embodiment as appropriate.
[0096] (Embodiment 3) In this embodiment, an information processing device that generates 3D shape data of a target region based on multiple X-ray images of the target region captured from multiple directions will be described. While the above-described embodiments 1 and 2 are configured to enable the generation of 3D shape data of a target region based on X-ray images captured from one direction, this embodiment is configured to generate 3D shape data of a target region based on X-ray images captured from multiple directions. In other words, the present disclosure includes a configuration in which 3D shape data of a target region is generated using additional X-ray images captured from a direction different from the one direction in embodiments 1 and 2. The information processing device 10 of this embodiment has the same configuration as the information processing device 10 of embodiment 1, and therefore a description of the configuration will be omitted.
[0097] Fig. 18 is a flowchart showing an example of the processing procedure for generating 3D shape data of a target region in embodiment 3, and Figs. 19A to 19C are explanatory diagrams of the 3D shape data generation processing. The processing shown in Fig. 18 is the processing shown in Fig. 10 with steps S81 to S82 added instead of step S61 and steps S83 to S84 added instead of step S65. Explanations of the same steps as in Fig. 10 will be omitted.
[0098] In the information processing device 10 of this embodiment, the control unit 11 acquires multiple X-ray images of a target area from multiple directions (S81). The multiple directions may be, for example, multiple directions from the front of the target area, the side, and oblique angles (first to fourth oblique angles). The control unit 11 may acquire the multiple X-ray images from an electronic medical record server, from a portable storage medium 10a using the reading unit 16, or directly from an X-ray device (X-ray device). The control unit 11 stores the acquired multiple X-ray images in the storage unit 12 and reads one of the stored X-ray images as the processing target (S82).
[0099] The control unit 11 performs the processes of steps S62 to S64 on the read X-ray image to generate silhouette images and depth images of the front and back surfaces of the region of interest (target region) (S62, S63), and then performs masking processing on the generated depth images of the front and back surfaces of the region of interest based on the silhouette images of the front and back surfaces of the region of interest (S64). This generates depth images (depth images of the front and back surfaces) in which areas other than the region of interest are masked. For example, as shown on the left side of FIG. 19A , based on an X-ray image captured with the first direction as the imaging direction, depth images of the front and back surfaces of the region of interest in the first direction are generated, as shown on the right side of FIG. 19A .
[0100] The control unit 11 determines whether processing has been completed for all X-ray images acquired in step S81 (S83). If it determines that processing has not been completed (S83: NO), the process returns to step S82 and repeats the processing of steps S82 and S62 to S64 for unprocessed X-ray images. As a result, for example, based on an X-ray image captured with the second direction as the imaging direction as shown on the left side of Fig. 19B, depth images of the front and back surfaces of the region of interest in the second direction are generated as shown on the right side of Fig. 19B. The control unit 11 processes all acquired X-ray images to generate depth images of the front and back surfaces for the imaging direction of each X-ray image.
[0101] If it is determined that processing has been completed for all X-ray images (S83: YES), the control unit 11 generates 3D shape data of the region of interest by stitching together the front and back depth images generated from each X-ray image (S84). For example, when stitching together depth images generated from two X-ray images as shown in FIGS. 19A and 19B, the control unit 11 generates 3D shape data (referred to as "first 3D shape data") representing the surface shape of the region of interest from the front and back depth images generated from one X-ray image. Again, the 3D shape data is generated using software such as 3D CAD, and the generated 3D shape data does not need to represent the front and back surfaces of the region of interest in a continuous manner. Similarly, the control unit 11 generates 3D shape data (referred to as "second 3D shape data") representing the surface shape of the region of interest from the front and back depth images generated from the other X-ray image. The control unit 11 then projects, for example, a first surface shape represented by the first 3D shape data onto a second surface shape represented by the second 3D shape data, optimizes projection parameters to minimize a difference (e.g., distance error) between the projected first surface shape and the second surface shape, and identifies the projection parameters that minimize the difference. The projection parameters include the projection position, projection direction (orientation), and projection scale (magnification ratio) of the first surface shape relative to the second surface shape. If the control unit 11 identifies the optimal projection parameters, it connects the first surface shape and the second surface shape projected according to the identified projection parameters to generate 3D shape data of the region of interest as shown in FIG. 19C . In FIG. 19C , the first surface shape is indicated by a thick solid line, and the second surface shape is indicated by a thick dashed line. By connecting the two 3D shape data, 3D shape data such as that shown in FIG. 19C can be generated. For the joint between the first surface shape and the second surface shape, one of the values may be adopted, or the average value of the two may be adopted, or a weighted average value weighted according to the shooting direction may be adopted.
[0102] The control unit 11 then performs the processes from step S66 onward on the 3D shape data generated in step S84. This generates and stores corrected 3D shape data in which missing portions in the 3D shape data have been corrected. Through the above-described process, 3D shape data of the region of interest (target region) can be generated (estimated) with high accuracy based on depth images of the front and back of the region of interest generated from X-ray images captured from multiple directions. As shown in FIGS. 19A and 19B , recesses such as dents or holes that cannot be identified in X-ray images captured in the first direction (or second direction) can be identified in X-ray images captured in the second direction (or first direction), thereby generating highly accurate 3D shape data. Furthermore, the 3D shape data of the region of interest is generated by appropriately stitching together depth images of the front and back from multiple directions, thereby obtaining 3D shape data in which the shape of the region of interest is estimated with high accuracy even at the seams. Also in this embodiment, the depth image is a 2D image on a 2D coordinate plane orthogonal to the imaging direction, and therefore can be calculated with fewer computational resources than when calculations are performed based on a 3D image.
[0103] 19A to 19C , a process for generating 3D shape data of a target region based on depth images generated from two X-ray images captured from two orthogonal directions has been described. However, the intersection angle between the two directions is not limited to 90 degrees. Alternatively, 3D shape data of the target region may be generated based on three or more X-ray images captured from three or more directions. In this case, depth images of the front and back of the target region are generated based on the respective X-ray images, and the surface shapes (3D shape data) of the target region shown in the generated front and back depth images are aligned and stitched together to generate 3D shape data of the target region.
[0104] (Embodiment 4) The above-described embodiments 1 to 3 are configured to generate 3D shape data of a target region (region of interest) by stitching together a depth image of the front and a depth image of the back of the target region. In this embodiment, an information processing device is described that generates 3D shape data of a target region based on at least one of a depth image of the front and a depth image of the back of the target region and a thickness image of the target region. The information processing device 10 of this embodiment has a similar configuration to the information processing device 10 of embodiment 1, and therefore a detailed description of the configuration will be omitted. Note that the storage unit 12 of the information processing device 10 of this embodiment stores a fourth learning model in addition to the data shown in FIG. 1.
[0105] FIG. 20 is an explanatory diagram of a thickness image. The thickness image is an image (2D image) on a 2D coordinate plane orthogonal to the imaging direction (X-ray irradiation direction) of the target region, and the pixel value of each pixel in the image indicates the thickness (width) of the target region on the X-ray path from the X-ray tube (imaging center) to the imaging surface in the imaging direction. FIG. 20 shows the pelvis and femur as viewed from the left side, and illustrates the state in which an X-ray image is formed on the imaging surface on the back side of the subject (right side in FIG. 20) by X-rays irradiated from a virtual X-ray source located on the front side (left side in FIG. 20) of the subject (image subject). The target region here is the pelvis. The solid line in FIG. 20 indicates the X-ray path from the virtual X-ray source to pixel A on the imaging surface. As indicated by arrow C1, the distance (more precisely, the horizontal distance) between a point on the front surface and a point on the back surface of the target region (pelvis) on the X-ray path is the thickness (thickness value) corresponding to pixel A, which is the pixel value of pixel A in the thickness image. As shown by arrow A1 in Fig. 20, the distance from the virtual X-ray source to a point on the front surface indicates the depth (depth value) of that point, which becomes the pixel value of pixel A in the depth image of the front surface of the target area. Also, as shown by arrow B1 in Fig. 20, the distance from the virtual X-ray source to a point on the back surface indicates the depth of that point, which becomes the pixel value of pixel A in the depth image of the back surface of the target area.
[0106] FIG. 21 is an explanatory diagram showing an example configuration of the fourth learning model M4. The fourth learning model M4 in FIG. 21 is a trained model that uses a frontal hip joint X-ray image as input data, performs calculations to predict thickness images of the pelvis and femur in the frontal hip joint X-ray image, and outputs the calculation results. The fourth learning model M4 of this embodiment predicts and outputs thickness images for the left and right halves of the pelvis and the left and right femur in the input frontal hip joint X-ray image. The fourth learning model M4 can be realized with a configuration similar to that of the first learning model M1. The fourth learning model M4 is a neural network having an encoder and a decoder, and may be configured using any algorithm or a combination of multiple algorithms as long as it has a configuration including an encoder and a decoder.
[0107] The fourth learning model M4 is generated by machine learning using training data that associates training X-ray images (frontal X-ray images of the hip joint) with thickness images of each training target region (left and right halves of the pelvis, left and right femurs). The training thickness images are generated for each target region from CT images obtained by using a CT scanner to capture an area including the pelvis and femur of the same subject as in the training X-ray images. The thickness images are images in which the subject's anterior-posterior thickness (thickness value) of each point on the target region is assigned to each pixel in a two-dimensional image of the subject viewed from the front. Such training data is stored, for example, in a training DB 12b provided in the memory unit 12 and used during the learning process.
[0108] The fourth learning model M4 learns to output correct thickness images (thickness images of each target region) included in the training data when an X-ray image included in the training data is input. During the learning process, the fourth learning model M4 performs calculations based on the input X-ray image and generates, as output data, thickness images of the target regions in the input X-ray image. The fourth learning model M4 then compares the generated thickness images of the target regions with the correct thickness images indicated in the training data and optimizes parameters used in the calculation process so that the generated thickness images approximate the correct thickness images. For example, the fourth learning model M4 optimizes parameters in the decoder and encoder using backpropagation, steepest descent, or the like. This results in a fourth learning model M4 that, when a frontal X-ray image of a hip joint is input, outputs thickness images of the target regions (pelvis and femur) in the X-ray image.
[0109] Figure 22 is a flowchart showing an example of a processing procedure for generating training data for the fourth learning model M4. The training data for the fourth learning model M4 can be generated by the same process as the training data for the first learning model M1 shown in Figure 5, so only the different processes will be described. The process shown in Figure 22 is the same as the process shown in Figure 5, except that step S91 is added between steps S16 and S17, and step S92 is added instead of step S18. Explanations of the same steps as in Figure 5 will be omitted.
[0110] When generating training data for the fourth learning model M4, the control unit 11 of the information processing device 10 also performs steps S11 to S16. That is, the control unit 11 aligns the region of interest in the X-ray image and CT image to be processed, and generates a depth image of the region of interest from the CT image aligned with the X-ray image. Note that in this embodiment, the control unit 11 does not need to generate depth images of both the front and back of the region of interest, and may be configured to generate a depth image of at least one of the front and back of the region of interest.
[0111] The control unit 11 generates a thickness image of the region of interest from the CT image aligned with the X-ray image (S91). Here, the control unit 11 generates the thickness image of the region of interest by calculating the pixel value of each pixel in the region of interest in the same direction as the X-ray image capture direction based on the CT image. For example, since the X-ray image shown in FIG. 6 includes a pelvis and femur, the control unit 11 generates thickness images of the left half of the pelvis, the right half of the pelvis, the left femur, and the right femur as the region of interest. The thickness at each position of the region of interest may be calculated, for example, by identifying the front and rear positions of the target region on the X-ray path corresponding to each position and calculating the difference between the depth at the front and rear positions. Alternatively, the thickness at each position of the region of interest may be determined by identifying the region of the target region on the X-ray path corresponding to each position based on the CT value (HU value) and then determining the width (thickness) of the identified region.
[0112] In step S17, the control unit 11 determines whether the process of generating depth images and thickness images for all regions of interest has been completed. If the control unit 11 determines that the process has been completed (S17: YES), the control unit 11 associates the X-ray images acquired in step S11, the depth images of the front and / or back surfaces of each region of interest generated in step S16, and the thickness images of each region of interest generated in step S91, and stores them as training data in the training DB 12b (S92). Through the above-described process, in addition to the training data used to train the first learning model M1, training data used to train the fourth learning model M4 can be generated and stored in the training DB 12b.
[0113] The fourth learning model M4 can be trained by a process similar to that shown in FIG. 7 . In the training process for the fourth learning model M4, the control unit 11 inputs a frontal hip joint X-ray image included in the training data into the fourth learning model M4 and obtains output data (thickness images of the target region) from the fourth learning model M4. The control unit 11 compares the output data with the correct thickness image indicated by the training data and optimizes parameters in the fourth learning model M4 using backpropagation algorithms or the like so that the output data approximates the correct thickness image. As a result, by inputting a frontal hip joint X-ray image, the fourth learning model M4 is generated, which outputs thickness images of the left half of the pelvis, the right half of the pelvis, the left femur, and the right femur captured in the X-ray image. The fourth learning model M4 also trains using training data in which 3D data of each region obtained by CT and X-ray images can be spatially associated with high accuracy using segmentation and registration techniques. Therefore, it is possible to realize a fourth learning model M4 (learning model for outputting thickness images) that can generate highly accurate thickness images without requiring a large number of cases (training data).
[0114] Fig. 23 is a flowchart showing an example of a processing procedure for generating 3D shape data of a target region in embodiment 4. The generation processing of this embodiment is similar to the generation processing of embodiment 1 shown in Fig. 10, so only the different processing will be described. The processing shown in Fig. 23 is the same as the processing shown in Fig. 10, except that step S101 is added between steps S63 and S64, and steps S102 to S103 are added instead of step S65. Descriptions of the same steps as in Fig. 10 will be omitted.
[0115] In this embodiment, the control unit 11 of the information processing device 10 also performs the processes of steps S61 to S63. Note that in this embodiment, the control unit 11 may be configured to generate a depth image of at least one of the front and back surfaces of the region of interest in step S63. After the processes of steps S61 to S63, the control unit 11 generates a thickness image of the region of interest in the hip joint front X-ray image based on the X-ray image (S101). Here, the control unit 11 inputs the hip joint front X-ray image into the fourth learning model M4 and obtains a thickness image of the region of interest captured in the X-ray image as output data from the fourth learning model M4. Here, the control unit 11 also generates thickness images of the left half of the pelvis, the right half of the pelvis, the left femur, and the right femur.
[0116] In step S64, the control unit 11 performs masking on the depth image of the front surface and / or the depth image of the back surface of the region of interest generated in step S63. After processing step S64, the control unit 11 performs masking on the thickness image of each region of interest generated in step S101 based on a silhouette image (S102). Step S102 is the same process as step S64. By processing step S102, each pixel in the region other than the region of interest in the thickness image of the region of interest is masked to 0 or 1, thereby obtaining a thickness image from which noise contained in the region other than the region of interest has been removed. Note that the processes of steps S62 to S63 and S101 may be performed in reverse order, and the processes of steps S64 and S102 may be performed in reverse order. Furthermore, the processes of steps S101 and S102 may be performed after the processes of steps S63 and S64.
[0117] Next, the control unit 11 generates 3D shape data of the region of interest based on the depth image of the front surface or the depth image of the rear surface of the region of interest obtained in steps S63 and S64 and the thickness image of the region of interest obtained in steps S101 and S102 (S103). Here, the control unit 11 generates the 3D shape data of the region of interest by reproducing thicknesses corresponding to each position on the front surface (or the rear surface) based on the depth image of the front surface (or the depth image of the rear surface) and the thickness image using a 3D CAD system or the like. Note that if depth images of both the front and rear surfaces of the region of interest are generated in steps S63 and S64, the control unit 11 may also generate 3D shape data of the region of interest based on the depth image and the thickness image of the front surface and the thickness image. For example, the control unit 11 generates first 3D shape data that reproduces thicknesses corresponding to each position on the front surface based on the depth image and the thickness image of the front surface, and generates second 3D shape data that reproduces thicknesses corresponding to each position on the rear surface based on the depth image and the thickness image of the rear surface. The control unit 11 then projects, for example, a first surface shape represented by the first 3D shape data onto a second surface shape represented by the second 3D shape data, and optimizes projection parameters to minimize a difference (e.g., distance error) between the projected first surface shape and the second surface shape. The projection parameters include a projection position, a projection direction (orientation), a projection scale (magnification ratio), etc., of the first surface shape relative to the second surface shape. When the control unit 11 identifies optimal projection parameters, it generates 3D shape data of the region of interest by stitching together the first surface shape and the second surface shape projected according to the identified projection parameters. Note that, for the seam between the first surface shape and the second surface shape, one of the values may be used, or an average value of the two values may be used, or a weighted average value weighted according to the front or rear side may be used.
[0118] The control unit 11 then performs the processes from step S66 onward. As a result, the control unit 11 corrects the 3D shape data generated based on the depth image and thickness image of the region of interest, stores the data in the electronic medical record server, and presents the estimated 3D shape data on a screen such as that shown in FIG. 11 . Through the above-described processes, in this embodiment, 3D shape data of the region of interest can be generated based on a depth image and a thickness image of the front surface of the region of interest, and 3D shape data of the region of interest can be generated based on a depth image and a thickness image of the rear surface of the region of interest. Furthermore, 3D shape data of the region of interest can be generated based on a depth image and a thickness image of the front surface and a depth image and a thickness image of the rear surface of the region of interest. In this case, 3D shape data can be generated with higher accuracy than when using a depth image of either the front or rear surface.
[0119] The configuration of this embodiment can be applied to the above-described first to third embodiments, and similar effects can be obtained even when applied to the first to third embodiments. For example, when the configuration of this embodiment is applied to the third embodiment, the information processing device 10 generates depth images and thickness images of the front or back surface of the target region based on multiple X-ray images of the target region captured from multiple directions, thereby generating 3D shape data of the target region. The information processing device 10 also generates a single 3D shape data set for the target region by stitching together the multiple 3D shape data sets generated from the X-ray images from each direction. For example, after processing step S63 of FIG. 18 , the control unit 11 generates a thickness image of the target region in the X-ray image, and after processing step S64, performs masking on the generated thickness image of the target region based on the silhouette image of the target region generated in step S62. Then, in step S84, the control unit 11 generates 3D shape data of the target region based on the depth images and thickness images of the front or back surface of the target region for each direction, and stitches together the 3D shape data sets generated for each direction to generate a single 3D shape data set. Except for the process of generating 3D shape data based on a depth image and a thickness image of the front or back surface generated from one X-ray image, the same processes as those in the third embodiment can be executed. In the present embodiment, the modified examples described in the first to third embodiments can also be applied as appropriate.
[0120] (Embodiment 5) In this embodiment, the target site is a recess formed in a site that can be photographed with X-ray images and CT images, and an information processing device that generates 3D shape data of the recess will be described. The information processing device 10 of this embodiment has the same configuration as the information processing device 10 of embodiment 1, so a detailed description of the configuration will be omitted.
[0121] FIG. 24 is an explanatory diagram showing an example of a recess. Similar to FIG. 20, FIG. 24 shows a recess formed in the pelvis with a thick solid line. The recess shown in FIG. 24 is the acetabulum (acetabulum) of the pelvis, called a socket, which has a structure that encases the femoral head. When such a recess is used as the target site, 3D shape data of the recess can be generated by defining the front surface of the subject (the surface closest to the subject in the X-ray irradiation direction) as the front surface and the back surface of the subject (the surface farther from the X-ray irradiation direction) as the rear surface at the contour (boundary) of the recess shown by the thick solid line. In the example of FIG. 24, the area on the front side of the boundary of the acetabulum is the front surface, and the distance (depth) from the virtual X-ray source to a point on the front surface, as indicated by arrow A2, is the pixel value of pixel B in the depth image of the front surface of the acetabulum. The area on the back side of the boundary line of the acetabulum becomes the posterior surface, and the distance (depth) from the virtual X-ray source to a point on the posterior surface, as indicated by arrow B2, becomes the pixel value of pixel B in the depth image of the posterior surface of the acetabulum. By stitching together the depth images of the front and posterior surfaces generated in this way, 3D shape data of the acetabulum can be generated.
[0122] FIG. 25A is an explanatory diagram showing a modified example of the first learning model M1, FIG. 25B is an explanatory diagram showing a modified example of the second learning model M2, and FIG. 25C is an explanatory diagram showing a modified example of the third learning model M3. When estimating 3D shape data of the acetabulum from a frontal hip joint X-ray image, the first learning model M1c to the third learning model M3c shown in FIGS. 25A to 25C are prepared. The first learning model M1c is a model that uses a frontal hip joint X-ray image as input data, performs a calculation to predict depth images of the anterior and posterior surfaces of the left and right acetabulum in the frontal hip joint X-ray image, and outputs the calculation results. The second learning model M2c is a model that uses 3D shape data of the acetabulum as input data, performs a calculation to correct missing areas in the 3D shape data, and outputs the calculation results (corrected 3D shape data). The third learning model M3c is a model that uses a frontal X-ray image of a hip joint as input data, performs calculations to recognize the front and rear surfaces of the left and right acetabulum in the frontal X-ray image of the hip joint, and outputs the calculation results (silhouette images of the front and rear surfaces of the acetabulum). The first learning model M1c to the third learning model M3c have the same configuration as the first learning model M1 to the third learning model M3 described in the first embodiment, and can be realized by the same learning process, except that the training data used for learning is different.
[0123] Training data for the first learning model M1c can be generated by a process similar to that shown in FIG. 5 . In step S14 of FIG. 5 , the control unit 11 extracts data on the left and right acetabula (recesses) from the CT image as data on the region of interest (target region). In step S16, the control unit 11 extracts coordinate values of each pixel in the acetabulum region in the same direction as the X-ray image from the CT image that has been aligned with the X-ray image, thereby generating depth images of the front and back of the acetabulum. Here, depth images of the front and back of the left and right acetabula are generated. As a result, the training data used for learning the first learning model M1c is stored in the training DB 12b.
[0124] The learning of the first learning model M1c can be realized by the same processing as that of Fig. 7. The control unit 11 performs learning using a frontal X-ray image of the hip joint and depth images of the front and rear surfaces of the left and right acetabulum as training data, and can generate the first learning model M1c that inputs a frontal X-ray image of the hip joint and outputs depth images of the front and rear surfaces of the left and right acetabulum captured in the X-ray image.
[0125] Training data for the second learning model M2c can be generated by a process similar to that shown in FIG. 8 . In step S34 of FIG. 8 , the control unit 11 extracts data on the left and right acetabulum (recesses) from the CT image as data on the region of interest (target region). In steps S35 and S36, the control unit 11 generates depth images of the front and rear of the acetabulum from the acetabulum data (CT image). In step S37, the control unit 11 generates 3D shape data of the acetabulum by connecting the contours of the acetabulum in the generated depth images of the front and rear of the acetabulum. In step S38, the control unit 11 generates 3D shape data of the acetabulum from the data on the left and right acetabulum extracted from the CT image. As a result, the training data used for training the second learning model M2c is stored in the training DB 12b. Training of the second learning model M2c can also be achieved by a process similar to that of FIG. 7 . Here, the control unit 11 performs learning using, as training data, 3D shape data generated from depth images of the front and back surfaces of the acetabulum and 3D shape data of the acetabulum generated from CT images (corrected 3D shape data of the correct answer). This allows for the generation of a second learning model M2c that outputs corrected 3D shape data in which missing parts in the 3D shape data have been corrected when the 3D shape data of the acetabulum is input.
[0126] Training data for the third learning model M3c can be generated by a process similar to that shown in FIG. 9 . In step S14 of FIG. 9 , the control unit 11 extracts data on the left and right acetabula (recesses) from the CT image as data on the region of interest (target region). In step S51, the control unit 11 generates silhouette images of the front and rear of the acetabulum from the CT image of the acetabulum. Here, silhouette images of the front and rear of the left and right acetabula are generated. As a result, training data used for training the third learning model M3c is stored in the training DB 12b. Training of the third learning model M3c can also be achieved by a process similar to that shown in FIG. 7 . Here, the control unit 11 performs training using a frontal hip joint X-ray image and silhouette images of the front and rear of the left and right acetabula as training data. As a result, by inputting a frontal hip joint X-ray image, the third learning model M3c can be generated, which outputs silhouette images of the front and rear of the left and right acetabula captured in the X-ray image.
[0127] The information processing device 10 of this embodiment can generate 3D shape data of a recess (e.g., acetabulum) by processing similar to that shown in FIG. 10 . In step S62 of FIG. 10 , the control unit 11 uses the third learning model M3c to generate silhouette images of the front and rear surfaces of the left and right acetabulum based on the front X-ray image of the hip joint. In step S63, the control unit 11 uses the first learning model M1c to generate depth images of the front and rear surfaces of the left and right acetabulum based on the front X-ray image of the hip joint. In step S64, the control unit 11 performs masking on the depth images of the front and rear surfaces of the left and right acetabulum based on the corresponding silhouette images. The control unit 11 then generates 3D shape data of the left and right acetabulum by connecting the contours of the acetabulum in the masked depth images of the front and rear surfaces of the left and right acetabulum.
[0128] As described above, for recesses, training data for learning is generated from X-ray images and CT images of the recess for which 3D shape data is to be generated, and a learning process using the training data is performed to generate the first learning model M1c to the third learning model M3c. Using such first learning model M1c to the third learning model M3c, 3D shape data of any recess can be generated (estimated) with high accuracy from an X-ray image of an area including the recess. Therefore, doctors and other professionals can grasp the shapes of not only the bone region, muscle region, organ, and other parts captured in X-ray and CT images, but also the recess itself. The configuration of this embodiment is not limited to recesses formed in bone regions, and 3D shape data can also be generated for recesses formed in muscles, fascia, organs, etc.
[0129] The configuration of this embodiment can be applied to the above-described embodiments 1 to 4, and similar effects can be obtained even when applied to embodiments 1 to 4. For example, when applied to embodiments 1 to 3, a depth image of the front surface and a depth image of the rear surface can be generated for a recess in each region captured in an X-ray image, and 3D shape data can be generated based on the depth image of the front surface and the depth image of the rear surface of the recess. Furthermore, when applied to embodiment 4, at least one of a depth image of the front surface and a depth image of the rear surface and a thickness image showing the thickness (size) of the recess can be generated for a recess in each region captured in an X-ray image, and 3D shape data can be generated based on the depth image and the thickness image of the front and / or rear surface of the recess.
[0130] In each of the above-described embodiments, simple X-ray images of the target region captured by an X-ray device (X-ray device) are used as input data for the first learning model M1 and the third learning model M3. However, this configuration is not limited to this. For example, X-ray images obtained by a DXA (Dual-energy X-ray Absorptiometry) device may be used as input data for the first learning model M1 and / or the third learning model M3. Even with this configuration, it is possible to perform processing similar to that performed when using normal X-ray images (X-ray images captured by an X-ray device) as input data, and similar effects can be obtained. Furthermore, in each of the above-described embodiments, the correct depth image (depth image of the region of interest) used as training data for the first learning model M1 is not limited to being generated from a CT image, but may also be generated from a three-dimensional image, such as an MRI image or an ultrasound image, from which 3D shape data of the region of interest can be generated. Furthermore, the 3D shape data and corrected 3D shape data used in the training data of the second learning model M2 are not limited to being generated from CT images, but may be generated from 3D images capable of generating 3D shape data of the region of interest, such as MRI images and ultrasound images. The correct silhouette image (silhouette image of the region of interest) used in the training data of the third learning model M3 is not limited to being generated from CT images, but may be generated from 3D images, such as MRI images and ultrasound images. Furthermore, the correct thickness image (thickness image of the region of interest) used in the training data of the fourth learning model M4 is not limited to being generated from CT images, but may be generated from 3D images capable of generating 3D shape data of the region of interest, such as MRI images and ultrasound images. Even with this configuration, it is possible to perform processing similar to that performed using the depth image, 3D shape data, corrected 3D shape data, silhouette image, and thickness image of the target region obtained from CT images, and similar effects can be obtained.
[0131] The matters described in each embodiment can be combined with each other. Furthermore, the independent claims and dependent claims described in the claims can be combined with each other in any and all combinations, regardless of the reference format. Furthermore, the claims use a format in which a claim references two or more other claims (multiple claim format), but this is not limited to this. A multiple claim (multi-multi claim) that references at least one other multiple claim may also be used.
[0132] The embodiments disclosed herein are to be considered in all respects as illustrative and not restrictive. The scope of the present invention is defined by the claims, not by the above meaning, and is intended to include all modifications within the meaning and scope of the claims.
[0133] REFERENCE SIGNS LIST 10 Information processing device 11 Control unit 12 Memory unit 13 Communication unit 14 Input unit 15 Display unit 16 Reading unit P Program (program product) 12a Medical image DB 12b Training DB M1 First learning model M2 Second learning model M3 Third learning model
Claims
1. A program that causes a computer to acquire an X-ray image of a target area, and input the acquired X-ray image into a learning model that outputs depth images of the front and back of the target area in the X-ray image when the X-ray image of the target area is input, thereby acquiring depth images of the front and back of the target area in the input X-ray image.
2. The program according to claim 1, which causes the computer to execute a process for generating three-dimensional shape data of the target region based on depth images of the front and back of the target region.
3. The program according to claim 1, which causes the computer to execute a process of acquiring a plurality of X-ray images of the target area taken from a plurality of directions, inputting each of the plurality of X-ray images into the learning model to acquire depth images of the front and back of the target area in each of the X-ray images, and generating three-dimensional shape data of the target area based on the depth images of the front and back of the target area acquired for each of the X-ray images.
4. A program as described in claim 2 or 3, which causes the computer to execute a process of obtaining corrected three-dimensional shape data by correcting the input three-dimensional shape data of the target part by inputting the generated three-dimensional shape data of the target part into a second learning model that outputs corrected three-dimensional shape data by correcting the three-dimensional shape data when the three-dimensional shape data of the target part is input.
5. A program as claimed in claim 2 or 3, which causes the computer to execute the following process: inputting an acquired X-ray image into a third learning model which, when an X-ray image of a target area is input, outputs silhouette images showing the shapes of the front and back of the target area in the X-ray image, thereby acquiring silhouette images of the front and back of the target area in the input X-ray image; masking background areas of depth images of the front and back of the target area using the silhouette images of the front and back of the target area; and generating three-dimensional shape data of the target area based on the depth images of the front and back of the target area with the background areas masked.
6. A program described in any one of claims 1 to 3, wherein the learning model is trained using training data including an X-ray image of the target area and depth images of the front and back of the target area obtained from a three-dimensional image of the target area, so that when the X-ray image is input, depth images of the front and back of the target area in the X-ray image are output.
7. The program described in claim 4, wherein the second learning model is trained to output corrected three-dimensional shape data obtained by correcting the three-dimensional shape data of the target area when the three-dimensional shape data of the target area is input, using training data including three-dimensional shape data of the target area generated based on depth images of the front and back of the target area obtained from the three-dimensional image of the target area, and three-dimensional shape data of the target area generated from the three-dimensional image of the target area.
8. The program according to any one of claims 1 to 3, wherein the target area includes at least one of a bone area and a muscle area.
9. An information processing method in which a computer acquires an X-ray image of a target area, and inputs the acquired X-ray image into a learning model that outputs depth images of the front and back of the target area in the X-ray image when the X-ray image of the target area is input, thereby acquiring depth images of the front and back of the target area in the input X-ray image.
10. An information processing device having a control unit, wherein the control unit acquires an X-ray image of a target area, and inputs the acquired X-ray image into a learning model that outputs depth images of the front and back of the target area in the X-ray image when the X-ray image of the target area is input, thereby acquiring depth images of the front and back of the target area in the input X-ray image.
11. A model generation method in which a computer acquires training data including an X-ray image of a target area and depth images of the front and back of the target area obtained from a 3D image of the target area, and uses the acquired training data to generate a learning model that, when an X-ray image of the target area is input, outputs depth images of the front and back of the target area in the X-ray image.
12. The model generation method according to claim 11, wherein the computer performs a process of aligning the position of the target area based on the X-ray image with the position of the target area based on the three-dimensional image, and acquiring the training data including depth images of the front and back of the target area obtained from the three-dimensional image after alignment.
13. A model generation method according to claim 11 or 12, wherein the three-dimensional images are three-dimensional images of multiple parts including the target part, the three-dimensional images are classified into parts, and the computer executes a process of acquiring the training data including depth images of the front and back of the target part obtained from the three-dimensional images of the classified target part.
14. A model generation method in which a computer executes a process of acquiring training data including three-dimensional shape data of a target area generated based on depth images of the front and back of the target area obtained from a three-dimensional image of the target area, and three-dimensional shape data of the target area generated from the three-dimensional image of the target area, and using the acquired training data to generate a second learning model that, when three-dimensional shape data of the target area is input, corrects the three-dimensional shape data and outputs corrected three-dimensional shape data.
15. A model generation method according to claim 11 or 14, wherein the three-dimensional image is a CT (Computed Tomography) image or an MRI (Magnetic Resonance Imaging) image of the target area.
16. The program described in claim 1 causes the computer to execute a process of acquiring a thickness image of the target area in the input X-ray image by inputting the acquired X-ray image into a fourth learning model that outputs a thickness image of the target area in the X-ray image when the X-ray image of the target area is input, and generating three-dimensional shape data of the target area based on at least one of the depth images of the front and back of the target area and the thickness image of the target area.
17. The program according to claim 1 or 16, wherein the target area is a recess formed in the area to be photographed in the X-ray image.
18. A model generation method according to claim 11, wherein the computer acquires training data including an X-ray image of a target area and a thickness image of the target area obtained from a three-dimensional image of the target area, and uses the acquired training data to generate a learning model for thickness image output that outputs a thickness image of the target area in an X-ray image when the X-ray image of the target area is input.
Citation Information
Patent Citations
Total hip replacement revision preoperative planning method and equipment based on deep learning
CN112971981A
Disease classification system based on deep learning and fiber bundle space statistical analysis
CN114494132A
Cloth calculation method and equipment for clothing model and storage medium
CN114758213A
Medical image generation method and device, computer equipment and storage medium
CN114913258A
Medical image processing device, x-ray diagnostic device, and medical information processing system
JP2020127600A