Estimation device, method, and program

A neural network trained with synthetic 2D images from 3D CT data enhances the accuracy of estimating images that emphasize specific components like bones and soft tissues in radiological images.

JP7758830B2Active Publication Date: 2025-10-22FUJIFILM CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024193174
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-11-01
Publication Date
2025-10-22
Estimated Expiration
2041-03-12

AI Technical Summary

Technical Problem

Existing methods for estimating images that emphasize specific components like bones and soft tissues in radiological images lack accuracy.

Method used

A neural network trained using synthetic 2D images derived from 3D CT images and training-enhanced images to enhance specific compositions, such as bones, soft tissues, muscle, and fat, is used to derive accurate estimation results from plain radiographic images.

Benefits of technology

Enables highly accurate estimation of images that emphasize specific components, such as bones and soft tissues, by utilizing a trained neural network to process plain radiographic images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007758830000001
    Figure 0007758830000001
  • Figure 0007758830000002
    Figure 0007758830000002
  • Figure 0007758830000003
    Figure 0007758830000003
Patent Text Reader

Abstract

To accurately predict a locomotorium disease in an estimation device, method, and program.SOLUTION: An estimation device includes at least one processor. The processor functions as a learned neural network for deriving an estimation result of at least one emphasis image that emphasizes a specific composition of a subject from a plain two-dimensional image acquired by executing plain radiography of the subject including a plurality of compositions. The learned neural network is learned by using, as teacher data, a composite two-dimensional image showing the subject derived by compositing three-dimensional CT images of the subject, and an emphasis image for learning that emphasizes a specific composition of the subject derived from the CT image.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an estimation device, a method, and a program. [Background technology]

[0002] Energy subtraction processing has been known for some time, utilizing the fact that the attenuation of transmitted radiation differs depending on the material constituting the subject, and using two radiological images obtained by irradiating the subject with two types of radiation with different energy distributions. Energy subtraction processing is a method of obtaining an image in which specific structures are emphasized by matching the pixels of the two radiological images obtained in this manner, multiplying the pixels by an appropriate weighting coefficient, and then performing subtraction. Energy subtraction processing has also been used to derive the composition of the human body, not only bones and soft tissues, but also soft tissues such as fat and muscle (see Patent Document 1).

[0003] Furthermore, various methods have been proposed for deriving a radiological image different from the radiological image acquired by photographing a subject using the radiological image acquired by photographing the subject. For example, Patent Document 2 proposes a method for deriving a bone image from a radiological image of a subject acquired by plain radiography by using a trained model constructed by training a neural network using a radiological image of the subject acquired by plain radiography and a bone image of the same subject as training data.

[0004] Note that simple radiography is a radiography method in which a subject is irradiated with radiation once to obtain a single two-dimensional image, which is a transmission image of the subject. In the following description, a two-dimensional image obtained by simple radiography will be referred to as a simple two-dimensional image. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Publication No. 2018-153605 [Patent Document 2] U.S. Patent No. 7,545,965 Summary of the Invention [Problem to be solved by the invention]

[0006] However, it is desired to estimate an image that emphasizes specific components of bones and the like with higher accuracy.

[0007] The present disclosure has been made in consideration of the above circumstances, and aims to enable highly accurate estimation of an image in which a specific composition is emphasized. [Means for solving the problem]

[0008] An estimation device according to the present disclosure includes at least one processor, The processor functions as a trained neural network that derives an estimation result of at least one enhanced image that emphasizes a specific composition of the subject from a simple two-dimensional image acquired by simply photographing the subject including a plurality of compositions; The trained neural network is trained using as training data a synthetic 2D image representing the subject derived by synthesizing 3D CT images of the subject, and a training-enhanced image that emphasizes specific components of the subject derived from the CT images.

[0009] In addition, in the estimation device according to the present disclosure, the composite two-dimensional image may be derived by deriving the radiation attenuation coefficient for the composition at each position in three-dimensional space and projecting the CT image in a predetermined direction based on the attenuation coefficient.

[0010] Furthermore, in the estimation device according to the present disclosure, the training-use emphasized image may be derived by identifying an area of ​​a specific composition in a CT image and projecting the CT image of the area of ​​the specific composition in a predetermined direction.

[0011] Furthermore, in the estimation device according to the present disclosure, the training emphasized image may be derived by weighted subtraction of two composite two-dimensional images that simulate imaging of a subject using radiation with different energy distributions, the composite two-dimensional images being derived by projecting a CT image in a predetermined direction.

[0012] In the estimation device according to the present disclosure, the specific composition may be at least one of soft tissue, bone, muscle, and fat of the subject.

[0013] The estimation method according to the present disclosure is an estimation method for deriving an estimation result of at least one enhanced image in which a specific composition of a subject is emphasized from a plain radiographic image using a trained neural network that derives an estimation result of at least one enhanced image in which a specific composition of the subject is emphasized from a plain radiographic image acquired by plain radiography of the subject, the estimation method comprising: The trained neural network is trained using as training data a synthetic 2D image representing the subject derived by synthesizing 3D CT images of the subject, and a training-enhanced image that emphasizes specific components of the subject derived from the CT images.

[0014] The estimation method according to the present disclosure may be provided as a program for causing a computer to execute the method. [Effects of the Invention]

[0015] According to the present disclosure, an image that emphasizes a specific composition can be estimated with high accuracy. [Brief explanation of the drawings]

[0016] [Figure 1] FIG. 1 is a schematic block diagram showing the configuration of a radiographic imaging system to which an estimation device according to a first embodiment of the present disclosure is applied. [Figure 2] FIG. 1 is a diagram showing a schematic configuration of an estimation device according to a first embodiment. [Figure 3] FIG. 1 is a diagram showing a functional configuration of an estimation device according to a first embodiment. [Figure 4]FIG. 1 is a diagram showing a schematic configuration of a neural network used in this embodiment. [Figure 5] Diagram showing training data [Figure 6] FIG. 1 is a diagram showing a schematic configuration of an information derivation device according to a first embodiment; [Figure 7] FIG. 1 is a diagram showing a functional configuration of an information derivation device according to a first embodiment; [Figure 8] A diagram for explaining the derivation of a synthetic 2D image. [Figure 9] A diagram for explaining the derivation of a synthetic 2D image. [Figure 10] Diagram to explain CT value [Figure 11] FIG. 1 is a diagram for explaining derivation of a bone image. [Figure 12] Image showing bones [Figure 13] FIG. 1 is a diagram for explaining derivation of a soft tissue image. [Figure 14] Soft tissue images [Figure 15] Diagram to explain neural network learning [Figure 16] Conceptual diagram of the processing performed by a trained neural network [Figure 17] Figure showing the display screen of the estimation results [Figure 18] 1 is a flowchart of a learning process performed in the first embodiment. [Figure 19] 1 is a flowchart of an estimation process performed in the first embodiment. [Figure 20] A diagram for explaining derivation of muscle images. [Figure 21] Figure showing muscle images [Figure 22] FIG. 1 is a diagram for explaining derivation of a fat image. [Figure 23] Diagram showing fat images [Figure 24] Another example of training data [Figure 25] FIG. 10 is a diagram showing a functional configuration of an information derivation device according to a third embodiment. [Figure 26] FIG. 10 is a diagram for explaining derivation of a composite two-dimensional image in the third embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0017] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. Fig. 1 is a schematic block diagram showing the configuration of a radiographic image capturing system to which an estimation device according to a first embodiment of the present disclosure is applied. As shown in Fig. 1, the radiographic image capturing system according to the first embodiment includes an imaging device 1, a CT device 7, an image storage system 9, an estimation device 10 according to the first embodiment, and an information derivation device 50. The imaging device 1, the CT (Computed Tomography) device 7, the estimation device 10, and the information derivation device 50 are connected to the image storage system 9 via a network (not shown).

[0018] The imaging device 1 is an imaging device that can acquire a simple radiographic image G0 of the subject H by irradiating a radiation detector 5 with radiation such as X-rays that are emitted from a radiation source 3 and transmitted through the subject H. The acquired simple radiographic image G0 is input to the estimation device 10. The simple radiographic image G0 is, for example, a front image including the groin area of ​​the subject H.

[0019] The radiation detector 5 is capable of repeatedly recording and reading out radiation images, and may be a so-called direct type radiation detector that generates electric charges upon direct exposure to radiation, or a so-called indirect type radiation detector that converts radiation into visible light and then converts the visible light into an electric charge signal. The radiation image signal readout method is preferably a TFT readout method in which the radiation image signal is read out by turning a TFT (thin film transistor) switch on and off, or an optical readout method in which the radiation image signal is read out by irradiating the detector with readout light, but is not limited to these, and other methods may also be used.

[0020] The CT device 7 performs CT imaging of the subject H to obtain a plurality of tomographic images representing a plurality of tomographic planes in the subject H as a three-dimensional CT image V0. The CT value of each pixel (voxel) in the CT image is a numerical representation of the radiation absorption rate in the components that make up the human body. The CT value will be described later.

[0021] The image storage system 9 is a system that stores image data of radiographic images acquired by the radiography device 1 and image data of CT images acquired by the CT device 7. The image storage system 9 extracts images from the stored radiographic images and CT images in response to requests from the estimation device 10 and the information derivation device 50, and transmits them to the requesting device. A specific example of the image storage system 9 is a PACS (Picture Archiving and Communication Systems). In this embodiment, the image storage system 9 stores a large amount of training data for training a neural network, which will be described later.

[0022] Next, an estimation device according to a first embodiment will be described. First, with reference to FIG. 2, the hardware configuration of the estimation device according to the first embodiment will be described. As shown in FIG. 2, the estimation device 10 is a computer such as a workstation, a server computer, or a personal computer, and includes a CPU (Central Processing Unit) 11, non-volatile storage 13, and memory 16 as a temporary storage area. The estimation device 10 also includes a display 14 such as a liquid crystal display, an input device 15 such as a keyboard and a mouse, and a network I / F (Interface) 17 connected to a network (not shown). The CPU 11, the storage 13, the display 14, the input device 15, the memory 16, and the network I / F 17 are connected to a bus 18. The CPU 11 is an example of a processor in the present disclosure.

[0023] The storage 13 is realized by a hard disk drive (HDD), a solid state drive (SSD), a flash memory, etc. The storage 13 as a storage medium stores the estimation program 12A and the learning program 12B installed in the estimation device 10. The CPU 11 reads out the estimation program 12A and the learning program 12B from the storage 13, expands them in the memory 16, and executes the expanded estimation program 12A and the learning program 12B.

[0024] The estimation program 12A and the learning program 12B are stored in a state accessible from the outside in a storage device of a server computer connected to a network or in a network storage, and are downloaded and installed in response to a request into a computer constituting the estimation device 10. Alternatively, they are recorded on a recording medium such as a DVD (Digital Versatile Disc) or a CD-ROM (Compact Disc Read Only Memory) and distributed, and are installed from the recording medium into a computer constituting the estimation device 10.

[0025] Next, the functional configuration of the estimation device according to the first embodiment will be described. Fig. 3 is a diagram showing the functional configuration of the estimation device according to the first embodiment. As shown in Fig. 3, the estimation device 10 includes an image acquisition unit 21, an information acquisition unit 22, an estimation unit 23, a learning unit 24, and a display control unit 25. The CPU 11 executes an estimation program 12A to function as the image acquisition unit 21, the information acquisition unit 22, the estimation unit 23, and the display control unit 25. The CPU 11 executes a learning program 12B to function as the learning unit 24.

[0026] The image acquisition unit 21 causes the imaging device 1 to perform simple imaging of the subject H, thereby acquiring from the radiation detector 5 a simple radiographic image G0, which is, for example, a frontal image of the area around the crotch of the subject H. When acquiring the simple radiographic image G0, imaging conditions are set, such as the imaging dose, radiation quality, tube voltage, SID (Source Image Receptor Distance) which is the distance between the radiation source 3 and the surface of the radiation detector 5, SOD (Source Object Distance) which is the distance between the radiation source 3 and the surface of the subject H, and the presence or absence of an anti-scatter grid.

[0027] The imaging conditions may be set by the operator through input device 15. The set imaging conditions are stored in storage 13. The plain radiographic image G0 and the imaging conditions are also transmitted to and stored in image storage system 9.

[0028] In this embodiment, the simple radiographic image G0 may be acquired by a program separate from the estimation program 12A and stored in the storage 13. In this case, the image acquisition unit 21 acquires the simple radiographic image G0 stored in the storage 13 by reading it from the storage 13 for processing.

[0029] The information acquisition unit 22 acquires training data for training a neural network, which will be described later, from the image storage system 9 via the network I / F 17.

[0030] The estimation unit 23 derives estimation results of a bone image in which bones included in the subject H are emphasized and a soft tissue image in which soft tissue is emphasized from the plain radiographic image G0. To this end, the estimation unit 23 derives estimation results of the bone image and the soft tissue image using a trained neural network 23A that outputs a bone image and a soft tissue image when the plain radiographic image G0 is input. Note that in this embodiment, the subject from which the estimation results of the bone image and the soft tissue image are derived is an image of the vicinity of the hip joint of the subject H, but is not limited to this. Furthermore, the high-grain image and the soft tissue image derived by the estimation unit 23 are examples of emphasized images.

[0031] The learning unit 24 constructs a trained neural network 23A by machine learning a neural network using training data. Examples of neural networks include a simple perceptron, a multilayer perceptron, a deep neural network, a convolutional neural network, a deep belief network, a recurrent neural network, and a probabilistic neural network. In this embodiment, a convolutional neural network is used as the neural network.

[0032] Fig. 4 is a diagram showing a neural network used in this embodiment. As shown in Fig. 4, the neural network 30 includes an input layer 31, an intermediate layer 32, and an output layer 33. The intermediate layer 32 includes, for example, a plurality of convolutional layers 35, a plurality of pooling layers 36, and a fully connected layer 37. In the neural network 30, the fully connected layer 37 is located before the output layer 33. In the neural network 30, the convolutional layers 35 and the pooling layers 36 are alternately arranged between the input layer 31 and the fully connected layer 37.

[0033] The configuration of the neural network 30 is not limited to the example shown in Fig. 4. For example, the neural network 30 may include one convolutional layer 35 and one pooling layer 36 between the input layer 31 and the fully connected layer 37.

[0034] 5 is a diagram showing an example of training data used for training a neural network. As shown in FIG. 5, training data 40 consists of training data 41 and correct answer data 42. In this embodiment, the data input to trained neural network 23A to obtain a bone mineral density estimation result is a plain radiographic image G0, but training data 41 includes a composite two-dimensional image C0 representing subject H derived by combining a CT image V0.

[0035] The correct answer data 42 is a bone image Gb and a soft tissue image Gs near the target bone (i.e., the femur) of the subject from which the learning data 41 was obtained. The bone image Gb and the soft tissue image Gs, which are the correct answer data 42, are derived from the CT image V0 by an information derivation device 50. The information derivation device 50 will be described below. The bone image Gb and the soft tissue image Gs derived from the CT image V0 are examples of learning-use emphasized images.

[0036] Fig. 6 is a schematic block diagram showing the configuration of an information derivation device according to a first embodiment. As shown in Fig. 6, information derivation device 50 according to the first embodiment is a computer such as a workstation, a server computer, or a personal computer, and includes a CPU 51, non-volatile storage 53, and memory 56 as a temporary storage area. Information derivation device 50 also includes a display 54 such as a liquid crystal display, an input device 55 including a keyboard, a pointing device such as a mouse, and a network I / F 57 connected to a network (not shown). CPU 51, storage 53, display 54, input device 55, memory 56, and network I / F 57 are connected to a bus 58.

[0037] The storage 53 is realized by an HDD, an SSD, a flash memory, or the like, similar to the storage 13. The storage 53 as a storage medium stores an information derivation program 52. The CPU 51 reads the information derivation program 52 from the storage 53, expands it in the memory 56, and executes the expanded information derivation program 52.

[0038] Next, the functional configuration of the information derivation device according to the first embodiment will be described. Fig. 7 is a diagram showing the functional configuration of the information derivation device according to the first embodiment. As shown in Fig. 7, the information derivation device 50 according to the first embodiment includes an image acquisition unit 61, a synthesis unit 62, and an emphasized image derivation unit 63. When the CPU 51 executes the information derivation program 52, the CPU 51 functions as the image acquisition unit 61, the synthesis unit 62, and the emphasized image derivation unit 63.

[0039] The image acquisition unit 61 acquires the CT image V0 for deriving the learning data 41 from the image storage system 9. Note that the image acquisition unit 61 may acquire the CT image V0 by causing the CT device 7 to capture an image of the subject H, similar to the image acquisition unit 21 of the estimation device 10.

[0040] The synthesis unit 62 synthesizes the CT images V0 to derive a synthesized two-dimensional image C0 representing the subject H. FIG. 8 is a diagram for explaining the derivation of the synthesized two-dimensional image C0. For the sake of explanation, FIG. 8 shows the three-dimensional CT image V0 in two dimensions. As shown in FIG. 8, the subject H is included in the three-dimensional space represented by the CT image V0. The subject H is composed of multiple components, including bones, fat, muscles, and internal organs.

[0041] Here, the CT value V0(x, y, z) of each pixel of the CT image V0 can be expressed by the following formula (1) using the attenuation coefficient μi of the composition at that pixel and the attenuation coefficient μw of water. (x, y, z) are coordinates that represent the pixel position of the CT image V0. In the following explanation, attenuation coefficient means the radiation source attenuation coefficient unless otherwise specified. The attenuation coefficient represents the degree (proportion) of attenuation of radiation due to absorption or scattering, etc. The attenuation coefficient differs depending on the specific composition (density, etc.) and thickness (mass) of the structure through which the radiation passes. V0(x,y,z)=(μi-μw) / μw×1000 (1)

[0042] The attenuation coefficient μw of water is known. Therefore, by solving equation (1) for μi, the attenuation coefficient μi of each composition can be calculated as shown in equation (2) below. μi=V0(x,y,z)×μw / 1000+μw (2)

[0043] As shown in FIG. 8, the synthesis unit 62 virtually irradiates the subject H with radiation at an exposure dose I0, and derives a synthetic two-dimensional image C0 by virtually detecting the radiation that has passed through the subject H using a radiation detector (not shown) installed on a virtual plane 64. The exposure dose I0 and radiation energy of the virtual radiation are set according to predetermined imaging conditions. At this time, the arrival dose I1(x, y) for each pixel of the synthetic two-dimensional image C0 passes through one or more components in the subject H. Therefore, the arrival dose I1(x, y) can be derived from the following equation (3) using the attenuation coefficient μi of one or more components through which the radiation of the exposure dose I0 passes. The arrival dose I1(x, y) becomes the pixel value of each pixel of the synthetic two-dimensional image C0. I1(x,y)=I0×exp(-∫μi·dt) (3)

[0044] If the irradiating radiation source is assumed to be a surface light source, the attenuation coefficient μi used in equation (3) can be derived from equation (2) using the CT values ​​of the pixels arranged in the vertical direction as shown in Fig. 8. If the irradiating radiation source is assumed to be a point light source, as shown in Fig. 9, pixels on the path of the radiation that reaches each pixel can be identified based on the geometric positional relationship between the point light source and each position on a virtual plane 64, and the attenuation coefficient derived from equation (2) using the CT values ​​of the identified pixels can be used.

[0045] The enhanced image derivation unit 63 uses the CT image V0 to derive a bone image Gb in which the bones of the subject are enhanced and a soft tissue image Gs in which the soft tissues are enhanced. Here, the CT value will be explained. FIG. 10 is a diagram for explaining the CT value. The CT value is a numerical representation of the X-ray absorption rate in the human body. Specifically, as shown in FIG. 10, the CT value is determined according to the composition of the human body, with water having a CT value of 0 and air having a CT value of -1000 (unit: HU).

[0046] The enhanced image derivation unit 63 first identifies a bone region in the CT image V0 based on the CT value of the CT image V0. Specifically, a region consisting of pixels with a CT value of 100 to 1000 is identified as a bone region by threshold processing. Note that instead of threshold processing, the bone region may be identified using a trained neural network that has been trained to detect bone regions from the CT image V0. Alternatively, the CT image V0 may be displayed on the display 54, and the bone region may be identified by manually specifying the bone region in the displayed CT image V0.

[0047] 11, the enhanced image derivation unit 63 projects the CT values ​​of the bone region Hb in the subject H included in the CT image V0 onto a virtual plane 64 in the same manner as when deriving the composite two-dimensional image C0, thereby deriving a bone image Gb. The bone image Gb is shown in FIG.

[0048] The enhanced image derivation unit 63 also identifies soft tissue regions in the CT image V0 based on the CT values ​​of the CT image V0. Specifically, a region consisting of pixels with CT values ​​between -100 and 70 is identified as a soft tissue region by threshold processing. Instead of threshold processing, a trained neural network that has been trained to detect soft tissue regions from the CT image V0 may be used to identify the soft tissue region. Alternatively, the CT image V0 may be displayed on the display 54, and the soft tissue region may be identified by manually specifying the soft tissue region in the displayed CT image V0.

[0049] 13, the enhanced image derivation unit 63 derives a soft tissue image Gs by projecting the CT values ​​of the soft tissue region Hs in the subject H included in the CT image V0 onto a virtual plane 64 in the same manner as when deriving the composite two-dimensional image C0. The soft tissue image Gs is shown in FIG.

[0050] The bone images Gb and soft tissue images Gs used as the correct answer data 42 are derived at the same time as the learning data 41 is acquired, and are transmitted to the image storage system 9. In the image storage system 9, the learning data 41 and the correct answer data 42 are associated with each other and stored as training data 40. To improve the robustness of learning, additional training data 40 may be created and stored, including images obtained by performing at least one of the following on the same image: enlargement / reduction, contrast change, translation, in-plane rotation, inversion, and noise addition.

[0051] Returning to the estimation device 10, the learning unit 24 trains the neural network using a large amount of training data 40. FIG. 15 is a diagram for explaining the training of the neural network 30. When training the neural network 30, the learning unit 24 inputs training data 41, i.e., a synthetic 2D image C0, to the input layer 31 of the neural network 30. The learning unit 24 then outputs bone images and soft tissue images as output data 47 from the output layer 33 of the neural network 30. The learning unit 24 then derives the difference between the output data 47 and the supervised data 42 as a loss L0. Note that the loss is derived between the bone images in the output data 47 and the bone images in the supervised data 42, and between the soft tissue images in the output data and the soft tissue images in the supervised data 42, but the loss will be referred to as L0.

[0052] The learning unit 24 trains the neural network 30 based on the loss L0. Specifically, the learning unit 24 adjusts the kernel coefficients in the convolutional layer 35, the connection weights between layers, the connection weights in the fully connected layer 37, and the like (hereinafter referred to as parameters 48) so as to reduce the loss L0. The parameters 48 can be adjusted, for example, by backpropagation. The learning unit 24 repeatedly adjusts the parameters 48 until the loss L0 becomes equal to or less than a predetermined threshold. In this way, when a simple radiographic image G0 is input, the parameters 48 are adjusted so that a bone image Gb and a soft tissue image Gs for the input simple radiographic image G0 are output, and a trained neural network 23A is constructed. The constructed trained neural network 23A is stored in the storage 13.

[0053] Fig. 16 is a conceptual diagram of the processing performed by the trained neural network 23A. As shown in Fig. 16, when a plain radiographic image G0 of a patient is input to the trained neural network 23A constructed as described above, the trained neural network 23A outputs a bone image Gb and a soft tissue image Gs for the input plain radiographic image G0.

[0054] The display control unit 25 displays the estimation results of the bone image Gb and soft tissue image Gs estimated by the estimation unit 23 on the display 14. FIG. 17 is a diagram showing a display screen for the estimation results. As shown in FIG. 17, the display screen 70 has a first image display area 71 and a second image display area 72. A simple radiographic image G0 of the subject H is displayed in the first image display area 71. Furthermore, the bone image Gb and soft tissue image Gs estimated by the estimation unit 23 are displayed in the second image display area 72.

[0055] Next, the processing performed in the first embodiment will be described. FIG. 18 is a flowchart showing the learning processing performed in the first embodiment. First, the information acquisition unit 22 acquires training data 40 from the image storage system 9 (step ST1). The learning unit 24 inputs training data 41 included in the training data 40 into the neural network 30, causing it to output bone images Gb and soft tissue images Gs, and trains the neural network 30 using a loss L0 based on the difference from the ground truth data 42 (step ST2), and returns to step ST1. The learning unit 24 then repeats the processing in steps ST1 and ST2 until the loss L0 reaches a predetermined threshold value, and then ends the learning processing. Note that the learning unit 24 may end the learning processing by repeating the learning a predetermined number of times. In this way, the learning unit 24 constructs a trained neural network 23A.

[0056] Next, the estimation process in the first embodiment will be described. Fig. 19 is a flowchart showing the estimation process in the first embodiment. It is assumed that the simple radiographic image G0 is acquired by imaging and stored in the storage 13. When an instruction to start the process is input from the input device 15, the image acquisition unit 21 acquires the simple radiographic image G0 from the storage 13 (step ST11). Next, the estimation unit 23 derives estimation results of the bone image Gb and the soft tissue image Gs from the simple radiographic image G0 (step ST12). Then, the display control unit 25 displays the estimation results of the bone image Gb and the soft tissue image Gs derived by the estimation unit 23 on the display 14 together with the simple radiographic image G0 (step ST13), and the process ends.

[0057] As described above, in this embodiment, estimated results of the bone image Gb and the soft tissue image Gs for the plain radiographic image G0 are derived using a trained neural network 23A constructed by training using the composite two-dimensional image C0 derived from the CT image V0 and the bone image Gb and the soft tissue image Gs derived from the CT image V0 as training data. In this embodiment, the neural network is trained using the composite two-dimensional image C0 derived from the CT image V0 and the bone image Gb and the soft tissue image Gs derived from the CT image V0. Therefore, compared to using a single radiographic image and information related to the bone image Gb and the soft tissue image Gs derived from the radiographic image as training data, the trained neural network 23A can more accurately derive estimated results of the bone image Gb and the soft tissue image Gs from the plain radiographic image G0. Therefore, this embodiment allows for more accurate estimation results of the bone image Gb and the soft tissue image Gs to be derived.

[0058] In the first embodiment, the estimation results of the bone image Gb and the soft tissue image Gs are derived, but this is not limited to this. The estimation results of either the bone image Gb or the soft tissue image Gs may be derived. In this case, the trained neural network 23A may be constructed by performing training using training data in which the correct answer data is either the bone image Gb or the soft tissue image Gs.

[0059] In the first embodiment, the estimation results of the bone image Gb and the soft tissue image are derived from the plain radiographic image G0, but this is not limited to this. For example, the estimation results of the muscle image and the fat image may be derived. This will be described below as the second embodiment.

[0060] The configurations of the estimation device and the information derivation device in the second embodiment are the same as those of the estimation device 10 and the information derivation device 50 in the first embodiment, and only the processing performed is different, so detailed description will be omitted here. In the second embodiment, the enhanced image derivation unit 63 of the information derivation device 50 derives a muscle image Gm and a fat image Gf instead of deriving a bone image Gb and a soft tissue image Gs as the correct answer data 42.

[0061] In the second embodiment, the enhanced image derivation unit 63 first identifies a muscle region in the CT image V0 based on the CT value of the CT image V0. Specifically, a region consisting of pixels with a CT value of 60 to 70 is identified as the muscle region by threshold processing. Note that instead of threshold processing, the muscle region may be identified using a trained neural network that has been trained to detect muscle regions from the CT image V0. Alternatively, the CT image V0 may be displayed on the display 54, and the muscle region may be identified by manually specifying the muscle region on the displayed CT image V0.

[0062] 20, the enhanced image derivation unit 63 derives a muscle image Gm by projecting the CT values ​​of a muscle region Hm in the subject H included in the CT image V0 onto a virtual plane 64 in the same manner as when deriving the composite two-dimensional image C0. The muscle image Gm is shown in FIG.

[0063] The enhanced image derivation unit 63 also identifies a fatty region in the CT image V0 based on the CT value of the CT image V0. Specifically, a region consisting of pixels with a CT value between -100 and -10 is identified as a fatty region by threshold processing. Note that instead of threshold processing, a trained neural network that has been trained to detect a fatty region from the CT image V0 may be used to identify the fatty region. Alternatively, the CT image V0 may be displayed on the display 54, and the fatty region may be identified by manually specifying the fatty region in the displayed CT image V0.

[0064] 22, the enhanced image derivation unit 63 derives a fat image Gf by projecting the CT value of the fat region Hf in the subject H included in the CT image V0 onto a virtual plane 64 in the same manner as when deriving the composite two-dimensional image C0. The fat image Gf is shown in FIG.

[0065] In the second embodiment, the muscle image Gm and fat image Gf derived by the information derivation device 50 are used as supervised data of the training data. Fig. 24 is a diagram showing the training data derived in the second embodiment. As shown in Fig. 24, the training data 40A consists of training data 41 including a composite two-dimensional image C0 and supervised data 42A including a muscle image Gm and a fat image Gf.

[0066] By training a neural network using the training data 40A shown in FIG. 24, a trained neural network 23A can be constructed that outputs a muscle image Gm and a fat image Gf as estimation results when a plain radiographic image G0 is input.

[0067] In the second embodiment, the estimation results of the muscle image Gm and the fat image Gf are derived, but this is not limited to this. The estimation results of either the muscle image Gm or the fat image Gf may be derived. In this case, the trained neural network 23A may be constructed by training using training data in which the correct answer data is either the muscle image Gm or the fat image Gf.

[0068] Next, a third embodiment of the present disclosure will be described. Fig. 25 is a diagram showing the functional configuration of an information derivation device according to the third embodiment. In Fig. 25, the same components as those in Fig. 7 are assigned the same reference numerals, and detailed description thereof will be omitted. In the first and second embodiments, training-use enhanced images (i.e., bone image Gb, soft tissue image Gs, muscle image Gm, and fat image Gf) that serve as ground truth data 42 are derived by projecting a specific composition within subject H in CT image V0. In the third embodiment, two composite two-dimensional images that simulate a low-energy radiation image and a high-energy radiation image are derived from CT image V0, and a training-use enhanced image is derived by performing weighted subtraction on the two derived composite two-dimensional images.

[0069] 25, the information derivation device 50A according to the third embodiment further includes a synthesis unit 62A and an enhanced image derivation unit 63A in addition to the components of the information derivation device 50 according to the first embodiment. In the following description, it is assumed that a bone image Gb and a soft tissue image Gs are derived as training enhanced images.

[0070] In the third embodiment, as shown in Fig. 26, the synthesis unit 62A virtually irradiates subject H with radiation of two doses IL0 and IH0 with different energy distributions, and derives two synthetic two-dimensional images CL0 and CH0 by virtually detecting radiation with transmission doses IL1 and IH1 that have passed through subject H using a radiation detector installed on a virtual plane 64. The synthetic two-dimensional image CL0 corresponds to a radiographic image of subject H taken using low-energy radiation that also includes so-called soft rays. The synthetic two-dimensional image CH0 corresponds to a radiographic image of subject H taken using high-energy radiation from which soft rays have been removed.

[0071] The enhanced image derivation unit 63A identifies bone regions and soft tissue regions in the composite 2D images CL0 and CH0. The bone regions and soft tissue regions have clearly different pixel values ​​in the composite 2D images CL0 and CH0. Therefore, the enhanced image derivation unit 63A identifies bone regions and soft tissue regions in the composite 2D images CL0 and CH0 by threshold processing. Note that instead of threshold processing, the bone regions and soft tissue regions may be identified using a trained neural network that has been trained to detect bone regions and soft tissue regions in the composite 2D images CL0 and CH0. Alternatively, the composite 2D images CL0 and CH0 may be displayed on the display 54, and the bone regions and soft tissue regions may be identified by manually specifying the bone regions and soft tissue regions in the displayed composite 2D images CL0 and CH0.

[0072] In the third embodiment, the enhanced image derivation unit 63A derives a bone image Gb and a soft-tissue image Gs by performing weighted subtraction between two composite two-dimensional images. To this end, the enhanced image derivation unit 63A derives, for the soft-tissue regions in the two composite two-dimensional images CL0 and CH0, the ratio CLs(x,y) / CHs(x,y) of pixel values ​​CLs(x,y) of composite two-dimensional image CL0 corresponding to a low-energy image to pixel values ​​CHs(x,y) of composite two-dimensional image CH0 corresponding to a high-energy image, as a weighting coefficient α used in performing weighted subtraction to derive the bone image Gb. The ratio CLs(x,y) / CHs(x,y) represents the ratio μls / μhs of the attenuation coefficient μls for low-energy radiation to the attenuation coefficient μhs for high-energy radiation in soft tissue.

[0073] Furthermore, for the bone regions in the two composite two-dimensional images CL0 and CH0, the enhanced image derivation unit 63A derives the ratio CLb(x,y) / CHb(x,y) of the pixel values ​​CLb(x,y) of the composite two-dimensional image CL0 corresponding to the low-energy image to the pixel values ​​CHb(x,y) of the composite two-dimensional image CH0 corresponding to the high-energy image as the weighting coefficient β when performing weighted subtraction to derive the soft tissue image Gs. Note that the ratio CLb(x,y) / CHb(x,y) represents the ratio μlb / μhb of the attenuation coefficient μlb for low-energy radiation to the attenuation coefficient μhb for high-energy radiation in the bone region.

[0074] In the third embodiment, the enhanced image derivation unit 63A uses the derived weighting coefficients α and β to perform weighted subtraction on the composite two-dimensional images CL0 and CH0 according to the following equations (4) and (5), thereby deriving a bone image Gb and a soft tissue image Gs. Gb(x,y)=α·CH0(x,y)-CL0(x,y) (4) Gs(x, y)=CH0(x, y)-β×CH0(x, y) (5)

[0075] In each of the above embodiments, the trained neural network 23A is constructed by training the neural network in the estimation device 10, but this is not limiting. A trained neural network 23A constructed in a device other than the estimation device 10 may be used in the estimation unit 23 of the estimation device 10 in this embodiment.

[0076] Furthermore, in each of the above embodiments, the process of estimating an enhanced image is performed using a radiographic image acquired in a system that uses a radiation detector 5 to capture an image of the subject H. However, the technology of the present disclosure can of course also be applied to cases where a stimulable phosphor sheet is used to acquire a radiographic image instead of a radiation detector.

[0077] Furthermore, the radiation in the above embodiment is not particularly limited, and in addition to X-rays, α rays, γ rays, etc. can be used.

[0078] Furthermore, in the above embodiment, the following various processors can be used as the hardware structure of processing units that perform various processes, such as the image acquisition unit 21, information acquisition unit 22, estimation unit 23, learning unit 24, and display control unit 25 of the estimation device 10, and the image acquisition unit 61, synthesis unit 62, and enhanced image derivation unit 63 of the information derivation device 50, etc. As described above, the various processors include a CPU, which is a general-purpose processor that executes software (programs) and functions as various processing units, as well as dedicated electrical circuits that are processors having a circuit configuration specifically designed to perform specific processes, such as a programmable logic device (PLD), a processor whose circuit configuration can be changed after manufacture, such as an FPGA (Field Programmable Gate Array), and an ASIC (Application Specific Integrated Circuit).

[0079] A single processing unit may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs or a combination of a CPU and an FPGA). Also, multiple processing units may be configured with a single processor.

[0080] Examples of configuring multiple processing units with a single processor include, first, a form in which one processor is configured with a combination of one or more CPUs and software, and this processor functions as multiple processing units, as typified by computers such as client and server. Second, a form in which a processor is used to realize the functions of an entire system including multiple processing units with a single IC (Integrated Circuit) chip, as typified by systems on chips (SoCs). In this way, various processing units are configured using one or more of the above-mentioned various processors as a hardware structure.

[0081] Furthermore, more specifically, the hardware structure of these various processors can be an electric circuit that combines circuit elements such as semiconductor elements.

[0082] The following are appendices to the present disclosure. (Additional note 1) at least one processor; The processor: The neural network functions as a trained neural network that derives an estimation result of at least one enhanced image in which a specific composition of a subject is enhanced from a simple two-dimensional image acquired by simply photographing the subject including the plurality of compositions, the trained neural network is trained using, as training data, a synthetic two-dimensional image representing the subject derived by synthesizing three-dimensional CT images of the subject, and a training-use enhanced image derived from the CT images in which a specific composition of the subject is enhanced. (Additional note 2) The estimation device according to appended claim 1, wherein the composite two-dimensional image is derived by deriving a radiation attenuation coefficient for a composition at each position in three-dimensional space and projecting the CT image in a predetermined direction based on the attenuation coefficient. (Additional note 3) 3. The estimation device according to claim 1, wherein the training-use weighted image is derived by identifying an area of ​​a specific composition in the CT image and projecting the CT image of the area of ​​the specific composition in a predetermined direction. (Additional note 4) The estimation device according to appendix 1 or 2, wherein the training-use emphasized image is derived by performing weighted subtraction on two composite two-dimensional images that simulate imaging of the subject using radiation with different energy distributions, the composite two-dimensional images being derived by projecting the CT image in a predetermined direction. (Additional note 5) 5. The estimation device according to any one of appendixes 1 to 4, wherein the specific composition is at least one of soft tissue, bone, muscle, and fat of the subject. (Additional note 6) An estimation method for deriving an estimation result of at least one enhanced image in which a specific composition of a subject is emphasized from a plain radiographic image obtained by simple radiography of the subject, using a trained neural network that derives an estimation result of at least one enhanced image in which a specific composition of the subject is emphasized from the plain radiographic image, the method comprising: an estimation method in which the trained neural network is trained using, as training data, a synthetic two-dimensional image representing the subject derived by synthesizing three-dimensional CT images of the subject, and a training-use enhanced image derived from the CT images in which a specific composition of the subject is enhanced. (Additional note 7) An estimation program that causes a computer to execute a procedure for deriving an estimation result of at least one enhanced image in which a specific composition of a subject is emphasized from a plain radiographic image obtained by plain radiography of the subject, using a trained neural network that derives an estimation result of at least one enhanced image in which a specific composition of the subject is emphasized from the plain radiographic image, the program comprising: The trained neural network is an estimation program that is trained using, as training data, a synthetic two-dimensional image representing the subject derived by synthesizing three-dimensional CT images of the subject, and a training-use enhanced image that emphasizes a specific composition of the subject derived from the CT images. [Explanation of symbols]

[0083] 1. Imaging device 3 Radiation source 5. Radiation detectors 7 CT device 9. Image Storage System 10 Estimation device 11, 51 CPUs 12 Estimation Processing Program 12B Study Program 13, 53 Storage 14, 54 display 15, 55 Input devices 16, 56 memory 17, 57 Network I / F Buses 18 and 58 21 Image acquisition unit 22 Information Acquisition Department 23 Estimation part 23A Trained Neural Network 24 Learning Department 25 Display control unit 30 Neural Networks 31 Input layer 32 Middle Class 33 Output layer 35 convolutional layers 36 Pooling Layer 37 Fully connected layer 40, 40A training data 41 Training data 42, 42A correct data 47 Output Data 48 parameters 50,50A Information Derivation Device 52 Information derivation program 61 Image acquisition unit 62, 62A synthesis section 63, 63A Enhanced image derivation unit 64 plane 70 display screen 71 first image display area 72 Second image display area C0, CL0, CH0 composite 2D image G0 plain radiographic image Gb Bone image Gf Fat Images Gm muscle images Gs Soft tissue images H Subject Hb bone area Hf fat area Hm muscle area Hs Soft region I0, IL0, IH0 exposure dose I1, IL1, IH1 achieved dose V0 CT image

Claims

1. at least one processor; The processor: The neural network functions as a trained neural network that derives an estimation result of at least one enhanced image in which at least one of muscle and fat compositions of a subject is enhanced from a plain radiographic image obtained by plain radiography of the subject including a plurality of compositions, the trained neural network is trained using, as training data, a composite two-dimensional image representing the subject derived by synthesizing three-dimensional CT images of the subject, and a training enhanced image derived from the CT images, in which a composition of the subject that is to be enhanced in the estimation result of the enhanced image output by the trained neural network is enhanced; The learning-use enhanced image is derived by weighted subtraction of two composite two-dimensional images that simulate imaging of the subject using radiation with different energy distributions, the composite two-dimensional images being derived by projecting the CT image in a predetermined direction.

2. The estimation device according to claim 1, wherein the composite two-dimensional image is derived by deriving a radiation attenuation coefficient for a composition at each position in three-dimensional space and projecting the CT image in a predetermined direction based on the attenuation coefficient.

3. An estimation method for deriving an estimation result of at least one enhanced image in which a specific composition of a subject is emphasized from a plain radiographic image obtained by simple radiography of the subject, the estimation result being at least one enhanced image in which at least one of muscle and fat compositions of the subject is emphasized from the plain radiographic image, the method comprising: the trained neural network is trained using, as training data, a composite two-dimensional image representing the subject derived by synthesizing three-dimensional CT images of the subject, and a training enhanced image derived from the CT images, in which a composition of the subject that is to be enhanced in the estimation result of the enhanced image output by the trained neural network is enhanced; The estimation method, in which the training-use emphasized image is derived by weighted subtraction of two composite two-dimensional images that simulate imaging of the subject with radiation having different energy distributions, which are derived by projecting the CT image in a predetermined direction.

4. An estimation program that causes a computer to execute a procedure for deriving an estimated result of at least one enhanced image in which a specific composition of a subject is emphasized from a plain radiographic image, using a trained neural network that derives an estimated result of at least one enhanced image in which at least one of muscle and fat compositions of the subject is emphasized from the plain radiographic image, the program comprising: the trained neural network is trained using, as training data, a composite two-dimensional image representing the subject derived by synthesizing three-dimensional CT images of the subject, and a training enhanced image derived from the CT images, in which a composition of the subject that is to be enhanced in the estimation result of the enhanced image output by the trained neural network is enhanced; The learning-use emphasized image is derived by weighted subtraction of two composite two-dimensional images that simulate imaging of the subject using radiation with different energy distributions, which are derived by projecting the CT image in a predetermined direction.

Citation Information

Patent Citations

  • Medical image deboning model construction method and bone information removal method

    CN111179373A

  • Method and apparatus which make artifact reduction easy

    JP2004188187A

  • Bed positioning system, radiotherapy system and bed positioning method

    JP2010187991A

  • Body fat percentage measuring apparatus, method, and program

    JP2018153605A

  • Medical image reconstruction method and device thereof

    JP2020093083A