Learning method, image processing method, learning device, image processing device, learning program, and image processing program
The learning method projects 3D CT images into 2D pseudo-X-ray images, using iterative comparisons and deep learning to enhance anatomical separation in medical images, addressing the challenge of distinguishing bones and soft tissues.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- FUJIFILM CORP
- Filing Date
- 2022-06-20
- Publication Date
- 2026-06-04
AI Technical Summary
Existing methods for separating anatomical structures in medical images, such as X-ray and CT images, face challenges in accurately distinguishing between bones and soft tissues, requiring specialized equipment or machine learning techniques that are not efficient in all scenarios.
A learning method that projects 3D CT images into 2D space to generate pseudo-X-ray images, using iterative comparisons and parameter updates to separate anatomical regions, incorporating techniques like deep learning and neural networks to enhance separation accuracy.
The method effectively separates anatomical regions in medical images, improving recognition and reducing errors through iterative learning processes, enabling precise identification of bones and soft tissues.
Smart Images

Figure 0007870280000001 
Figure 0007870280000002 
Figure 0007870280000003
Abstract
Description
Technical Field
[0001] The present invention relates to the learning of medical images and medical image processing.
Background Art
[0002] In the field of handling medical images such as X-ray images and CT images (CT: Computed Tomography) (sometimes also referred to as medical images), the captured images are decomposed into a plurality of anatomical regions. When separating such images, in the case of simple X-ray images, in addition to organs and blood vessels, bones are superimposed and projected, so it is not easy to recognize anatomical structures and diseases. On the other hand, by imaging using X-ray sources in a plurality of energy bands, bones and soft tissues can be separated, but dedicated imaging equipment is required. In addition, in such a field, a method for separating simple X-ray images by machine learning has been developed. For example, Non-Patent Documents 1 and 2 describe separating simple X-ray images into bones and soft tissues using a CNN (Convolutional Neural Network).
[0003]
Prior Art Documents
Non-Patent Documents
[0004]
Non-Patent Document 1
[0005] One embodiment of the technology of this disclosure provides a learning method, learning apparatus, and learning program for separating a two-dimensional X-ray image into various regions, as well as an image processing method, image processing apparatus, and image processing program using the learning results. Means for solving the problem
[0006] To achieve the above-mentioned objectives, the learning method according to the first aspect of the present invention inputs a 3D CT image of a subject, projects the 3D CT image into a 2D space to generate a first pseudo-X-ray image, generates a second pseudo-X-ray image from the first pseudo-X-ray image which is a plurality of images in which a plurality of anatomical regions of the subject are each separated, obtains a label image in which a plurality of anatomical regions are labeled on the 3D CT image, generates a masked CT image which is a 3D image in which a plurality of anatomical regions are each separated, based on the 3D CT image and the label image, and projects the masked CT image into a 2D space. The system projects the image to generate a third pseudo-X-ray image, which consists of multiple images showing separate anatomical regions. The second and third pseudo-X-ray images are then compared, and based on the comparison, the parameter values used to generate the second pseudo-X-ray image are updated. The process of inputting a 3D CT image, generating the first pseudo-X-ray image, generating the second pseudo-X-ray image, acquiring a labeled image, generating a masked CT image, generating the third pseudo-X-ray image, comparing the second and third pseudo-X-ray images, and updating the parameter values is repeated until predetermined conditions are met.
[0007] In the learning method according to the second embodiment, in the first embodiment, when comparing the second pseudo-X-ray image with the third pseudo-X-ray image, the error of the second pseudo-X-ray image with respect to the third pseudo-X-ray image is calculated, and when updating the parameter values, the error is reduced by updating the parameters.
[0008] The learning method according to the third embodiment, in the first embodiment, compares the second pseudo-X-ray image with the third pseudo-X-ray image, calculates the error of the second pseudo-X-ray image relative to the third pseudo-X-ray image, distinguishes between the third pseudo-X-ray image and the second pseudo-X-ray image, integrates the error and the discrimination result, and in updating the parameter values, reduces the result of the integration by updating the parameters.
[0009] The learning method according to the fourth embodiment, in the first embodiment, compares the second pseudo-X-ray image with the third pseudo-X-ray image, calculates the error of the second pseudo-X-ray image relative to the third pseudo-X-ray image, distinguishes between the pair data of the first pseudo-X-ray image and the third pseudo-X-ray image and the pair data of the first pseudo-X-ray image and the second pseudo-X-ray image, integrates the error and the discrimination result, and in updating the parameter values, reduces the result of the integration by updating the parameters.
[0010] The learning method relating to the fifth aspect involves acquiring a label image by generating a label image from a 3D CT image in any one of the first to fourth aspects.
[0011] The learning method relating to the sixth aspect, in any one of the first to fifth aspects, generates a masked CT image by generating a mask image from extracted CT images obtained by region extraction from a 3D CT image, and then generates a masked CT image by multiplying the extracted CT image by the mask image.
[0012] The learning method relating to the seventh aspect involves, in any one of the first to sixth aspects, inputting a simple X-ray image in addition to the 3D CT image input, and generating a fourth pseudo-X-ray image from the simple X-ray image, which is a set of images in which multiple anatomical regions are shown separately. death, The system distinguishes between the second and fourth pseudo-X-ray images, and in updating the parameter values, it further updates the parameter values based on the result of distinguishing between the second and fourth pseudo-X-ray images.
[0013] The learning method according to the eighth aspect, in the seventh aspect, converts the fourth pseudo-X-ray image into a first segmentation label, obtains a second segmentation label for the simple X-ray image, compares the first segmentation label with the second segmentation label, integrates the results of the comparison between the second pseudo-X-ray image and the third pseudo-X-ray image with the results of the comparison between the first segmentation label and the second segmentation label, and updates the parameter values based on the integrated comparison results.
[0014] The learning method according to the ninth aspect, in any one of the first to sixth aspects, inputs a simple X-ray image and dual-energy X-ray images, which are multiple X-ray images taken with X-rays of different energies and in which different anatomical regions of the subject are highlighted. A fourth pseudo-X-ray image is generated from the simple X-ray image, which are multiple images in which multiple anatomical regions are shown separately. The fourth pseudo-X-ray image is reconstructed to generate a fifth pseudo-X-ray image, which are multiple images in which the same anatomical regions as those in the X-ray images constituting the dual-energy X-ray image are highlighted. The dual-energy X-ray image and the fifth pseudo-X-ray image are compared. The results of the comparison between the dual-energy X-ray image and the fifth pseudo-X-ray image are integrated with the results of the comparison between the second pseudo-X-ray image and the third pseudo-X-ray image. In updating the parameter values, the parameters are updated based on the integrated comparison results.
[0015] To achieve the above-mentioned objectives, the image processing method according to the tenth aspect of the present invention acquires a simple X-ray image or a pseudo-X-ray image, transforms the acquired simple X-ray image or pseudo-X-ray image using parameters updated by a learning method according to any one of the first to ninth aspects, and generates separated images, which are multiple images in which multiple anatomical regions of a subject are shown separately.
[0016] The image processing method according to the 11th embodiment converts separated images into segmentation labels.
[0017] To achieve the above-mentioned objectives, a learning device according to a twelfth aspect of the present invention is a learning device comprising a processor, the processor inputs a 3D CT image of a subject, projects the 3D CT image into a 2D space to generate a first pseudo-X-ray image, generates a second pseudo-X-ray image from the first pseudo-X-ray image which is a plurality of images in which a plurality of anatomical regions of the subject are separated and depicted, obtains a label image in which a plurality of anatomical regions are labeled on the 3D CT image, and generates a masked CT image based on the 3D CT image and the label image which is a 3D image in which a plurality of anatomical regions are separated and depicted, The masked CT image is projected into a two-dimensional space to generate a third pseudo-X-ray image, which consists of multiple images showing multiple anatomical regions separated from each other. The second pseudo-X-ray image and the third pseudo-X-ray image are compared, and based on the comparison, the values of the parameters used to generate the second pseudo-X-ray image are updated. The process of inputting a three-dimensional CT image, generating the first pseudo-X-ray image, generating the second pseudo-X-ray image, acquiring a labeled image, generating a masked CT image, generating the third pseudo-X-ray image, comparing the second and third pseudo-X-ray images, and updating the parameter values is repeated until predetermined conditions are met.
[0018] To achieve the above-mentioned objectives, an image processing apparatus according to the 13th aspect of the present invention is an image processing apparatus comprising a processor, the processor acquires a simple X-ray image or a pseudo-X-ray image, and transforms the acquired simple X-ray image or pseudo-X-ray image with parameters updated by a learning method according to any one of the first to 9 aspects to generate a separated image, which is a plurality of images in which a plurality of anatomical regions of a subject are separately depicted.
[0019] In the 14th embodiment, the image processing apparatus, in the 13th embodiment, includes a processor that converts separated images into segmentation labels.
[0020] To achieve the above-mentioned objectives, a learning program according to a 15th aspect of the present invention is a learning program that causes a learning device equipped with a processor to execute a learning method, the learning method being: inputting a 3D CT image of a subject; projecting the 3D CT image into a 2D space to generate a first pseudo-X-ray image; generating a second pseudo-X-ray image from the first pseudo-X-ray image, which is a plurality of images in which a plurality of anatomical regions of the subject are each separated and depicted; obtaining a label image in which a plurality of anatomical regions are labeled on the 3D CT image; and based on the 3D CT image and the label image, obtaining a masked CT image, which is a 3D image in which a plurality of anatomical regions are each separated and depicted. The system generates a third pseudo-X-ray image, which consists of multiple images showing multiple anatomical regions separated, by projecting the generated and masked CT image into a two-dimensional space. The second pseudo-X-ray image and the third pseudo-X-ray image are then compared, and based on the comparison, the values of the parameters used to generate the second pseudo-X-ray image are updated. The processor is then instructed to repeatedly input a 3D CT image, generate the first pseudo-X-ray image, generate the second pseudo-X-ray image, acquire a labeled image, generate a masked CT image, generate the third pseudo-X-ray image, compare the second and third pseudo-X-ray images, and update the parameter values until predetermined conditions are met.
[0021] To achieve the above-mentioned objectives, an image processing program according to the 16th aspect of the present invention is an image processing program that causes an image processing apparatus equipped with a processor to execute an image processing method, wherein the image processing method acquires a simple X-ray image or a pseudo-X-ray image, transforms the acquired simple X-ray image or pseudo-X-ray image with parameters updated by a learning method according to any one of the 1st to 9th aspects, and generates separated images which are multiple images in which multiple anatomical regions of a subject are shown separately.
[0022] In the 17th embodiment, the image processing program further causes the processor to convert the separated images into segmentation labels. [Brief explanation of the drawing]
[0023] [Figure 1] FIG. 1 is a diagram showing a schematic configuration of a learning device according to a first embodiment. [Figure 2] FIG. 2 is a diagram showing a functional configuration of a processor. [Figure 3] FIG. 3 is a diagram showing a state of learning in the first embodiment. [Figure 4] FIG. 4 is a diagram schematically showing a three-dimensional CT image. [Figure 5] FIG. 5 is a diagram schematically showing a first pseudo X-ray image. [Figure 6] FIG. 6 is a diagram schematically showing second to fourth pseudo X-ray images. [Figure 7] FIG. 7 is a diagram schematically showing a label image. [Figure 8] FIG. 8 is a diagram showing a state of generating a masked CT image. [Figure 9] FIG. 9 is a diagram showing a functional configuration of a processor in a second embodiment. [Figure 10] FIG. 10 is a diagram showing a state of learning in the second embodiment. [Figure 11] FIG. 11 is a diagram schematically showing a simple X-ray image. [Figure 12] FIG. 12 is a diagram showing a state of learning in a modification. [Figure 13] FIG. 13 is a diagram showing a functional configuration of a processor in a third embodiment. [Figure 14] FIG. 14 is a diagram showing a state of learning in the third embodiment. [Figure 15] FIG. 15 is a diagram schematically showing a dual energy X-ray image. [Figure 16] FIG. 16 is a diagram showing a state of image processing using a learning result.
MODE FOR CARRYING OUT THE INVENTION
[0024] Embodiments of a learning method, image processing method, learning apparatus, image processing apparatus, learning program, and image processing program according to one aspect of the present invention will be described below. In the description, the accompanying drawings will be referenced as necessary. Note that, for the sake of clarity, some components may be omitted from the accompanying drawings.
[0025] A learning method according to a first aspect of the present invention includes: an image input step of inputting a 3D CT image of a subject; a first image generation step of projecting the 3D CT image into a 2D space to generate a first pseudo-X-ray image; a second image generation step of generating a second pseudo-X-ray image from the first pseudo-X-ray image, which is a plurality of images in which a plurality of anatomical regions of the subject are shown separately; a label image acquisition step of acquiring a label image in which a plurality of anatomical regions are labeled on the 3D CT image; a masked CT image generation step of generating a masked CT image, which is a 3D image in which a plurality of anatomical regions are shown separately, based on the 3D CT image and the label image; and a masked CT image The process includes a third image generation step of projecting the image into a two-dimensional space to generate a third pseudo-X-ray image which is a set of images in which multiple anatomical regions are shown separately; a first comparison step of comparing the second pseudo-X-ray image with the third pseudo-X-ray image; and an update step of updating the parameter values used to generate the second pseudo-X-ray image based on the results of the comparison in the first comparison step. The processor is instructed to repeat the image input step, the first image generation step, the second image generation step, the labeled image acquisition step, the masked CT image generation step, the third image generation step, the first comparison step, and the update step until predetermined conditions are met.
[0026] In the first embodiment, a masked CT image is generated from a 3D CT image, and a third pseudo-X-ray image is generated from the masked CT image. Then, this third pseudo-X-ray image is used as a reference and compared with a second pseudo-X-ray image generated by projecting the original CT image, and the parameters are updated based on the comparison result.
[0027] In the first embodiment, the processor repeats the processing of each step, and learning (parameter updates) progresses. The order of processing during repetition does not necessarily have to be in the order described. For example, various methods such as sequential processing, batch processing, and mini-batch processing can be used. In addition, in the first embodiment, the "predetermined conditions" for terminating the processing can be various conditions such as the number of parameter updates, the comparison result meeting the criteria, or the completion of processing for all images.
[0028] Once the learning process is complete (i.e., the "predetermined conditions" are met and the parameter updates are finished), the parameters described above can be used to separate a two-dimensional X-ray image into various regions.
[0029] In the first embodiment and the following embodiments, multiple images in which multiple anatomical regions are depicted separately may be referred to as "separated images," and images in which multiple anatomical regions are depicted without separation may be referred to as "unseparated images." Furthermore, in the first embodiment and the following embodiments, machine learning techniques such as deep learning may be used for processing such as image generation, label acquisition, and image comparison.
[0030] The learning method according to the second embodiment, in the first embodiment, calculates the error of the second pseudo-X-ray image based on the third pseudo-X-ray image in the first comparison step, and reduces the error by updating the parameters in the update step. According to the second embodiment, the error (the error of the second pseudo-X-ray image based on the third pseudo-X-ray image) is reduced by updating the parameters, and the second pseudo-X-ray image (separated image) can be generated with high accuracy.
[0031] The learning method according to the third embodiment, in the first embodiment, further comprises, in a first comparison step, a step of calculating the error of the second pseudo-X-ray image with respect to the third pseudo-X-ray image; a discrimination step of distinguishing between the third pseudo-X-ray image and the second pseudo-X-ray image; and a step of integrating the error and the discrimination result, and in an update step, the result of the integration is reduced by updating the parameters.
[0032] The learning method according to the fourth embodiment, in the first embodiment, further comprises: a first comparison step of calculating the error of a second pseudo-X-ray image with respect to a third pseudo-X-ray image; a discrimination step of distinguishing between paired data of the first pseudo-X-ray image and the third pseudo-X-ray image and paired data of the first pseudo-X-ray image and the second pseudo-X-ray image; and a step of integrating the error and the discrimination result, wherein the result of the integration is reduced by updating the parameters in the update step.
[0033] The learning method relating to the fifth aspect acquires a label image in any one of the first to fourth aspects by generating a label image from a 3D CT image during the label image acquisition step. The fifth aspect defines one aspect of the label image generation method.
[0034] The learning method according to the sixth embodiment, in any one of the first to fifth embodiments, generates a mask image from extracted CT images obtained by region extraction from a 3D CT image in the masked CT image generation step, and generates a masked CT image by multiplying the CT image by the mask image. The sixth embodiment defines one embodiment of a masked CT image generation method.
[0035] The learning method according to the seventh embodiment further comprises, in any one of the first to sixth embodiments, an image input step in which a simple X-ray image is further input, and a fourth image generation step in which a fourth pseudo-X-ray image is generated from the simple X-ray image, which is a plurality of images in which a plurality of anatomical regions are separately depicted, and a discrimination step in which the second pseudo-X-ray image and the fourth pseudo-X-ray image are distinguished, and in the update step, the parameter values are updated based on the result of the discrimination.
[0036] In the seventh embodiment, the CT image and the plain X-ray image may be taken in different positions. For example, CT images are often taken in a supine position, and plain X-ray images are often taken in an upright position, but even in such cases, as in the seventh embodiment, the plain X-ray image may be taken in a different position. lineBy incorporating image-based pseudo-X-ray images (fourth pseudo-X-ray image) into parameter updates, the effects of domain shift can be reduced.
[0037] In the seventh embodiment, it is preferable that the 3D CT image and the plain X-ray image are images taken of the same subject; however, the seventh embodiment can also be applied when these images are images taken of different subjects.
[0038] The learning method according to the eighth embodiment, in the seventh embodiment, further causes the processor to perform a first label acquisition step of acquiring a first segmentation label for a simple X-ray image, a second label acquisition step of converting a fourth pseudo-X-ray image into a second segmentation label, a label comparison step of comparing the first segmentation label and the second segmentation label, and a first integration step of integrating the comparison results from the first comparison step and the comparison results from the label comparison step, and in the update step, updates the parameters based on the comparison results integrated in the first integration step. In the eighth embodiment, the method for integrating the comparison results is not particularly limited, but for example, simple addition, weighted addition, averaging, etc., can be performed.
[0039] The learning method according to the ninth embodiment, in any one of the first to sixth embodiments, further involves the processor performing the following steps: an image input step inputs a simple X-ray image and a dual-energy X-ray image, which is a set of X-ray images taken with different energy levels and highlighting different anatomical regions of the subject; a fourth image generation step generates a fourth pseudo-X-ray image from the simple X-ray image, which is a set of images in which multiple anatomical regions are shown separately; a fifth image generation step reconstructs the fourth pseudo-X-ray image to generate a fifth pseudo-X-ray image, which is a set of images in which the same anatomical regions as those in the X-ray images constituting the dual-energy X-ray image are highlighted; a second comparison step compares the dual-energy X-ray image and the fifth pseudo-X-ray image; and a second integration step integrates the results of the comparison in the second comparison step with the results of the comparison in the first comparison step; and in the update step, the parameters are updated based on the integrated comparison results.
[0040] According to the ninth aspect, X-rays with different energies, covered Dual-energy X-ray images, which are multiple X-ray images emphasizing different anatomical regions of the specimen (e.g., bone and soft tissue), are input, and these dual-energy X-ray images are compared with a fifth pseudo-X-ray image (second comparison step). In this case, both images can be taken in an upright position, in which case no domain shift occurs due to posture. The comparison results from the second comparison step are then integrated with the comparison results from the first comparison step, and the parameters are updated based on the integrated comparison results. According to the seventh embodiment, by combining dual-energy X-ray images and pseudo-X-ray images complementaryly in this way, it is possible to separate the X-ray image into many regions while reducing domain shift due to posture. In the ninth embodiment, the method for integrating the comparison results is not particularly limited, but for example, simple addition, weighted addition, averaging, etc., can be used.
[0041] To achieve the above-mentioned objectives, the image processing method according to the tenth aspect of the present invention includes an image input step of acquiring a simple X-ray image or a pseudo-X-ray image, and processing the acquired simple X-ray image or pseudo-X-ray image from first to third... 9The processor is made to perform a separated image generation step, which involves converting a plain X-ray image or a pseudo-X-ray image using parameters updated by a learning method according to any one of the embodiments, thereby generating separated images, which are multiple images in which multiple anatomical regions of a subject are shown separately. In the tenth embodiment, a plain X-ray image or a pseudo-X-ray image is converted using parameters updated by a learning method according to any one of the first to ninth embodiments. Convert This allows for the acquisition of separate images of multiple anatomical regions.
[0042] The image processing method according to the 11th embodiment further causes the processor to perform a label acquisition step that converts the separated images into segmentation labels, in accordance with the 10th embodiment. The 11th embodiment applies the image processing method according to the 10th embodiment to segmentation.
[0043] To achieve the above-mentioned objectives, a learning device according to a twelfth aspect of the present invention is a learning device comprising a processor, the processor comprising: an image input process for inputting a 3D CT image of a subject; a first image generation process for projecting the 3D CT image into a 2D space to generate a first pseudo-X-ray image; a second image generation process for generating a second pseudo-X-ray image from the first pseudo-X-ray image, which is a plurality of images in which a plurality of anatomical regions of the subject are separated and depicted; a label image acquisition process for acquiring a label image in which a plurality of anatomical regions are labeled on the 3D CT image; and a masked CT image, which is a 3D image in which a plurality of anatomical regions are separated and depicted, based on the 3D CT image and the label image. The system performs the following steps: generating a masked CT image; a third image generation process that projects the masked CT image into a two-dimensional space to generate a third pseudo-X-ray image, which is a set of images in which multiple anatomical regions are separated; a first comparison process that compares the second pseudo-X-ray image with the third pseudo-X-ray image; an update process that updates the parameter values used to generate the second pseudo-X-ray image based on the results of the comparison in the first comparison process; and a loop process that repeats the image input process, the first image generation process, the second image generation process, the label image acquisition process, the masked CT image generation process, the third image generation process, the first comparison process, and the update process until predetermined conditions are met.
[0044] According to the twelfth embodiment, similar to the first embodiment, a two-dimensional X-ray image can be separated into various regions. In addition, in the twelfth embodiment, the processor may be made to perform further processing similar to that in the second to ninth embodiments.
[0045] To achieve the above-mentioned objectives, an image processing apparatus according to a thirteenth aspect of the present invention is an image processing apparatus comprising a processor, wherein the processor performs an image input process to acquire a simple X-ray image or a pseudo-X-ray image, and processes the acquired simple X-ray image or pseudo-X-ray image from first to third... 9The system performs a separation image generation process, which involves converting parameters updated by a learning method relating to any one of the embodiments to generate separate images, which are multiple images in which multiple anatomical regions of the subject are shown separately. According to the 13th embodiment, separate images of multiple anatomical regions can be obtained, similar to the 10th embodiment.
[0046] The image processing apparatus according to the 14th embodiment, in the 13th embodiment, further causes a processor to perform a label acquisition process that converts separated images into segmentation labels. The 14th embodiment is similar to the 11th embodiment in that the image processing apparatus according to the 13th embodiment is applied to segmentation.
[0047] To achieve the above-mentioned objectives, a learning program according to a 15th aspect of the present invention is a learning program that causes a learning device equipped with a processor to execute a learning method, the learning method comprising: an image input step of inputting a 3D CT image of a subject; a first image generation step of projecting the 3D CT image into a 2D space to generate a first pseudo-X-ray image; a second image generation step of generating a second pseudo-X-ray image from the first pseudo-X-ray image, which is a plurality of images in which a plurality of anatomical regions of the subject are separated and depicted; a label image acquisition step of acquiring a label image in which a plurality of anatomical regions are labeled on the 3D CT image; and a masked C image, which is a 3D image in which a plurality of anatomical regions are separated and depicted based on the 3D CT image and the label image. The process involves a masked CT image generation step that generates a T image, a third image generation step that projects the masked CT image into a two-dimensional space to generate a third pseudo-X-ray image which is a set of images in which multiple anatomical regions are shown separately, a first comparison step that compares the second pseudo-X-ray image with the third pseudo-X-ray image, an update step that updates the parameter values used to generate the second pseudo-X-ray image based on the results of the comparison in the first comparison step, and causing the processor to repeat the image input step, the first image generation step, the second image generation step, the label image acquisition step, the masked CT image generation step, the third image generation step, the first comparison step, and the update step until predetermined conditions are met.
[0048] According to the 15th aspect, similar to the first and tenth aspects, a two-dimensional X-ray image can be separated into various regions. The learning method executed by the learning program according to the present invention may have the same configuration as the second to ninth aspects. Furthermore, a non-temporary recording medium on which the computer-readable code of the learning program described above is recorded can also be cited as an aspect of the present invention.
[0049] To achieve the above-mentioned objectives, an image processing program according to the 16th aspect of the present invention is an image processing program that causes an image processing device equipped with a processor to execute an image processing method, the image processing method comprising: an image input step of acquiring a simple X-ray image or a pseudo-X-ray image; and a separated image generation step of converting the acquired simple X-ray image or pseudo-X-ray image with parameters updated by a learning method according to any one of the first to 9th aspects to generate separated images, which are multiple images in which multiple anatomical regions of a subject are separately depicted.
[0050] According to the 16th aspect, separate images of multiple anatomical regions can be obtained, similar to the 10th and 13th aspects.
[0051] The image processing program according to the 17th embodiment, in the 16th embodiment, further causes the processor to execute a label acquisition step that converts separated images into segmentation labels. The 17th embodiment, like the 11th and 14th embodiments, applies the image processing apparatus according to the 16th embodiment to segmentation. The image processing method executed by the image processing program according to the present invention may have the same configuration as the 2nd to 9th embodiments. Furthermore, a non-temporary recording medium on which the computer-readable code of the learning program described above is recorded can also be cited as an embodiment of the present invention.
[0052] [First Embodiment] [Configuration of the learning device] Figure 1 is a diagram showing the schematic configuration of a learning device 10 (learning device, image processing device) according to the first embodiment. The learning device 10 includes a processor 100 (processor), a storage device 200, a display device 300, an operation unit 400, and a communication unit 500. The connections between these components may be wired or wireless. Furthermore, these components may be housed in a single enclosure or in separate enclosures.
[0053] [Processor Functional Configuration] Figure 2 shows the functional configuration of the processor 100. As shown in the figure, the processor 100 includes an image input unit 102, a first image generation unit 104, a second image generation unit 106, a label image acquisition unit 108, a masked CT image generation unit 110, a third image generation unit 112, a first comparison unit 114, a parameter update unit 116, a learning control unit 118, a display control unit 120, and a recording control unit 122. Learning in the learning device 10 is mainly performed using the image input unit 102 to the learning control unit 118. The display control unit 120 causes the display device 300 to display various images (simple X-ray images, pseudo-X-ray images, 3D CT images, etc.), and the recording control unit 122 controls the recording of images and data from or to the storage device 200.
[0054] The functions of the processor 100 described above are various processors and recording media. This can be achieved using various types of processors, such as the CPU (Central Processing Unit), which is a general-purpose processor that executes software (programs) to perform various functions. These are processors specialized in image processing, such as GPUs (Graphics Processing Units) and FPGAs (Field Programmable Gate Arrays), which allow for changes to the circuit configuration after manufacturing. This also includes a certain programmable logic device (PLD). Each function may be implemented by a single processor, or by multiple processors of the same or different types (for example, multiple FPGAs, or a combination of CPU and FPGA, or a combination of CPU and GPU). Alternatively, multiple functions may be implemented by a single processor. More specifically, the hardware structure of these various processors is an electrical circuit made up of circuit elements such as semiconductor devices.
[0055] When the aforementioned processor or electrical circuit executes software (program), it stores code readable by the computer (for example, the various processors and electrical circuits constituting the processor 100, and / or combinations thereof) in a non-temporary recording medium (memory) such as flash memory or ROM (Read Only Memory) (not shown), and the computer refers to the software. The program to be executed includes a program (learning program, image processing program) that executes a method (learning method, image processing method) according to one aspect of the present invention. Also, when the software is executed, information (images and other data) stored in the storage device 200 is used as needed. Also, when executing, for example, RAM (Random Access Memory) (not shown) is used as a temporary storage area.
[0056] [Information stored in memory] The memory device 200 is composed of various magneto-optical recording media, semiconductor memories, and their control units, and stores CT images and X-ray images (simple X-ray images, real X-ray images) actually taken by photographing a subject, pseudo-X-ray images generated artificially (1st to 5th pseudo-X-ray images), and software executed by the aforementioned processor (including a learning program and an image processing program according to one aspect of the present invention).
[0057] [Configuration of display device, control unit, and communication unit] The display device 300 consists of a device such as a liquid crystal monitor and can display CT images, X-ray images, learning and image processing results, etc. The operation unit 400 consists of a mouse, keyboard, etc. (not shown), and the user can give instructions necessary for executing medical image processing methods and learning methods via the operation unit 400. The user can give instructions via the screen displayed on the display device 300. The display device 300 may also be configured as a touch panel type monitor, and the user may give instructions via the touch panel. The communication unit 500 can acquire CT images, X-ray images, and other information from other systems connected via a network (not shown).
[0058] [Learning method using a learning device] Next, the learning method (machine learning method) using the learning device 10 will be described. Figure 3 shows the learning process in the first embodiment. In Figure 3, various images are shown with solid lines, and the processing of the images is shown with dotted lines (the same applies to the figures described later).
[0059] [Image Input] The image input unit 102 (processor) inputs a 3D CT image of the subject (step S100; image input step, image input processing). Figure 4 is a schematic diagram of the 3D CT image 700 (3D CT image, X-ray CT image), showing the lungs, heart, and aorta as anatomical regions. In the 3D CT image 700, sections 700S, 700C, and 700A are cross-sections in the sagittal, coronal, and axial directions, respectively.
[0060] [Generation of the first pseudo-X-ray image] The first image generation unit 104 (processor) projects the 3D CT image 700 into a 2D space to generate a first pseudo-X-ray image (step S102; first image generation step, first image generation process). The first image generation unit 104 can perform image projection using various known methods. Figure 5 is a schematic diagram showing the first pseudo-X-ray image. The first pseudo-X-ray image is a non-separated image in which multiple anatomical regions (bones, lungs, heart, and aorta in Figure 5) are shown together.
[0061] [Generation of a second pseudo-X-ray image] The second image generation unit 106 (processor) generates a second pseudo-X-ray image 720 (second pseudo-X-ray image) from the first pseudo-X-ray image 710, which is a set of images (separated images) showing multiple anatomical regions of the subject separated from each other (step S104; second image generation step, second image generation process). Figure 6 shows an example of a pseudo-X-ray image, with parts (a) to (d) of the figure showing images of the bone, heart, lungs, and aorta, respectively. The third pseudo-X-ray image 750 and the fourth pseudo-X-ray image 770, which will be described later, are also images showing the bone, heart, lungs, and aorta, similar to Figure 6.
[0062] The second image generation unit 106 is a generator or converter that receives the input of the first pseudo-X-ray image 710 and generates a second pseudo-X-ray image 720, and can be configured using machine learning techniques such as deep learning. Specifically, for example, a neural network such as CNN (Convolutional Neural Network), SVM (Support Vector Machine), or U-Net (a type of FCN (Fully Convolutional Network)) used in pix2pix can be used to generate the second image. It is possible to configure the adult unit 106. Furthermore, it is based on ResNet (Deep Residual Network). A neural network may be used. These neural networks are merely examples of the configuration of the second image generation unit 106 and do not limit the methods used for image generation or transformation in one embodiment of the present invention.
[0063] An example of a layer configuration when the second image generation unit 106 is constructed using a CNN is described below. The CNN includes an input layer, an intermediate layer, and an output layer. The input layer takes the first pseudo-X-ray image 710 as input and outputs features. The intermediate layer includes a convolutional layer and a pooling layer, and calculates other features by taking the features output by the input layer as input. These layers have a structure in which multiple "nodes" are connected by "edges" and hold multiple weight parameters. The values of the weight parameters change as learning progresses. The output layer outputs the recognition result of which anatomical region (in this example, bone, heart, lung, or aorta) each pixel of the input image belongs to, based on the features output from the intermediate layer. By combining only the pixels belonging to a specific anatomical region into a single image, images for each anatomical region, as shown in Figure 6, can be obtained.
[0064] [Generation of a third pseudo-X-ray image] The label image acquisition unit 108 (processor) extracts multiple anatomical regions of the subject from the 3D CT image 700 and labels them pixel by pixel (region extraction), thereby acquiring the extracted CT image 730 (3D image; a label image in which multiple anatomical regions are labeled) (step S108; label image acquisition step, label image acquisition process). Figure 7 is a schematic diagram of a label image, showing labels attached to different anatomical regions in different shaded states.
[0065] The label image acquisition unit 108 can perform labeling using known image processing techniques. Alternatively, the label image acquisition unit 108 may acquire pre-labeled images (which may be images labeled by a device other than the learning device 10, or images labeled based on user operations).
[0066] The masked CT image generation unit 110 (processor) generates a masked CT image, which is a three-dimensional image in which multiple anatomical regions are shown separately, based on the three-dimensional CT image 700 and the extracted CT image 730 (label image) described above (step S110; masked CT image generation step, masked CT image generation process).
[0067] In step S110, the masked CT image generation unit 110 generates a mask image from the extracted CT image 730. The mask image is a three-dimensional image in which, for example, pixels belonging to a specific anatomical region have a pixel value of 1, and pixels in other regions have a pixel value of 0 (zero). The masked CT image generation unit 110 creates this mask image for each anatomical region. Then, the masked CT image generation unit 110 multiplies the extracted CT image 730 and the mask image for each anatomical region to generate a masked CT image, which is a three-dimensional image in which multiple anatomical regions are shown separately.
[0068] Figure 8 shows the process of generating a masked CT image. Part (a) of Figure 8 shows the masked image 735 for the lung region 736, in which the pixel value for the lung region 736 is 1, and the pixel value for the other regions is 0 (zero). In Figure 8, regions with a pixel value of 0 are indicated by diagonal lines. Sections 735S, 735C, and 735A are cross-sections in the sagittal, coronal, and axial directions, respectively. Part (b) of Figure 8 is a masked CT image 742 of the lung, generated by multiplying the 3D CT image 700 by the masked image 735, in which only the lung region 743 is extracted from the 3D CT image 700.
[0069] The masked CT image generation unit 110 can generate a masked CT image 740, which is a three-dimensional image in which multiple anatomical regions are shown separately, by performing the same processing not only on the lungs but also on other anatomical regions such as bones, the heart, and the aorta.
[0070] The third image generation unit 112 (processor) projects the masked CT image 740 into a two-dimensional space to generate a third pseudo-X-ray image 750, which is a set of images in which multiple anatomical regions are separated (step S112; third image generation step, third image generation process). The third pseudo-X-ray image 750 is a set of two-dimensional images in which the same anatomical regions as the second pseudo-X-ray image 720 described above are separated (separation is performed for the same anatomical regions in order to compare the second pseudo-X-ray image 720 and the third pseudo-X-ray image 750, as will be described later).
[0071] [Image comparison] The first comparison unit 114 (processor) compares the second pseudo-X-ray image 720 and the third pseudo-X-ray image 750 (step S106; first comparison step, first comparison process). The first comparison unit 114 can compare the images by taking the difference in pixel values for each pixel and adding the differences, or by weighted addition, etc.
[0072] [Parameter update] The parameter update unit 116 (processor) updates the parameter values used by the second image generation unit 106 to generate the second pseudo-X-ray image 720 based on the comparison results of step S106 (step S114; update step, update process). The parameter update unit 116 can reduce the error by, for example, calculating the error of the second pseudo-X-ray image 720 relative to the third pseudo-X-ray image 750 in step S106 (first comparison step) and updating the parameters in step S114 (update step). Such a method is equivalent to learning without using a configuration such as a Generative Adversarial Network (GAN).
[0073] In the learning according to the first embodiment, the parameters may be updated by considering the discrimination result in addition to the image error. In this case, the parameter update unit 116 (or the first comparison unit 114) may, for example, in step S106 (first comparison step), calculate the error of the second pseudo-X-ray image 720 with respect to the third pseudo-X-ray image 750, and discriminate between the third pseudo-X-ray image 750 and the second pseudo-X-ray image 720, integrate the error and the discrimination result, and update the parameters in step S114 (update step) to reduce the integrated result. Such a method is used in learning using a Conditional GAN (when using a classifier without correspondence). It is equivalent to (combined).
[0074] Furthermore, in the learning according to the first embodiment, the parameters may be updated by considering the pair data discrimination result in addition to the image error. In this case, the parameter update unit 116 (or the first comparison unit 114) may, for example, in step S106 (first comparison step), calculate the error of the second pseudo-X-ray image 720 with respect to the third pseudo-X-ray image 750, and discriminate between the pair data of the first pseudo-X-ray image 710 and the third pseudo-X-ray image 750 and the pair data of the first pseudo-X-ray image 710 and the second pseudo-X-ray image 720, integrate the error and the discrimination result, and update the parameters in step S114 (update step) to reduce the integrated result. Such a method is used in learning with Conditional GAN (recognition with correspondence). This is equivalent to using a separate device.
[0075] [Image discrimination and integration with comparison results] In the image "discrimination" described above, it is possible to determine whether "the images are identical or different" or "whether the images are genuine or fake." However, there are variations in the methods used for discrimination "image by image" or "patch by patch." The latter is called a "patch discriminator" and is proposed in the non-patent literature below.
[0076] [Non-Patent Literature 3] “Image-to-Image Translation with Conditional Adversarial Networks”, ICCV, 2017, Philip Isola et al., [Retrieved June 24, 2021], Internet (https: / / openaccess.thecvf.com / content_cvpr_2017 / papers / Isola_Image-To-Image_Translation_With_CVPR_2017_paper.pdf) In "image comparison" as in step S106, for example, the first comparison unit 114 calculates the absolute error (or squared error, etc.) for each pixel, calculates its average (or sum, etc.), and outputs a scalar value. On the other hand, in "image discrimination" or "pair data discrimination," in the case of a patch discriminator, the classification error of "real or fake" is calculated for each patch, calculates its average (or sum, etc.), and outputs a scalar value. Then, in the "integration" described above, the parameter update unit 116 (or the first comparison unit 114) calculates the weighted sum of these scalar values (the result of integration). The parameter update unit 116 reduces this weighted sum by updating the parameters.
[0077] The comparison and classification of images and the integration of results described above can also be performed in the second and third embodiments described later (see, for example, steps S106, S116, S118 in Figure 10, and steps S106, S122, S124 in Figure 14).
[0078] [Learning control] The learning control unit 118 (processor) causes each part of the processor 100 to repeat the processing of steps S100 to S114 until predetermined conditions are met. By repeating these processes, learning (parameter updates) progresses. Various methods can be used for repeating the processing, such as sequential processing, batch processing, and mini-batch processing. The "predetermined conditions" for terminating learning (parameter updates) can include the number of parameter updates, whether the comparison results meet the criteria, or the completion of processing for all images.
[0079] Once the learning process is complete (i.e., the "predetermined conditions" are met and the parameter updates are finished), the two-dimensional X-ray image can be separated into various regions using the second image generation unit 106 and the updated parameters, as will be described later. In other words, the second image generation unit 106 in the completed learning state can be understood as a "trained model".
[0080] [Second Embodiment] A second embodiment of the learning method, learning device, and learning program of the present invention will now be described. Figure 9 shows the functional configuration of the processor 100A (processor) in the second embodiment. In addition to the functions of the processor 100 in the first embodiment (image input unit 102 to recording control unit 122), the processor 100A has a fourth image generation unit 124 and an image discrimination unit 126. In the second embodiment, the same reference numerals are used for components and processing contents as in the first embodiment, and detailed explanations are omitted.
[0081] [Learning in the second embodiment] Figure 10 shows the learning process in the second embodiment. In the second embodiment, in addition to the 3D CT image 700, a simple X-ray image 760 is used for learning.
[0082] The image input unit 102 receives a plain X-ray image 760 (plain X-ray image) in addition to the 3D CT image 700 (step S100A; image input step, image input processing). The plain X-ray image 760 is a 2D real image obtained by actually photographing the subject, as shown in the schematic diagram in Figure 11.
[0083] [Reducing domain shifts] Note that plain X-ray images are often taken in an upright position, while 3D CT images are often taken in a supine position. Therefore, processing to reduce the domain shift between the 3D CT image 700 and the plain X-ray image 760 may be performed. Domain shift reduction can be achieved, for example, by the following method (medical image processing method).
[0084] (Method 1) A medical image processing method performed by a medical image processing apparatus equipped with a processor, wherein the processor A reception process for receiving input of the first medical image actually taken in the first posture, An image generation process for generating a second medical image from a first medical image, wherein the second medical image is taken from a second medical image in a second pose different from the first medical image, and the second medical image is generated using a deformation vector field that transforms the first medical image into the second medical image. A medical image processing method that performs the following.
[0085] In the image generation step of the above method, the processor is a generator that outputs a deformed vector field when a first medical image is input, and it is preferable to generate the deformed vector field using a generator learned by machine learning. Furthermore, in the image generation step, it is preferable that the processor applies the deformed vector field to the first medical image to generate a second medical image. The processor may also further perform a modality conversion step in which it pseudo-converts the pseudo-generated second medical image into a third medical image using a different modality than the second medical image.
[0086] In the above method, for example, the first medical image can be a 3D CT image actually taken in a supine position (first position: supine), and the second medical image can be a pseudo-CT image in an upright position (second position: upright). Furthermore, the third medical image can be a pseudo-X-ray image in an upright position. The "deformation vector field" is a collection of vectors representing the displacement (image deformation) from each voxel (or pixel) in the first medical image to each voxel in the second medical image.
[0087] Furthermore, processor 100A can be used as the "processor" in the above-described method. Also, the "generation unit" in the above-described method can be configured using a neural network, similar to the second image generation unit 106.
[0088] While it is preferable that the plain X-ray image 760 is an image of the same subject as the 3D CT image 700, one aspect of the present invention can also be applied when the two images are of different subjects. This is because, in general, it is rare to actually obtain images of the same part of the same subject from different poses. Furthermore, it is preferable that the plain X-ray image 760 and the 3D CT image 700 are images of the same part.
[0089] [Generation of the fourth pseudo-X-ray image] As described above in the first embodiment, the second image generation unit 106 generates a second pseudo-X-ray image 720 from the first pseudo-X-ray image 710 (step S104). In the second embodiment, the fourth image generation unit 124 (processor) further generates a fourth pseudo-X-ray image 770 (fourth pseudo-X-ray image) from the simple X-ray image 760, which is a plurality of images (separated images) in which a plurality of anatomical regions of the subject (the same anatomical regions as the second pseudo-X-ray image 720) are separately depicted (step S104A; fourth image generation step, fourth image generation process). The fourth image generation unit 124 can be configured using machine learning methods such as deep learning, similar to the second image generation unit 106, and can use the same network configuration and the same parameters (weight parameters, etc.) as the second image generation unit 106. In other words, the fourth image generation unit 124 can be considered identical to the second image generation unit 106.
[0090] [Image Recognition] The image discrimination unit 126 (processor) distinguishes between the third pseudo-X-ray image 750 and the fourth pseudo-X-ray image 770 described above (step S116: discrimination step, discrimination process). Specifically, the image discrimination unit 126 attempts to distinguish (identify) between the third pseudo-X-ray image 750 and the fourth pseudo-X-ray image as "ground truth images," and operates like a discriminator in a Generative Adversarial Network (GAN).
[0091] [Parameter update] The parameter update unit 116 (processor) integrates the comparison result in step S106 and the discrimination result in step S116 (step S118; first integration step, first integration process), and updates the parameters used by the second image generation unit 106 and the fourth image generation unit 124 based on the integration result (step S114A; update step, update process). The parameter update unit 116 updates the parameter values, for example, so that the image discrimination unit 126 can no longer distinguish between the third pseudo-X-ray image 750 and the fourth pseudo-X-ray image. That is, the second image generation unit 106 and the fourth image generation unit 124 operate like generators in a GAN. Also, similar to the first embodiment described above, the second image generation unit 106 and the fourth image generation unit 124 in a state where training has been completed can be understood as a "trained model".
[0092] Thus, in the second embodiment, by updating the parameter values based on the discrimination results, the generation of the separated images described later can be performed with high accuracy.
[0093] [Modified version of the second embodiment] A modified version of the second embodiment described above will now be explained. In the second embodiment, a fourth pseudo-X-ray image 770 generated from a simple X-ray image 760 is used for training, but in this modified version, labeled images are used for training. Figure 12 shows the training process in the modified version. In the example shown in the figure, the processing of the 3D CT image 700 is the same as in Figure 10, and the illustration and detailed explanation are omitted, but the generation and comparison of the second pseudo-X-ray image 720 and the third pseudo-X-ray image 750 are performed in the same way as in Figure 10.
[0094] In the modified example shown in Figure 12, a label image generation unit (not shown) of the processor 100A generates a first label image 772 (first segmentation label) from the fourth pseudo-X-ray image 770 (step S105; first label acquisition step, first label acquisition process). Meanwhile, the label image generation unit segments the simple X-ray image 760 using methods such as snakes (dynamic contour method) or graph cuts (step S107; segmentation step, segmentation process) to generate a second label image 774 (second segmentation label) (step S109; second label acquisition step, second label acquisition process).
[0095] Then, the second comparison unit (not shown) of the processor 100A compares the first label image 772 with the second label image 774 (step S117; label comparison step, label comparison process). This comparison result is integrated with the comparison result of the second pseudo-X-ray image 720 with the third pseudo-X-ray image 750, similar to step S106 in Figure 10 (step S118A; first integration step, first integration process), and the parameter update unit 116 updates the parameters based on the integration result (step S114A; update step, update process).
[0096] In the modified version, learning can be performed to separate two-dimensional X-ray images into various regions, similar to the first and second embodiments described above. In the modified version, instead of using dual-energy X-ray images as in the third embodiment described later, constraints on the outline of anatomical regions can be added using segmentation labels of simple X-ray images, thereby reducing the effect of domain shift due to posture. In steps S105 and S109 of the modified version, segmentation labels processed by a device other than the learning device 10 or segmentation labels assigned by the user may be obtained.
[0097] [Third Embodiment] A third embodiment of the learning method, learning device, and learning program of the present invention will now be described. The third embodiment further utilizes dual-energy X-ray images, which are multiple X-ray images taken with X-rays of different energies, each highlighting a different anatomical region of the subject.
[0098] [Processor Functional Configuration] Figure 13 shows the functional configuration of the processor 100B (processor) according to the third embodiment. In addition to the functions of the processor 100 in the first embodiment (image input unit 102 to recording control unit 122), the processor 100B further includes a fourth image generation unit 124, a fifth image generation unit 140, a second comparison unit 142, and a second integration unit 144. In the third embodiment, the same reference numerals are used for components and processing contents as in the first and second embodiments, and detailed explanations are omitted.
[0099] [Learning process] Figure 14 shows the learning process in the third embodiment. The image input unit 102 further inputs a simple X-ray image 760 (simple X-ray image) and a dual-energy X-ray image 790 in addition to the 3D CT image 700 (step S100B; image input step, image input processing). As shown in the schematic diagram of Figure 11, the simple X-ray image 760 is a 2D image obtained by actually photographing the subject, and the dual-energy X-ray image 790 is a plurality of X-ray images taken with X-rays of different energies, each highlighting different anatomical regions of the subject. The simple X-ray image 760 and the dual-energy X-ray image 790 can be taken in the same posture (e.g., standing). Furthermore, with the dual-energy X-ray image 790, depending on the energy of the X-rays used for photography, it is possible to obtain an image in which bone is highlighted (part (a) in the same figure) and an image in which soft tissue (part (b) in the same figure) is highlighted, for example, as shown in the schematic diagram of Figure 15.
[0100] While it is preferable that the simple X-ray image 760 and the dual-energy X-ray image 790 are images of the same subject as the three-dimensional CT image 700, embodiments of the present invention can also be applied when the two images are of different subjects.
[0101] The fourth image generation unit 124 (processor) acquires a fourth pseudo-X-ray image, similar to the second embodiment (step S104B; fourth image generation step, fourth image generation process). As described above for the second embodiment, the fourth image generation unit 124 can be configured using machine learning techniques such as deep learning, similar to the second image generation unit 106, and can use the same network configuration and the same parameters (weight parameters, etc.) as the second image generation unit 106. In other words, the fourth image generation unit 124 can be considered identical to the second image generation unit 106.
[0102] The fifth image generation unit 140 (processor) reconstructs the fourth pseudo-X-ray image 770 to generate a fifth pseudo-X-ray image 780, which is a set of images in which the same anatomical regions as the X-ray images constituting the dual-energy X-ray image are emphasized (step S120; fifth image generation step, fifth image generation process). Specifically, for example, if the fourth pseudo-X-ray image 770 is an image of bone, heart, lungs, and aorta as shown in the schematic diagram of Figure 6, and the dual-energy X-ray image 790 is an image of bone and soft tissue as shown in the schematic diagram of Figure 15, then the fifth image generation unit 140 generates the fifth pseudo-X-ray image 780, which consists of images of bone and soft tissue, in accordance with the separation of the dual-energy X-ray image 790.
[0103] The second comparison unit 142 (processor) compares the dual-energy X-ray image 790 with the fifth pseudo-X-ray image 780 (step S122; second comparison step, second comparison process), and the second integration unit 144 (processor) integrates the comparison result with the result of the comparison in step S106 described above (comparison of the third pseudo-X-ray image with the second pseudo-X-ray image) (step S124; second integration step, second integration process). The second integration unit 144 can integrate the comparison results by simple addition, weighted addition, etc.
[0104] The dual-energy X-ray image 790 and the fifth pseudo-X-ray image 780 are, for example, separated images of "bone and soft tissue." When comparing such images, the second comparison unit 142 (processor) calculates the absolute error (or squared error, etc.) for each pixel in both "bone" and "soft tissue," calculates the average (or sum, etc.) of these errors, and outputs a scalar value. Furthermore, the second comparison unit 142 averages the scalar values in "bone" and "soft tissue" to calculate a final scalar value.
[0105] On the other hand, the second pseudo-X-ray image 720 and the third pseudo-X-ray image 750 are separated images of, for example, "bones, lungs, heart, and aorta." In comparing such images (step S106), the parameter update unit 116 (or the first comparison unit 114) calculates scalar values in the same way as in the case of "bones and soft tissues," and the second integration unit 144 (processor) calculates, for example, a weighted sum of these scalar values (the result of the integrated comparison). Note that the "weights" here are not learned parameters, but empirically determined hyperparameters.
[0106] The parameter update unit 116 (processor) updates the parameters used by the second image generation unit 106 to generate the pseudo-X-ray image based on the results of the integrated comparison (step S114B; update step, update process). By updating the parameters, the parameter update unit 116 can bring the fifth pseudo-X-ray image 780 closer to the dual-energy X-ray image 790.
[0107] The method using pseudo-X-ray images generated from 3D CT images can separate not only bone and soft tissue but also numerous other regions, while the method using dual-energy X-ray images (which can be acquired in the same posture as simple X-ray images) does not cause domain shift due to posture. In the third embodiment, by combining dual-energy X-ray images and pseudo-X-ray images complementaryly in this way, it is possible to learn how to separate 2D X-ray images into various regions while reducing domain shift caused by the acquisition posture. As with the first and second embodiments, the second image generation unit 106 in a state where learning has been completed can be recognized as a trained model.
[0108] [Application to image processing] A learning method, learning apparatus, and learning program according to one aspect of the present invention can be applied to image processing (image processing method, image processing apparatus, and image processing program). Specifically, by applying parameters updated by the learning described in the first to third embodiments and inputting a two-dimensional X-ray image to the second image generation unit 106 (processor, converter), separated images can be obtained by separating the two-dimensional X-ray image into various regions.
[0109] Figure 16 shows an example of generating separated images. In the example shown in Figure 16, a separated image (corresponding to the second pseudo-X-ray image 720) is generated by inputting a simple X-ray image 760 into the second image generation unit 106 (processor, converter, trained model) which has completed training (step S114; separated image generation step, separated image generation process). Note that a pseudo-X-ray image may be input instead of a simple X-ray image 760 (real X-ray image).
[0110] Thus, the learning device 10, once the learning process is complete, can also function as an image processing device that generates separated images. Alternatively, an image processing device that generates separated images can be configured by transferring the second image generation unit 106, which has completed the learning process, to a separate image processing device (using the same network configuration and the same parameters). In such an image processing device, separated images can be generated using an image processing method and an image processing program according to one aspect of the present invention.
[0111] Furthermore, an image processing method, image processing apparatus, and image processing program according to one aspect of the present invention can also be applied to image segmentation. Specifically, a processor such as the second image generation unit 106 (processor, converter, trained model) can convert separated images into segmentation labels (label acquisition step, label acquisition process).
[0112] While embodiments of the present invention have been described above, the present invention is not limited to the embodiments described above, and various modifications are possible without departing from the spirit of the invention. [Explanation of Symbols]
[0113] 10 Learning device 100 processors 100A Processor 100B Processor 102 Image Input Section 104 First image generation unit 106 Second image generation unit 108 Label image acquisition unit 110 Masked CT Image Generation Unit 112 Third image generation unit 114 First Comparison Section 116 Parameter update section 118 Learning Control Unit 120 Display Control Unit 122 Recording Control Unit 124 Fourth Image Generation Unit 126 Image discrimination unit 140 Fifth Image Generation Unit 142 Second Comparison Section 144 Second Integration Section 200 Storage device 300 display device 400 Control unit 500 Communications Department 700 3D CT images 700A cross section 700C cross section 700S cross section 710 First pseudo-X-ray image 720 Second pseudo-X-ray image 730 CT images 735 Mask Images 735A cross section 735C cross section 735S cross section 736 Lung area 740 Masked CT images 742 Masked CT images 743 Lung area 750 Third pseudo-X-ray image 760 Simple X-ray images 770 Fourth pseudo-X-ray image 772 First label image 774 Second label image 780 Fifth pseudo-X-ray image 790 Dual-energy X-ray imaging S100~S124 Steps in the learning method and image processing method
Claims
1. Input the 3D CT image of the subject, The three-dimensional CT image is projected into a two-dimensional space to generate a first pseudo-X-ray image. From the first pseudo-X-ray image, a neural network generates a second pseudo-X-ray image, which is a set of images in which multiple anatomical regions of the subject are shown separately, using a set of weight parameters that have been set in advance or updated. A labeled image is obtained in which the multiple anatomical regions are labeled to the three-dimensional CT image. Based on the three-dimensional CT image and the label image, a masked CT image is generated, which is a three-dimensional image in which the multiple anatomical regions are each depicted separately. The masked CT image is projected into a two-dimensional space to generate a third pseudo-X-ray image, which is a set of images in which the multiple anatomical regions are shown separately. The second pseudo-X-ray image and the third pseudo-X-ray image are compared, In the comparison described above, the absolute error or squared error between the pixel values of the second pseudo-X-ray image and the pixel values of the third pseudo-X-ray image is calculated for each pixel, and the average or weighted sum of the absolute error or squared error is taken as the result of the comparison. Based on the results of the comparison, the values of the multiple weight parameters used to generate the second pseudo-X-ray image are updated. A learning method that repeats the following until predetermined conditions are met: input of the 3D CT image, generation of the first pseudo-X-ray image, generation of the second pseudo-X-ray image, acquisition of the label image, generation of the masked CT image, generation of the third pseudo-X-ray image, comparison of the second pseudo-X-ray image and the third pseudo-X-ray image, and updating of the values of the multiple weight parameters, The learning method wherein the predetermined conditions are one of the following: the number of updates of the multiple weight parameters has reached a predetermined number; the result of the comparison has met the criteria; or processing for all images has been completed.
2. The learning method according to claim 1, wherein the average or weighted sum of the absolute error or the squared error, which is the result of the comparison, is reduced by updating the plurality of weight parameters.
3. In comparing the second pseudo-X-ray image and the third pseudo-X-ray image, As a result of the above comparison, the absolute error or the squared error is calculated, The third pseudo-X-ray image and the second pseudo-X-ray image are distinguished, The absolute error or the squared error and the result of the discrimination are integrated, In the aforementioned discrimination process, the classification error is calculated for each patch, the average or sum of the classification errors is calculated, and a scalar value representing the calculated average or sum is output. In the aforementioned integration, the weighted sum of the scalar values for each patch is calculated, The learning method according to claim 1, wherein, in updating the values of the plurality of weight parameters, the weighted sum which is the result of the integration is reduced by updating the plurality of weight parameters.
4. In comparing the second pseudo-X-ray image and the third pseudo-X-ray image, As a result of the above comparison, the absolute error or the squared error is calculated, The pair data of the first pseudo-X-ray image and the third pseudo-X-ray image is distinguished from the pair data of the first pseudo-X-ray image and the second pseudo-X-ray image. The absolute error or the squared error and the result of the discrimination are integrated, In the aforementioned discrimination process, the classification error is calculated for each patch, the average or sum of the classification errors is calculated, and a scalar value representing the calculated average or sum is output. In the aforementioned integration, the weighted sum of the scalar values for each patch is calculated, The learning method according to claim 1, wherein, in updating the values of the plurality of weight parameters, the weighted sum which is the result of the integration is reduced by updating the plurality of weight parameters.
5. The learning method according to any one of claims 1 to 4, wherein the acquisition of the label image is performed by generating the label image from the three-dimensional CT image.
6. The learning method according to any one of claims 1 to 4, wherein, in generating the masked CT image, a mask image is generated from the extracted CT image obtained by region extraction from the three-dimensional CT image, and the masked CT image is generated by multiplying the extracted CT image by the mask image.
7. In the input of the aforementioned 3D CT image, a simple X-ray image is further input. From the aforementioned simple X-ray image, a fourth pseudo-X-ray image is generated, which is a set of images in which the multiple anatomical regions are each shown separately. Determine whether the fourth pseudo-X-ray image is identical to the third pseudo-X-ray image which is the correct image. The learning method according to any one of claims 1 to 4, wherein the updating of the values of the plurality of weight parameters is performed by further updating the values of the plurality of weight parameters based on the result of the discrimination between the second pseudo-X-ray image and the fourth pseudo-X-ray image, thereby causing it to be determined that the fourth pseudo-X-ray image is identical to the third pseudo-X-ray image which is the correct image.
8. The fourth pseudo-X-ray image is converted into the first label image, A second label image is obtained for the aforementioned simple X-ray image. The first label image and the second label image are compared, The first scalar value, which is the result of the comparison between the second pseudo-X-ray image and the third pseudo-X-ray image, and the second scalar value, which is the result of the comparison between the first label image and the second label image, are integrated. In updating the values of the aforementioned multiple weight parameters, the multiple weight parameters are updated based on the results of the integrated comparison. In comparing the first label image and the second label image, the absolute error or squared error between the pixel values of the first label image and the pixel values of the second label image is calculated for each pixel, and the average or sum of the absolute error or squared error is calculated to output the second scalar value. In the aforementioned integration, a weighted sum of the first scalar value and the second scalar value is calculated. In updating the values of the multiple weight parameters, the sum of the weights is reduced by updating the multiple weight parameters based on the result of the integrated comparison. The learning method according to claim 7.
9. In the input of the three-dimensional CT image, a simple X-ray image and dual-energy X-ray images, which are multiple X-ray images taken with X-rays of different energies and highlighting different anatomical regions of the subject, are input. From the aforementioned simple X-ray image, a fourth pseudo-X-ray image is generated, which is a set of images in which the multiple anatomical regions are each shown separately. The fourth pseudo-X-ray image is reconstructed to generate a fifth pseudo-X-ray image, which is a set of images in which the same anatomical regions as those in the dual-energy X-ray image are emphasized. The dual-energy X-ray image and the fifth pseudo-X-ray image are compared, The third scalar value, which is the result of comparing the dual-energy X-ray image and the fifth pseudo-X-ray image, is integrated with the first scalar value, which is the result of comparing the second pseudo-X-ray image and the third pseudo-X-ray image. In updating the values of the aforementioned multiple weight parameters, the multiple weight parameters are updated based on the results of the integrated comparison. In comparing the dual-energy X-ray image with the fifth pseudo-X-ray image, the absolute error or squared error between the pixel values of the dual-energy X-ray image and the pixel values of the fifth pseudo-X-ray image is calculated for each pixel, and the average or sum of the absolute error or squared error is calculated to output the third scalar value. In the aforementioned integration, a weighted sum of the first scalar value and the third scalar value is calculated. In updating the values of the multiple weight parameters, the sum of the weights is reduced by updating the multiple weight parameters based on the result of the integrated comparison. The learning method according to any one of claims 1 to 4.
10. Obtain a simple X-ray image or a pseudo-X-ray image, The acquired simple X-ray image or the pseudo-X-ray image is transformed using parameters updated by the learning method described in any one of claims 1 to 4 to generate separated images, which are multiple images in which multiple anatomical regions of the subject are shown separately. Image processing methods.
11. The image processing method according to claim 10, which converts the separated images into segmentation labels.
12. A learning device equipped with a processor, The aforementioned processor, Input the 3D CT image of the subject, The three-dimensional CT image is projected into a two-dimensional space to generate a first pseudo-X-ray image. From the first pseudo-X-ray image, a neural network generates a second pseudo-X-ray image, which is a set of images in which multiple anatomical regions of the subject are shown separately, using a set of weight parameters that have been set in advance or updated. A labeled image is obtained in which the multiple anatomical regions are labeled to the three-dimensional CT image. Based on the three-dimensional CT image and the label image, a masked CT image is generated, which is a three-dimensional image in which the multiple anatomical regions are each depicted separately. The masked CT image is projected into a two-dimensional space to generate a third pseudo-X-ray image, which is a set of images in which the multiple anatomical regions are shown separately. The second pseudo-X-ray image and the third pseudo-X-ray image are compared, In the comparison described above, the absolute error or squared error between the pixel values of the second pseudo-X-ray image and the pixel values of the third pseudo-X-ray image is calculated for each pixel, and the average or weighted sum of the absolute error or squared error is taken as the result of the comparison. Based on the results of the comparison, the values of the multiple weight parameters used to generate the second pseudo-X-ray image are updated. A learning device that repeatedly performs the following until predetermined conditions are met: input of the 3D CT image, generation of the first pseudo-X-ray image, generation of the second pseudo-X-ray image, acquisition of the label image, generation of the masked CT image, generation of the third pseudo-X-ray image, comparison of the second pseudo-X-ray image and the third pseudo-X-ray image, and updating of the values of the plurality of weight parameters, The learning device is one of the following predetermined conditions: the number of parameter updates has reached a predetermined number; the comparison result has met the criteria; or processing for all images has been completed.
13. An image processing apparatus comprising a processor, The aforementioned processor, Obtain a simple X-ray image or a pseudo-X-ray image, The acquired simple X-ray image or the pseudo-X-ray image is transformed using parameters updated by the learning method described in any one of claims 1 to 4 to generate separated images, which are multiple images in which multiple anatomical regions of the subject are shown separately. Image processing device.
14. The image processing apparatus according to claim 13, wherein the processor converts the separated images into segmentation labels.
15. A learning program that causes a learning device equipped with a processor to execute a learning method, The aforementioned learning method is Input the 3D CT image of the subject, The three-dimensional CT image is projected into a two-dimensional space to generate a first pseudo-X-ray image. From the first pseudo-X-ray image, a neural network generates a second pseudo-X-ray image, which is a set of images in which multiple anatomical regions of the subject are shown separately, using a set of weight parameters that have been set in advance or updated. A labeled image is obtained in which the multiple anatomical regions are labeled to the three-dimensional CT image. Based on the three-dimensional CT image and the label image, a masked CT image is generated, which is a three-dimensional image in which the multiple anatomical regions are each depicted separately. The masked CT image is projected into a two-dimensional space to generate a third pseudo-X-ray image, which is a set of images in which the multiple anatomical regions are shown separately. The second pseudo-X-ray image and the third pseudo-X-ray image are compared, In the comparison described above, the absolute error or squared error between the pixel values of the second pseudo-X-ray image and the pixel values of the third pseudo-X-ray image is calculated for each pixel, and the average or weighted sum of the absolute error or squared error is taken as the result of the comparison. Based on the results of the comparison, the values of the multiple weight parameters used to generate the second pseudo-X-ray image are updated. A learning program that causes the processor to repeatedly input the three-dimensional CT image, generate the first pseudo-X-ray image, generate the second pseudo-X-ray image, acquire the label image, generate the masked CT image, generate the third pseudo-X-ray image, compare the second pseudo-X-ray image and the third pseudo-X-ray image, and update the values of the plurality of weight parameters until predetermined conditions are met, The learning program is characterized by the predetermined conditions being one of the following: the number of parameter updates has reached a predetermined number; the comparison result has met the criteria; or processing for all images has been completed.
16. An image processing program that causes an image processing device equipped with a processor to execute an image processing method, The aforementioned image processing method is: Obtain a simple X-ray image or a pseudo-X-ray image, The acquired simple X-ray image or the pseudo-X-ray image is transformed by the plurality of weight parameters updated by the learning method described in any one of claims 1 to 4 to generate separated images, which are a plurality of images in which a plurality of anatomical regions of the subject are shown separately. Image processing program.
17. The image processing program according to claim 16, further causing the processor to convert the separated images into segmentation labels.