Medical image processing method, medical image processing program, recording medium, medical image processing device, and learning method
The method addresses domain shift in medical image conversion by using a deformation vector field trained through machine learning to generate pseudo-images in different postures, enhancing annotation and machine learning applications.
Patent Information
- Application Number
- JP2025153554
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-12-24
- Filing Date
- 2025-09-16
- Publication Date
- 2025-11-14
AI Technical Summary
Existing medical image domain conversion techniques fail to adequately reduce domain shift due to varying image acquisition postures, leading to inaccuracies in annotation and machine learning applications.
A medical image processing method that generates pseudo-images in a different posture using a deformation vector field, trained through machine learning, to transform images while minimizing domain shift.
Reduces domain shift by generating high-quality pseudo-images that maintain image quality and accuracy, enabling effective annotation and machine learning using converted medical images.
Smart Images

Figure 2025170152000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a technique for performing domain transformation of medical images. [Background technology]
[0002] In the field of medical imaging (sometimes called medical imaging), it is common to generate pseudo-medical images taken in one domain (modality, etc.) from actual medical images taken in another domain, and then use the generated images for various purposes (for example, utilizing label data attached to the medical images, using them as images for machine learning, observing and diagnosing lesions, etc.).
[0003] For example, Patent Document 1 describes generating a virtual fluoroscopic image based on three-dimensional volume data reconstructed from a tomographic image acquired by a CT device (CT: Computed Tomography), and using the generated image as data for machine learning. Also, Non-Patent Document 1 describes converting a pseudo X-ray image generated from a CT image to resemble an actual X-ray image, and extracting (labeling) organs from the converted image. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 2019-198376 [Non-patent literature]
[0005] [Non-Patent Document 1] "Task-driven generative modeling for unsupervised domain adaptation: Application to x-ray image segmentation.", MICCAI 2018, Zhang, Yue et al., [Retrieved December 8, 2020], Internet (https: / / arxiv.org / abs / 1806.07201) Summary of the Invention [Problem to be solved by the invention]
[0006] Domain conversion of medical images allows annotation data from the original image to be used in other images, or the converted images to be used as data for machine learning. However, the number of medical images acquired varies greatly depending on the domain, and the subject's posture during imaging is often fixed. In such situations, simply converting medical images with different postures can result in domain shift. For example, in the aforementioned Patent Document 1, posture conversion is performed using a simple conversion table based on fat mass and muscle mass, which does not sufficiently reduce domain shift. Furthermore, in Non-Patent Document 1, the pseudo X-ray image generated from the CT image is in the supine position, while the converted X-ray image is in the upright position, resulting in domain shift due to posture.
[0007] As described above, it has been difficult to reduce the domain shift in the domain conversion of medical images using conventional techniques.
[0008] The present invention has been made in consideration of the above circumstances, and aims to provide a medical image processing method, a medical image processing program, a recording medium, a medical image processing device, and a learning method that can reduce domain shift. [Means for solving the problem]
[0009] To achieve the above-mentioned object, a medical image processing method according to a first aspect of the present invention is a medical image processing method executed by a medical image processing device including a processor, wherein the processor executes the following steps: a receiving step of receiving an input of a first medical image actually captured in a first posture; and an image generating step of generating, from the first medical image, a pseudo-second medical image captured in a second posture different from the first posture and showing the same part as the first medical image, wherein the second medical image is generated using a deformation vector field that transforms the first medical image into the second medical image. In the medical image processing method according to the first aspect, domain shift can be reduced by generating the second medical image from the first medical image using the deformation vector field, rather than directly transforming medical images.
[0010] In the first aspect and each of the following aspects, the "deformation vector field" is a collection of vectors that indicate the displacement (image deformation) from each voxel (or pixel; the same applies hereinafter) of the first medical image to each voxel of the second medical image. Note that the "same" part includes not only the case where the parts are completely the same, but also the case where at least a part of the part is common.
[0011] A medical image processing method according to a second aspect is the first aspect, wherein in the image generating step, a processor generates a generator that outputs a deformation vector field when a first medical image is input, and the deformation vector field is generated using a generator trained by machine learning. The second aspect defines one aspect of a generator construction method, and in machine learning, the generator can be constructed by training using the first medical image as training data. Machine learning includes deep learning. In addition, in the second aspect, a generator constructed by a training method according to any of the thirteenth to seventeenth aspects (described below) may be used.
[0012] A medical image processing method according to a third aspect is the first or second aspect, wherein in the image generating step, a processor applies a deformation vector field to the first medical image to generate a second medical image.
[0013] A medical image processing method according to a fourth aspect is any one of the first to third aspects, wherein in the image generation process, the processor converts the resolution of a first medical image to a lower resolution than the resolution before conversion, generates a deformation vector field from the first medical image converted to a lower resolution, converts the resolution of the generated deformation vector field to a higher resolution than the resolution before conversion, and applies the deformation vector field converted to a higher resolution to the first medical image to generate a high-resolution second medical image.
[0014] When attempting to generate a second medical image while keeping the first medical image at high resolution, it may be difficult to generate a deformation vector field. Even in such cases, in the fourth aspect, the difficulty can be avoided by generating a deformation vector field from the first medical image converted to low resolution and then converting this deformation vector field to high resolution, thereby generating a high-resolution second medical image. The degree of resolution conversion can be determined depending on the processing load and the final required resolution of the medical image.
[0015] A medical image processing method according to a fifth aspect is any one of the first to fourth aspects, wherein the processor accepts a CT image in a supine position as the first posture as the first medical image in the receiving step, and generates a CT image in an upright position as the second posture as the second medical image in the image generating step. The fifth aspect defines one aspect of the first and second postures and the first and second medical images.
[0016] A sixth aspect of the medical image processing method is any one of the first to fifth aspects, wherein the processor further performs a modality conversion step of converting the pseudo-generated second medical image into a third medical image using a modality different from that of the second medical image. The modality conversion can be performed, for example, from a CT image to an X-ray image, or vice versa.
[0017] A medical image processing method according to a seventh aspect is the sixth aspect, wherein in the modality conversion step, the processor generates an X-ray fluoroscopic image in the second posture as the third medical image.
[0018] The medical image processing method according to an eighth aspect is any one of the first to seventh aspects, wherein in the image generating step, the processor uses a deformation vector field to convert first label data corresponding to a first medical image into second label data corresponding to a second medical image. According to the eighth aspect, the deformation vector field can be used to convert the label data in addition to generating the second medical image. The label data is, for example, a segmentation label (a label attached to an organ).
[0019] A medical image processing method according to a ninth aspect is any one of the first to eighth aspects, wherein the processor receives, in the receiving step, a T1-weighted MR image and a T2-weighted MR image in which one of an upright position and a supine position is set as a first position as a first medical image, and generates, in the image generating step, a pseudo-T1-weighted MR image and a pseudo-T2-weighted MR image in which the other of the upright position and the supine position is set as a second position. The ninth aspect specifies conversion between MR images in different positions. The MR images are images acquired by an MR (Magnetic Resonance) device.
[0020] A medical image processing method according to a tenth aspect is any of the first to eighth aspects, in which the processor accepts, as a first medical image, a chest CT image in which one of an expiratory posture and an inhalation posture is set as a first posture in the receiving step, and generates a pseudo chest CT image in which the other of the expiratory posture and the inhalation posture is set as a second posture in the image generating step. The tenth aspect defines conversion between chest CT images in different postures, taking into account that the shape of the lungs changes between the expiratory posture and the inhalation posture.
[0021] To achieve the above-mentioned object, a medical image processing program according to an eleventh aspect of the present invention is a medical image processing program that causes a processor of a medical image processing apparatus to execute each step of a medical image processing method, the medical image processing program causing the processor to execute the following steps: a receiving step of receiving input of a first medical image actually captured in a first posture; and an image generating step of generating, from the first medical image, a pseudo-second medical image captured in a second posture different from the first posture of the same region as the first medical image, the second medical image being generated using a deformation vector field that transforms the first medical image into the second medical image. According to the eleventh aspect, it is possible to reduce domain shift as in the first aspect. Note that the medical image processing method executed by the medical image processing program according to the present invention may have the same configuration as the second to tenth aspects. A non-transitory recording medium on which computer-readable code of the above-mentioned medical image processing program is recorded can also be cited as an aspect of the present invention.
[0022] To achieve the above-mentioned object, a medical image processing apparatus according to a twelfth aspect of the present invention is a medical image processing apparatus including a processor, the processor performing the following steps: receiving an input of a first medical image actually captured in a first posture; and image generation processing for generating, from the first medical image, a pseudo-second medical image captured in a second posture different from the first posture of the same region as the first medical image, the second medical image being generated using a deformation vector field that transforms the first medical image into the second medical image. According to the twelfth aspect, it is possible to reduce domain shift, as in the first and twelfth aspects. The medical image processing apparatus according to the present invention may have the same configuration as the second to tenth aspects.
[0023] In order to achieve the above-mentioned object, a learning method for a medical image processing device according to a thirteenth aspect of the present invention includes a generator that receives an input of a first medical image that has been actually captured in a first posture, and artificially generates a second medical image from the first medical image, the second medical image being captured in a second posture different from the first posture, the generator generating the second medical image using a deformation vector field that converts the first medical image into the second medical image; and a classification unit that accepts input of a medical image of and classifies whether the input medical image is a second medical image or a fourth medical image, the method comprising: a generation unit learning step of updating generation parameters used by the generation unit to generate a deformation vector field so as to maximize the classification error of the classification unit while maintaining the parameters of the classification unit without updating them; and a classification unit learning step of updating the parameters of the classification unit so as to minimize the classification error of the classification unit while maintaining the generation parameters of the generation unit without updating them. The thirteenth aspect defines a learning method (machine learning method) for generating a deformation vector field, and the deformation vector field generated by this method can be used in each of the above-mentioned aspects of the present invention.
[0024] A learning method according to a fourteenth aspect is the same as that of the thirteenth aspect, but includes a step of inputting a first medical image to a generation unit and smoothing the generated deformation vector field by imposing constraints on generation parameters in the generation unit learning step. According to the fourteenth aspect, smoothing the deformation vector field makes it possible to generate a smooth second medical image. In the fourteenth aspect, the "constraint" may be, for example, a constraint on changes in the direction or magnitude of adjacent deformation vectors.
[0025] The learning method according to the 15th aspect is the 13th or 14th aspect, in which the generation unit learning step and the classification unit learning step are performed by inputting a first medical image to the generation unit and a fourth medical image to the classification unit. The 15th aspect specifies one aspect of medical images used for learning.
[0026] The learning method according to the 16th aspect is the same as that of the 15th aspect, in which the first medical image and the fourth medical image are medical images of the same part of different subjects. Generally, it is rare to acquire images of the same part of the same subject in different postures, and therefore the 16th aspect specifies a learning method for such cases.
[0027] A learning method according to a seventeenth aspect is any one of the thirteenth to sixteenth aspects, wherein the generating unit and the identifying unit are configured by neural networks. The seventeenth aspect defines one aspect of the generating unit and the identifying unit.
[0028] In addition, a program (learning program) that causes a medical image processing device to execute the learning methods of the 13th to 17th aspects, and a non-transitory recording medium on which computer-readable code of the program is recorded, can also be cited as an aspect of the present invention. [Effects of the Invention]
[0029] As described above, the medical image processing method, medical image processing program, recording medium, medical image processing apparatus, and learning method for a medical image processing apparatus according to the present invention can reduce domain shift. [Brief explanation of the drawings]
[0030] [Figure 1] FIG. 1 is a diagram showing a schematic configuration of a medical image processing apparatus according to the first embodiment. [Figure 2] FIG. 2 is a diagram showing a schematic diagram of a deformation vector field. [Figure 3] FIG. 3 is a diagram showing the state of the generation unit learning process. [Figure 4] FIG. 4 is a diagram showing a schematic view of three-dimensional CT images in the standing and lying positions. [Figure 5] FIG. 5 is a diagram showing the state of the recognition unit learning process. [Figure 6] FIG. 6 is a diagram showing how a medical image is transformed using a deformation vector field. [Figure 7]FIG. 7 is a diagram schematically showing label data corresponding to a three-dimensional CT image. [Figure 8] FIG. 8 is a diagram illustrating the transformation of label data using a deformation vector field. [Figure 9] FIG. 9 is a diagram showing a schematic configuration of a medical image processing apparatus according to the second embodiment. [Figure 10] FIG. 10 is a diagram schematically showing an X-ray fluoroscopic image. [Figure 11] FIG. 11 is a diagram showing the state of learning (execution of the learning method) in the second embodiment. [Figure 12] FIG. 12 is a diagram showing a schematic configuration of a medical image processing apparatus according to the third embodiment. [Figure 13] FIG. 13 is a diagram showing how a high-resolution CT image is converted into a low-resolution image. [Figure 14] FIG. 14 is a diagram showing the learning process using a low-resolution CT image. [Figure 15] FIG. 15 is a diagram showing how the resolution of a deformation vector field is converted. [Figure 16] FIG. 16 is a diagram showing how a medical image is transformed using a high-resolution transformed deformation vector field. [Figure 17] FIG. 17 is a diagram showing a modified example of the configuration of a medical image processing apparatus. DETAILED DESCRIPTION OF THE INVENTION
[0031] The following describes embodiments of a medical image processing method, a medical image processing program, a recording medium, a medical image processing device, and a learning method according to the present invention. In the description, reference is made to the accompanying drawings as necessary. For the sake of convenience, some components may be omitted from the accompanying drawings.
[0032] [First embodiment] [Configuration of medical image processing device] 1 is a diagram showing a schematic configuration of a medical image processing apparatus 10 (medical image processing apparatus) according to the first embodiment. The medical image processing apparatus 10 includes an image processing unit 100 (a generation unit, a classification unit, a deformation vector field, a learning control unit, a projection unit, and a processor), a storage device 200, a display device 300, an operation unit 400, and a communication unit 500. These components may be connected to each other by wire or wirelessly. Furthermore, these components may be housed in a single housing, or may be housed separately in multiple housings.
[0033] [Image processing unit configuration] As shown in FIG. 1, the image processing unit 100 (processor) includes a generation unit 110 (generation unit), a classifier 120 (classification unit), and a learning control unit 130. The generation unit 110 includes a generator 112 (generator), a deformation vector field 114 (deformation vector field), and a converter 116. The generator 112 is a network that receives an input of a medical image (first medical image) and generates a deformation vector field, and can be configured using a neural network such as U-Net used in pix2pix. A ResNet (Deep Residual Network)-based neural network can also be used (see Non-Patent Document 2). Any network structure used in the field of super-resolution can basically be applied to the present invention.
[0034] [Non-Patent Document 2] "Perceptual Losses for Real-Time Style Transfer and Super-Resolution", ECCV, 2016, Justin Johnson et al. [Retrieved December 8, 2020], Internet (https: / / arxiv.org / abs / 1603.08155) The deformation vector field 114 (deformation vector field) is a collection of three-dimensional vectors (vectors indicating the direction and amount of deformation) that transform each voxel of an input medical image (in the case of a three-dimensional image) into each voxel of an output medical image. FIG. 2 is a diagram schematically showing a deformation vector. Part (a) of FIG. 2 shows the entire deformation vector field (deformation vector field 114), and as shown in part (b) of the same figure, deformation vectors 114B exist corresponding to individual small regions 114A of the deformation vector field 114. In the case of a two-dimensional image, the deformation vector field is a collection of two-dimensional vectors (vectors indicating the direction and amount of deformation) that transform each pixel of an input medical image into each pixel of an output medical image.
[0035] The converter 116 applies the deformation vector field 114 to the input medical image (first medical image captured in a first posture) to pseudo-generate a second medical image (a medical image captured of the same region as the first medical image in a second posture different from the first posture) (image generation process, image generation step). Note that the "same" region does not necessarily mean that the region is completely the same between the first medical image and the second medical image, but also includes the case where at least a part of the region is common (the same applies to each of the following aspects).
[0036] The classifier 120 (classification unit) classifies whether the input medical image (second medical image, fourth medical image) is an actually captured medical image (fourth medical image) or a pseudo-generated medical image (second medical image) by the generation unit 110. Similar to the above-mentioned generator 112, the patch classifier used in pix2pix can be used as the classifier 120.
[0037] The learning control unit 130 updates the parameters of the generator 112 and the classifier 120 based on the classification result of the classifier 120 (generator training step, classifier training step). That is, the generator 112 and the classifier 120 are trained (configured) by machine learning. Details of the parameter update of the generator 112 and the classifier 120 (each step of the learning method) will be described later.
[0038] The functions of the image processing unit 100 described above can be realized using various processors and recording media. The various processors include, for example, a CPU (Central Processing Unit), which is a general-purpose processor that executes software (programs) to realize various functions, a GPU (Graphics Processing Unit), which is a processor specialized for image processing, and a programmable logic device (PLD), such as an FPGA (Field Programmable Gate Array), whose circuit configuration can be changed after manufacturing. Each function may be realized by a single processor, or by multiple processors of the same or different types (e.g., multiple FPGAs, a combination of a CPU and an FPGA, or a combination of a CPU and a GPU). Furthermore, multiple functions may be realized by a single processor. The hardware structure of these various processors is, more specifically, an electric circuit made up of a combination of circuit elements such as semiconductor elements.
[0039] When the processor or electrical circuit described above executes software (programs), computer-readable code for the software to be executed (for example, various processors and electrical circuits constituting the image processing unit 100, and / or a combination thereof) is stored in a non-transitory recording medium (memory) such as a flash memory or a ROM (Read Only Memory), and the computer references the software. The programs to be executed include programs (medical image processing programs, learning programs) that execute methods (medical image processing methods, learning methods) according to one aspect of the present invention. Furthermore, when the software is executed, information (medical images, etc.) stored in the storage device 200 is used as necessary. Furthermore, when the software is executed, for example, a RAM (Random Access Memory) is used as a temporary storage area.
[0040] In addition to the above components, the image processing unit 100 may also include a display control unit and an image acquisition unit (not shown).
[0041] [Information stored in storage device] The storage device 20 is composed of various types of magneto-optical recording media, semiconductor memories, and their control units, and stores actually captured medical images, pseudo-generated medical images (first to fourth medical images), software (programs) executed by the above-mentioned processor, etc.
[0042] [Configuration of display device, operation unit, and communication unit] The display device 300 is configured with a device such as a liquid crystal monitor and is capable of displaying data such as medical images. The operation unit 400 is configured with a mouse, keyboard, etc. (not shown), and a user can issue instructions necessary for executing the medical image processing method and learning method via the operation unit 400. The user can issue instructions via a screen displayed on the display device 300. The display device 300 may be configured with a touch panel monitor, and the user can issue instructions via the touch panel. The communication unit 500 can acquire medical images and other information from other systems connected via a network.
[0043] [Learning method for medical image processing equipment] Next, we will explain the learning method (machine learning method) of the medical image processing apparatus 10. Learning is performed by dividing it into a generation unit learning process that updates the parameters (generation parameters) of the generator 112 and a classification unit learning process that updates the parameters of the classifier 120.
[0044] [Generator (generation part) training] 3 is a diagram showing the generator learning process. The generator 112 (generator 110, processor) receives input of a first medical image (medical image for learning) actually captured in a first position (reception process, reception processing). Here, the "first medical image" is a CT image (supine position actual CT image 700) actually captured of a certain part of a subject (for example, the chest) in a supine position (an example of the "first position"). The "supine position" may be a supine position or a prone position. When the generator 112 receives a first medical image as input, it outputs a deformation vector field 114, and the converter 116 (generator 110, processor) applies the deformation vector field 114 to the supine real CT image 700 (first medical image) to generate a pseudo upright pseudo CT image 710 (second medical image) in which the same body part as the supine real CT image 700 is captured in an upright position (an example of a "second position different from the first position") (image generation step, image generation process). Note that randomness may be imparted to the generated medical image by adding a noise component (e.g., random noise) to the generator 112 (the same applies to each of the following embodiments).
[0045] FIG. 4 is a diagram showing three-dimensional CT images in the upright and supine positions. Part (a) of FIG. 4 shows a CT image 600 in the upright position, and part (b) of FIG. 4 shows a CT image 610 in the supine position. In part (a) of FIG. 4, cross sections 600S, 600C, and 600A are cross sections in the sagittal direction, coronal direction, and axial direction, respectively. Similarly, in part (b) of FIG. 4, cross sections 610S, 610C, and 610A are cross sections in the sagittal direction, coronal direction, and axial direction, respectively.
[0046] The classifier 120 (classifier, classification unit) receives input of a pseudo-generated upright pseudo CT image 710 and an upright actual CT image 720 (fourth medical image), which is a medical image actually captured in an upright position (second posture) of the same region as the supine actual CT image 700 (first medical image), and classifies whether the input medical image is the upright pseudo CT image 710 (second medical image) or the upright actual CT image 720 (fourth medical image). For example, the classifier 120 outputs the probability that the input medical image is the upright pseudo CT image 710 and / or the probability that the input medical image is the upright actual CT image 720 (one aspect of the classification result).
[0047] The supine actual CT image 700 (first medical image) and the upright actual CT image 720 (fourth medical image) may not be medical images of the same subject, but may be medical images of the same region of different subjects. This is because, in general, it is rare to actually acquire images of the same region of the same subject in different postures. According to this embodiment, even when the subjects in the supine actual CT image 700 and the upright actual CT image 720 are different, learning can be performed.
[0048] Based on the classification result, the learning control unit 130 (processor) updates the generation parameters (parameters used by the generator 112 to generate the deformation vector field 114) so as to maximize the classification error of the classifier 120, while maintaining the parameters of the classifier 120 without updating them (generator learning step). Note that the input to the classifier 120 is shown by a dotted line in Fig. 3, which indicates that the parameters of the classifier 120 are maintained without being updated.
[0049] [Classifier (classification part) training] 5 is a diagram showing the process of the classifier learning step. In the classifier learning step, similar to the generator learning step described above, a pseudo upright pseudo CT image 710 (second medical image) is generated (image generation step, image generation processing), and the classifier 120 classifies whether the input medical image is the upright pseudo CT image 710 or the upright actual CT image 720. Based on the classification result, the learning control unit 130 (processor) updates the parameters of the classifier 120 so as to minimize the classification error of the classifier 120, while maintaining the generation parameters of the generator 112 without updating them (classifier learning step). Note that the input to the generator 112 is indicated by a dotted line in FIG. 5, indicating that the parameters of the generator 112 are maintained without being updated.
[0050] The above-described learning may be terminated after the generator learning step and the classifier learning step have been performed a predetermined number of times, or may be terminated when the fluctuations in the parameters of the generator 112 and the classifier 120 have converged. Furthermore, the learning control unit 130 (processor) may alternately (sequentially) perform the generator learning step and the classifier learning step, or may repeat the steps in batch units or mini-batch units. Furthermore, the image processing unit 100 may display the learning process (e.g., the state of change in parameters) on the display device 300.
[0051] [Smoothing of deformation vector field] In addition, in the generator learning process, the image processing unit 100 (processor) may input a supine actual CT image (first medical image) to the generator 112 and smooth the generated deformation vector field by adding constraints to the generation parameters (smoothing process).
[0052] [Medical image transformation using deformation vector fields] 6 is a diagram showing how a medical image is transformed using a deformation vector field. Once the deformation vector field 114 is obtained by the above-described learning method, the transformer 116 (processor) applies the deformation vector field 114 to the supine real CT image 700A (first medical image) to generate an upright pseudo CT image 710A (second medical image). The image processing unit 100 may display the input supine real CT image 700A and / or the upright pseudo CT image 710A on the display device 300 (in response to a user's operation via the operation unit 400 or automatically) or may store the input in the storage device 200.
[0053] As described above, according to the first embodiment, instead of directly converting medical images, an upright pseudo CT image (second medical image) is generated from a supine real CT image 700 (first medical image) using the deformation vector field 114, thereby reducing domain shift.
[0054] [Transformation of label data using deformation vector fields] According to the first embodiment, the deformation vector field 114 generated by the above-described learning method can be used to convert label data in a certain posture into label data in a different posture. The label data is data in which segmentation labels are assigned to each organ (lungs 900A and heart 900B in the example of FIG. 7) corresponding to a medical image 900 (a 3D CT image in the example of FIG. 7) acquired in a certain posture (e.g., supine position), as schematically shown in FIG. 7 . When converting such label data, as shown in FIG. 8 , the image processing unit 100 (processor) applies the deformation vector field 114 to supine position CT label data 702 (first label data) corresponding to a supine position actual CT image 700 (first medical image) to generate upright position CT label data 704 (second label data) corresponding to an upright position pseudo CT image 710 (second medical image). The image processing unit 100 may display the input supine CT label data 702 and / or upright CT label data 704 for the medical image and motion on the display device 300 (in response to user operation via the operation unit 400 or automatically) or may store them in the storage device 200.
[0055] [Variations in medical images] In the first embodiment described above, a medical image in an upright position (a pseudo upright CT image) is generated from a medical image in a supine position (a real CT image in a supine position 700). However, the relationship between the postures of the input medical image and the generated medical image is not limited to this. In addition to the above, a supine image may be generated from a standing image. The images used may also be X-ray images or MR images. For example, the image processing unit 100 (processor) may accept, as the first medical images, T1-weighted MR images and T2-weighted MR images in which one of the upright and supine positions is set as a first posture in the reception step (reception processing), and may generate pseudo T1-weighted MR images and T2-weighted MR images in which the other of the upright and supine positions is set as a second posture in the image generation step. Furthermore, the image processing unit 100 (processor) may accept a chest CT image in which one of the expiratory posture and the inhalation posture is set as a first posture as a first medical image in the reception step (reception processing), and may generate a pseudo chest CT image in which the other of the expiratory posture and the inhalation posture is set as a second posture in the image generation step. Furthermore, the region captured in the medical image is not particularly limited. Furthermore, the medical image is not limited to a three-dimensional image as in the above-described embodiment, but may be a two-dimensional image such as an X-ray fluoroscopic image. When a two-dimensional image is used, a two-dimensional deformation vector field is used correspondingly.
[0056] [Second embodiment] Next, a second embodiment of the present invention will be described. In the second embodiment, a modality conversion step (modality conversion processing) is further executed to convert a second medical image pseudo-generated by the above-described method into a third medical image generated by a modality different from that of the second medical image. FIG. 9 is a diagram showing a schematic configuration of a medical image processing apparatus 11 according to the second embodiment, which includes a projection unit 140. In the medical image processing method, medical image processing apparatus, and medical image processing program according to the present invention, a medical image acquired by a certain modality may be pseudo-converted into a medical image generated by a different modality (execution of the modality conversion step). In the second embodiment, the projection unit 140 pseudo-converts a CT image (second medical image) into an X-ray fluoroscopic image (third medical image). In other words, generating an X-ray fluoroscopic image (two-dimensional image) by projecting a CT image (three-dimensional volume data) is one aspect of modality conversion, and the X-ray fluoroscopic image is one aspect of the third medical image. FIG. 10 is a diagram schematically showing an X-ray fluoroscopic image 750. As shown in FIG.
[0057] The components of the medical image processing apparatus 11 other than the projection unit 140 are the same as those of the medical image processing apparatus 10 in the first embodiment, and therefore detailed description thereof will be omitted.
[0058] 11 is a diagram showing the learning process (execution of the learning method) in the second embodiment. In the second embodiment, the projection unit 140 converts an upright pseudo CT image 710 (second medical image) in a pseudo manner to generate an upright pseudo X-ray image 730 (third medical image) (modality conversion step, modality conversion process). The classifier 120 (classifier, classification unit) receives input of the pseudo-generated upright pseudo X-ray image 730 and an upright actual X-ray image 740 (fourth medical image), which is a medical image actually captured in an upright position (second posture) of the same region as the supine actual CT image 700 (first medical image), and classifies whether the input medical image is the upright pseudo X-ray image 730 (second medical image) or the upright actual X-ray image 740 (fourth medical image). For example, the classifier 120 outputs the probability that the input medical image is a pseudo upright X-ray image 730 and / or a real upright X-ray image 740 (one aspect of the classification result).
[0059] The learning control unit 130 (processor) updates the parameters of the generator 112 and the parameters of the classifier 120 based on the classification result (generation unit learning step, classification unit learning step). As described above in the first embodiment, the learning control unit 130 (processor) performs this by maintaining one of the parameters of the generator 112 and the parameters of the classifier 120 without updating the other. Note that for convenience, in FIG. 11 , the fact that one parameter is maintained without updating is not distinguished, and the generation unit learning step and the classification unit learning step are shown together in one diagram, but in reality, parameter updating is performed as separate steps, as in the first embodiment (see FIGS. 4 and 5).
[0060] As described above, in the second embodiment, the domain shift can be reduced by using a deformation vector field in the same way as in the first embodiment. Note that in the second embodiment, the conversion of medical images during actual operation (inference), the conversion of label data, variations of medical images, etc. can be performed in the same way as described above for the first embodiment. Also in the second embodiment, modality conversion can be performed in the same way as in the first embodiment.
[0061] [Third embodiment] In the above-described embodiment, a generator and a classifier are used, but it is known that learning may not go well if the resolution of the image input to the generator is high (see, for example, Non-Patent Documents 3 and 4 below).
[0062] [Non-Patent Document 3] "Conditional Image Synthesis with Auxiliary Classifier GANs", Augustus Odena et al., [Retrieved December 8, 2020], Internet (https: / / arxiv.org / abs / 1610.09585) [Non-Patent Document 4] "Progressive Growing of GANs for Improved Quality, Stability, and Variation", ICLR 2018, Tero Karras et al., [Retrieved December 8, 2020], Internet (https: / / arxiv.org / abs / 1710.10196) From this perspective, in the third embodiment, learning is performed after converting the resolution of the pre-conversion medical image. Fig. 12 is a diagram showing a schematic configuration of a medical image processing device 12 (medical image processing device) according to the third embodiment, and the medical image processing device 12 includes a resolution conversion unit 150 (processor). Note that the components of the medical image processing device 12 other than the projection unit 140 are the same as those of the medical image processing device 10 in the first embodiment, and therefore detailed description thereof will be omitted.
[0063] When performing learning and medical image processing using the medical image processing device 12, for example, as shown in FIG. 13, a high-resolution supine real CT image 760 (first medical image) is converted into a low-resolution supine real CT image 770 (first medical image) by the resolution conversion unit 150 (image generation process, resolution conversion process). That is, the resolution of the low-resolution supine real CT image 770 is lower than the resolution of the high-resolution supine real CT image 760 before conversion. The resolution conversion unit 150 can convert the resolution by thinning or averaging voxels or pixels. Then, as shown in FIG. 14, this low-resolution supine real CT image is input to the generator 112, which generates a low-resolution deformation vector field 117A (deformation vector field) by machine learning. The learning procedure is similar to that described above for the first and second embodiments, and includes a generator learning process for updating the parameters of the generator 112 and a classifier learning process for updating the parameters of the classifier 120. The classifier 120 receives as input a low-resolution upright pseudo CT image 780 (second medical image) pseudo-generated by the generator 110 and a low-resolution upright actual CT image 790 (fourth medical image).
[0064] 15, the resolution conversion unit 150 generates a high-resolution deformation vector field 117B (deformation vector field) by interpolation, enlargement, etc. from the low-resolution deformation vector field 117A generated by such learning (image generation process, resolution conversion process). That is, the resolution of the high-resolution deformation vector field 117B is higher than that of the low-resolution deformation vector field 117A, which is the deformation vector field before conversion.
[0065] 16, the converter 116 (processor) applies the high-resolution deformation vector field 117B thus obtained to the high-resolution supine position real CT image 760 to generate a high-resolution upright position pseudo CT image 800 (second medical image) (image generation process). According to the third embodiment, by performing such resolution conversion, it is possible to generate a high-precision image while avoiding problems that may occur during learning.
[0066] The resolution may be converted according to the progress of learning. For example, in the early stages of learning, high-resolution medical images may be converted to low-resolution and used for learning, and as learning progresses, the resolution of the medical images used for learning may be increased. When changing the resolution of the medical images in this way, it is preferable to change the resolution of the deformation vector field and the network configuration of the generator 112 and the classifier 120 (e.g., the number and size of convolutional layers) accordingly (see Non-Patent Document 4 above).
[0067] Furthermore, the resolution of the original medical image at which resolution conversion is performed and the degree to which the resolution is to be lowered can be determined taking into consideration the purpose of use of the medical image, the processing load, etc. The image processing unit 100 (processor) may accept the user's settings of the resolution conversion conditions via the operation unit 400, etc., and perform the above-mentioned processing based on the accepted conditions. Also in the third embodiment, conversion of medical images during actual operation (inference), conversion of label data, variations of medical images, etc. can be performed in the same manner as described above for the first embodiment. Also in the third embodiment, conversion of modality may be performed in the same manner as in the first embodiment.
[0068] <Modification of the configuration of the medical image processing device> In the first to third embodiments described above, the posture is converted in one direction (for example, a medical image in a supine position is converted into a medical image in an upright position). However, the present invention also allows for a posture to be converted in two directions (for example, a posture from a supine position to a standing position and a posture from a standing position to a supine position). FIG. 17 shows a modified configuration of a medical image processing device. The medical image processing device 13 (medical image processing device) generates a pseudo-standing pseudo CT image 710A (second medical image) from a supine actual CT image 700A (first medical image) using a generator 110A (generator 112A, deformation vector field 115A, converter 116A; processor). The pseudo-standing CT image 710A and the upright actual CT image 710B are then input to a classifier 120A, which calculates a classification error. The learning control unit 130A (processor) updates the parameters of the generator 112A and the classifier 120A based on the calculated classification error (generator learning step, classifier learning step), in the same manner as described above in the first embodiment.
[0069] Similarly, a pseudo-supine position pseudo CT image 700B (second medical image) is generated from an upright position real CT image 710B (first medical image) by the generation unit 110B (generator 112B, deformation vector field 115B, converter 116B; processor). Then, the pseudo-supine position CT image 700B and the supine position real CT image 700A (fourth medical image) are input to the classifier 120B to calculate a classification error. The learning control unit 130B (processor) updates the parameters of the generator 112B and the classifier 120B based on the calculated classification error, as described above for the first embodiment (generator learning step, classifier learning step).
[0070] When training the generator 112B and the classifier 120B, the upright pseudo CT image 710A may be input instead of the upright real CT image 710B. Similarly, when training the generator 112A and the classifier 120A, the upright pseudo CT image 700B may be input instead of the supine real CT image 700A. The supine real CT image 700A and the supine pseudo CT image 700B constitute a supine CT image domain 701A, and the upright pseudo CT image 710A and the upright real CT image 710B constitute an upright CT image domain 701B. As described above, noise components may be input to the generators 112A and 112B.
[0071] In the above-described modified example, similarly to the first to third embodiments, conversion of medical images during actual operation (during inference), conversion of label data, variation of medical images, and conversion of modality may be performed.
[0072] Although the embodiments of the present invention have been described above, the present invention is not limited to the above-described aspects, and various modifications are possible without departing from the spirit of the present invention. [Explanation of symbols]
[0073] 10 Medical image processing device 11 Medical image processing equipment 12 Medical image processing equipment 13 Medical image processing equipment 20 Storage device 100 Image processing unit 110 Generation part 110A generation section 110B Generation section 112 Generator 112A generator 112B Generator 114 Deformed Vector Fields 114A Small area 114B Deformation Vector 115A Deformed Vector Field 115B Deformed vector field 116 Converter 116A Converter 116B Converter 117A Low-resolution deformation vector field 117B High-resolution deformation vector field 120 Classifier 120A discriminator 120B Discriminator 130 Learning control unit 130A Learning control unit 130B Learning control unit 140 Projection section 150 Resolution conversion unit 200 Storage device 300 display device 400 Control unit 500 Communications Department 600 CT images 600A cross section 600C cross section 600S cross section 610 CT images 610A cross section 610C cross section 610S cross section 700 supine CT images 700A supine CT image 700B Supine position pseudo CT image 701A Supine CT Image Domain 701B Upright CT Image Domain 702 supine CT label data 704 Standing CT Label Data 710 Upright Pseudo CT Images 710A Upright Pseudo CT Image 710B Upright CT image 720 upright CT images 730 Standing Pseudo X-ray Image 740 Standing X-ray Images 750 X-ray images 760 high-resolution supine CT images 770 low-resolution supine CT images 780 Low-resolution upright pseudo-CT images 790 Low-resolution upright CT images 800 high-resolution upright pseudo-CT images 900 Medical Images 900A Lung 900B Heart
Claims
1. A medical image processing method executed by a medical image processing apparatus including a processor, the processor comprising: a receiving step of receiving an input of a first medical image actually captured in a first posture; an image generating step of generating a pseudo-second medical image from the first medical image, the second medical image being obtained by photographing the same part as the first medical image in a second posture different from the first posture, the image generating step generating the second medical image using a deformation vector field that transforms the first medical image into the second medical image; Run A medical image processing method in which the processor further performs a modality conversion step of pseudo-converting the pseudo-generated second medical image into a third medical image using a modality different from that of the second medical image.
2. The medical image processing method according to claim 1 , wherein in the image generating step, the processor generates the second medical image by applying the deformation vector field to the first medical image.
3. In the image generating step, the processor converting the resolution of the first medical image to a lower resolution than the resolution before conversion; generating the deformation vector field from the first medical image converted to the low resolution; converting the resolution of the generated deformation vector field to a higher resolution than the resolution before conversion; 3. The medical image processing method according to claim 1, further comprising applying the high-resolution transformed deformation vector field to the first medical image to generate the high-resolution second medical image.
4. The processor: In the receiving step, a CT image in a supine position as the first posture is received as the first medical image; 4. The medical image processing method according to claim 1, wherein the image generating step generates a CT image in which a standing position is used as the second posture as the second medical image.
5. The medical image processing method according to claim 1 , wherein in the modality conversion step, the processor generates an X-ray fluoroscopic image in the second posture as the third medical image.
6. 6. The medical image processing method according to claim 1, wherein in the image generation step, the processor converts first label data corresponding to the first medical image into second label data corresponding to the second medical image using the deformation vector field.
7. The processor: In the receiving step, a T1-weighted MR image and a T2-weighted MR image in which one of an upright posture and a supine posture is set as the first posture are received as the first medical images; 7. The medical image processing method according to claim 1, wherein in the image generation step, a T1-weighted MR image and a T2-weighted MR image are generated in a pseudo manner, with the other of the standing posture and the lying posture being the second posture.
8. The processor: In the receiving step, a chest CT image in which one of an exhalation posture and an inhalation posture is set as the first posture is received as the first medical image; 8. The medical image processing method according to claim 1, wherein the image generating step generates a pseudo chest CT image in which the other of the expiratory posture and the inhalation posture is the second posture.
9. A medical image processing program that causes a processor of a medical image processing device to execute each step of a medical image processing method, the processor, a receiving step of receiving an input of a first medical image actually captured in a first posture; an image generating step of generating a pseudo-second medical image from the first medical image, the second medical image being obtained by photographing the same part as the first medical image in a second posture different from the first posture, the image generating step generating the second medical image using a deformation vector field that transforms the first medical image into the second medical image; a modality conversion step of converting the pseudo-generated second medical image into a pseudo-third medical image using a modality different from that of the second medical image; A medical image processing program that executes the above.
10. A non-transitory computer-readable recording medium having the program according to claim 9 recorded thereon.
11. A medical image processing device including a processor, the processor: a receiving process for receiving an input of a first medical image actually captured in a first posture; an image generation process for generating a pseudo-second medical image from the first medical image, the second medical image being obtained by capturing the same part as the first medical image in a second posture different from the first posture, the image generation process generating the second medical image using a deformation vector field that transforms the first medical image into the second medical image; a modality conversion process for converting the pseudo-generated second medical image into a pseudo-third medical image using a modality different from that of the second medical image; A medical image processing device that performs the above.
12. a generating unit that receives an input of a first medical image that has been actually captured in a first posture, and generates a pseudo-second medical image from the first medical image, the second medical image being captured in a second posture different from the first posture, the generating unit generating the second medical image using a deformation vector field that converts the first medical image into the second medical image; an identification unit that receives input of the second medical image and a fourth medical image that is an actual image of the same region as the first medical image in the second posture, and identifies whether the input medical image is the second medical image or the fourth medical image; A learning method for a medical image processing apparatus comprising: The medical image processing device, a modality conversion step of converting the pseudo-generated second medical image into a pseudo-third medical image using a modality different from that of the second medical image; a generator learning step of updating generation parameters used by the generator to generate the deformation vector field so as to maximize the classification error of the classifier while maintaining the parameters of the classifier without updating them; a classifier learning step of updating the parameters of the classifier so as to minimize the classification error of the classifier, while maintaining the generation parameters of the generation step without updating them; Learning how to do it.
13. The learning method according to claim 12, wherein the generation unit learning step includes a step of inputting the first medical image to the generation unit and smoothing the generated deformation vector field by adding constraints to the generation parameters.
14. The learning method according to claim 12 or 13, wherein the generation unit learning step and the identification unit learning step are performed by inputting the first medical image to the generation unit and inputting the fourth medical image to the identification unit.
15. The learning method according to claim 14 , wherein the first medical image and the fourth medical image are medical images of the same region of different subjects.
16. The learning method according to any one of claims 12 to 15, wherein the generating unit and the identifying unit are configured by neural networks.
Citation Information
Patent Citations
Image processing apparatus and processing method and program therefor
JP2012217770A
Image processing device, control method and program thereof
JP2015073799A
Medical imaging apparatus and method of operating same
US20170011509A1
Medical image processor, medical image processing method, and medical image processing system
JP2019198376A