Medical image processing method and medical image processing device

The medical image processing method addresses the challenge of motion artifacts in CT imaging by using feature maps and four-dimensional motion field estimation for motion-compensated reconstruction, resulting in improved image quality and reduced costs.

JP2025081284APending Publication Date: 2025-05-27CANON MEDICAL SYST CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024199870
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-15
Filing Date
2024-11-15
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Current CT image reconstruction methods struggle with motion artifacts, particularly due to respiratory and cardiac motion, which degrade image quality and pose challenges in lung diagnosis and quantitative chest CT.

Method used

A medical image processing method that involves acquiring a dataset from a 3D region of a subject, generating feature maps to estimate motion, and using these feature maps to calculate a four-dimensional motion field, which is then used for motion-compensated reconstruction to generate images without motion artifacts.

Benefits of technology

This method effectively reduces motion artifacts, improves CT image quality, and reduces computational time and hardware costs, leading to more accurate lung and heart imaging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025081284000001_ABST
    Figure 2025081284000001_ABST
Patent Text Reader

Abstract

To achieve reduction of motion artifacts at low cost.SOLUTION: A medical image processing method acquires a data set obtained by the scan of a three-dimensional region of a subject; generates a pair of feature maps indicating features of an image reconstructed from a corresponding part, which are feature maps for estimating the motion at a time point on the basis of the aforesaid part corresponding to the time point of the data set for each time point of a plurality of time points in the scan; estimates a four-dimensional motion field indicating a change with time of the motion of the subject in the three-dimensional region on the basis of a plurality of pairs of the feature maps corresponding to the plurality of time points; and generates a reconstruction image of the subject on the basis of the four-dimensional motion field and the data set.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed in this specification and the drawings relate to a medical image processing method and a medical image processing apparatus.

Background Art

[0002] The description of the background art is for schematically showing the background of the present disclosure. It is not explicitly or implicitly recognized that the current achievements (work) of the inventors in the scope described in this background art column are regarded as the prior art for the present disclosure, and similarly, the achievements of the inventors in aspects not recognized as the prior art at the time of filing are regarded as the prior art for the present disclosure.

[0003] Computer tomography (CT) systems and methods are widely used particularly in medical imaging and diagnosis. A CT system generally creates images of one or more cross-sectional slices across the entire body of a subject. An irradiation source such as an X-ray source irradiates radiation from one side of the body. At least one detector on the opposite side of the body receives the radiation that has propagated through the body. By processing the electrical signals received from the detector, the attenuation of the radiation that has passed through the body is measured.

[0004] X-ray CT has been recognized for extensive clinical applications in imaging of cancer, the heart, and the brain. CT has been increasingly used for various applications such as discrimination tests for cancer and pediatric imaging, and accordingly, it has been required to reasonably reduce the radiation dose of clinical CT scans as much as possible. In the case of low-dose CT, the image quality may be degraded by many factors such as high quantum noise, problems with scanning geometric shapes (i.e., large cone angles, high helical pitches, truncation errors, etc.), and other non-ideal physical phenomena (i.e., scattering, beam hardening, crosstalk, metal, etc.).

[0005] Removing motion artifacts is one of the most difficult problems in CT imaging. Artifacts arise from two main types of motion: respiratory motion and cardiac motion. In uncooperative patients or emergency cases, the image quality of CT scans may be impaired due to respiratory motion or skull motion. Motion artifacts in lung structures are inevitable in the normal practice of chest CT scans. The motion of the heartbeat or the involuntary motion of the diaphragm is the main cause of lung motion that occurs even when holding one's breath during a CT scan.

[0006] As a result, motion artifacts of various lung structures such as lung parenchyma, pulmonary vessels, or airways are often seen in normal CT images. These artifacts resemble various lung diseases, including bronchiectasis caused by edge duplication, cysts, emphysema, or ground glass opacity (GGO) nodules due to CT value bias, which poses difficulties in lung diagnosis using CT. In addition, motion artifacts are one of the main challenges in quantitative chest CT. Therefore, motion estimation in cardiac CT images is important for the evaluation of the anatomical structure and function of the human heart.

[0007] Conventional CT image reconstruction methods generally assume that the subject is stationary during data acquisition. Artifacts have a significant impact on the diagnosis using these reconstructed images, especially when the features being imaged are small. Correction methods are needed to detect motion artifacts and correct them during image reconstruction. For example, plaques formed in the coronary arteries generally indicate the risk of potential heart attacks, but they are difficult to image due to their small size.

[0008] Developing an efficient correction method can be difficult due to the accurate modeling of the forward model and the difficulty of analyzing complex inverse problems. For motion correction of coronary vessels, a local motion model is required, but for organs such as the lungs or skull, an overall motion model is sufficient.

[0009] Over the past few decades, many innovative techniques for improving CT image quality, such as model-based iterative image reconstruction for correcting respiratory motion or lung motion, have been developed. However, conventional imaging techniques use brute-force approaches to reduce the impact of motion artifacts in CT imaging. Some of these techniques include using two X-ray tube or detector pairs at different angles to each other, rotating the gantry faster using a heavier or higher-output power tube, or combining data from consecutive cardiac cycles. However, these techniques and methods are often time-consuming, computationally challenging, and require expensive hardware. In particular, in some difficult scenarios, the image quality is still inferior.

[0010] Therefore, there is a need to develop an efficient method for motion estimation and generation of CT images without motion artifacts using motion compensation reconstruction techniques. Furthermore, an improved method is desired to reduce computational time, hardware costs, and further improve CT image quality. The disclosed embodiments satisfy such needs.

Prior Art Documents

Non-Patent Documents

[0011]

Non-Patent Document 1

Non-Patent Document 2

[0012] One of the problems to be solved by the embodiments disclosed in this specification and the drawings is to reduce motion artifacts at low cost. However, the problems to be solved by the embodiments disclosed in this specification and the drawings are not limited to the above problems. The problems corresponding to the respective effects of each configuration shown in the embodiments described later can also be regarded as other problems.

Means for Solving the Problems

[0013] The medical image processing method according to the embodiment includes acquiring a data set obtained by scanning a three-dimensional region of a subject, and for each of a plurality of time points during the scan, based on a part of the data set corresponding to the time point, generating a pair of feature maps that represent the features of the image reconstructed from the part and are for estimating the motion at the time point, estimating a four-dimensional motion field that shows the change in the motion of the subject in the three-dimensional region over time based on the plurality of pairs of feature maps corresponding to the plurality of time points, and generating a reconstructed image of the subject based on the four-dimensional motion field and the data set.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2A

Figure 2B

Figure 2C

Figure 3

Figure 4

Figure 5

Figure 6A

Figure 6B

Figure 6C

Figure 6D

Figure 7

DETAILED DESCRIPTION OF THE INVENTION

[0015] Hereinafter, embodiments of a medical image processing method and a medical image processing apparatus will be described with reference to the accompanying drawings. The embodiments generally relate to a method or apparatus for estimating motion in a Computed Tomography (CT) scan image using a feature map and compensating for the estimated motion during reconstruction of the CT scan image. For example, the medical image processing method and medical image processing apparatus according to the embodiments estimate motion from the feature map of the CT scan image using image registration and a Deep Learning (DL) network.

[0016] To address the above-identified problems of known reconstruction methods for medical images, the methods described herein estimate motion and perform reconstruction to generate an image without motion artifacts, and are further developed to improve the image quality of medical images such as Computed Tomography (CT) images of the lungs and heart. In the embodiments, an image without motion artifacts means an image in which the influence of motion artifacts has been removed or the influence of motion artifacts has been suppressed. Further, each example presented herein applying these methods is non-limiting. These methods described herein can also be applied to other medical imaging modalities, such as MRI, PET / SPECT, etc., by adapting the framework described herein.

[0017] Therefore, the discussion herein merely discloses and describes exemplary implementations of the present disclosure. As will be understood by those skilled in the art, the present disclosure may be embodied in other specific forms without departing from its spirit or essential characteristics. Accordingly, the present disclosure is intended to be illustrative and not intended to limit the scope of the present invention and other claims. The present disclosure, including easily understandable variations of the teachings herein, partially defines the scope of the terms of the appended claims so as not to publicly present an inventive subject matter.

[0018] As described above, CT image quality can be degraded by many factors such as movement due to breathing and heart movement, as well as other non-ideal physical phenomena. As a result, developing an efficient correction method is often difficult due to several computational challenges that are time-consuming and require expensive hardware with existing techniques. In particular, in some difficult scenarios, the resulting image quality is still inferior.

[0019] To address the above-identified issues of known methods, the methods described herein use feature map-based motion estimation and motion compensation reconstruction. Feature maps can represent the edges and high-frequency information of the scanned subject instead of conventional CT images. For example, in certain implementations, filters are utilized to obtain the feature maps of the scanned images in the methods described herein. The feature maps are further utilized for the generation of a motion field, which is ultimately used for motion compensation during reconstruction and the generation of an image without motion artifacts. In particular, various implementations of the methods described herein offer several advantages over conventional methods of image reconstruction.

[0020] In certain implementations of the methods described herein, the generation of the motion field is realized via image registration. First, image registration is performed using the feature maps, and then the motion field is generated using the results of the image registration. Image registration, also known as image fusion or image matching, is a process of aligning two or more images based on the appearance of the images. Medical image registration aims to find the optimal spatial transformation that best aligns the underlying anatomical structures within the images. Medical image registration is used in many clinical applications such as image guidance, motion tracking, segmentation, dose accumulation, and image reconstruction.

[0021] In certain implementations of the methods described herein, the generation of the motion field is realized via a deep learning (DL) network. In certain implementations, an offline training of the DL network is performed and then the trained network is incorporated into the reconstruction step. Generally, DL networks have been previously adapted for image processing to improve the spatial resolution of images and reduce noise. Compared to conventional methods, deep learning does not require accurate modeling and instead relies on learning from a training dataset. Thus, the methods described herein can achieve better image quality than conventional methods from the perspective of images without motion artifacts.

[0022] For example, the methods disclosed herein can affect the improvement of various research areas where a DL-based convolutional neural network (CNN) may be used to generate reconstructed CT images without motion artifacts. By using training data corresponding to different CT scan methods and scan conditions, various CNN networks can be trained that will be trained on projection data corresponding to specific CT scan methods, procedures, uses, and conditions by using the training data selected to match the specific CT scan methods, procedures, uses, and conditions. In this way, each CNN network can be customized and adjusted according to specific conditions and methods related to CT scans.

[0023] Furthermore, the customization of the CNN network can be extended to the motion field of the projection data and can also be extended to the anatomical structure or region of the body to be imaged. The methods described herein can be applied to the motion estimation and compensation of the reconstructed CT images. Further, by using the redundancy of information in adjacent slices of the 3D CT image and using the kernels of the convolutional layers of the DL network that are extended to the pixels of the slices above and below the slice to be denoised, volumetric-based DL can be implemented. Generally, DL can be adapted to image processing for motion estimation and generation of motionless images.

[0024] Accordingly, feature maps, image registration, and convolutional neural networks are used in the embodiments herein to obtain the motion field of the scan image. The motion field is further utilized for motion compensation of the scan image, resulting in a high-quality motionless scanned CT image.

[0025] Referring now to the drawings, throughout several figures, the same reference numerals denote the same or corresponding parts, and FIG. 1 shows a flowchart of method 100. In the embodiments disclosed herein, method 100 is a method for motion estimation using scan data and motion-compensated image reconstruction.

[0026] In step 120, unprocessed data 115, which is unprocessed projection data, is obtained from the projection view. The set of unprocessed data 115 collected through the scan is also simply referred to as the data set. In the embodiments, the unprocessed data means the data before the reconstruction process and is not limited to the meaning of data that has not received any processing. Also, in this embodiment, the projection data collected by a CT device is described as an example of the unprocessed data 115, but it is also possible to obtain data collected by another type of scan such as MRI, PET, or SPECT as the unprocessed data 115.

[0027] Note that the data obtained in step 120 may be image data obtained by a computer tomography (CT) device through a back-projection process. Back-projection is a process of mapping unprocessed projection data to an image. Generally, several cycles of projection data are combined to reconstruct the final image. In the embodiments disclosed herein, partial unprocessed projection data is obtained in step 120 for further processing.

[0028] In step 140, feature map reconstruction is performed using the unprocessed data 115, and a feature map 135 of the scanned subject is generated. The feature map can represent, for example, the edges and high-frequency information of the scanned subject. The feature map generally consists of a set of connected regions each having basic features such as edges, frequency components related to some special image patterns, image intensity / color, prior shape, and texture, and includes information in the image that can usually be recognized. The purpose of feature map reconstruction is to divide the image into specific regions that are adjacent to each other but do not overlap with each other according to different features of the image, divide these specific regions into different classes, group regions having the same or similar features into the same class, and make the basic features different between classes. In the embodiments disclosed herein, feature map generation is used to realize an understanding of any movement of the subject occurring during scanning.

[0029] In the embodiments disclosed herein, feature map reconstruction utilizes a feature enhancement filter and known parameters of the scanned subject (e.g., information regarding the parameters of the heart if the scanned subject is a patient's heart). For example, feature enhancement filters such as high-pass filters, low-pass filters, and band-pass filters are used. In examples of the embodiments disclosed herein, various segmentation methods, including conventional segmentation and DL-based segmentation methods, can be used for feature map reconstruction. Histogram-based image value conversion can also be utilized to enhance features and either expand the differences in values between organs and tissues within the image or reduce the differences between the same special tissues and organs. Generally, feature enhancement filters are used not only to enhance specific information, e.g., features within an image, but also to reduce or remove any unnecessary information such as noise based on the known parameters of the scanned subject. The enhanced information is represented as a feature map of the scanned subject.

[0030] In step 160, the feature map 135 is used and motion estimation is performed to obtain a motion field 155 of the scanned subject. The motion field provides a detailed analysis of all detected motion within the scanned image. The motion field is, for example, a set of the positions of all features of the scanned image at all time points of the scan. Using this information, information about any changes in position within a three-dimensional region for each increment of time, i.e., signs of motion within the scan data, is determined.

[0031] In step 180, motion-compensated reconstruction is performed using the motion field 155 from the output of step 160. Information from the motion field 155 is used in the motion-compensated reconstruction to obtain a motion-corrected image 175. The motion-corrected image 175 in the embodiments disclosed herein utilizes information from the motion field 155 and removes the corresponding motion artifacts to create an image without motion artifacts.

[0032] Figures 2A, 2B, and 2C show implementation forms of feature map reconstruction. Figure 2A is a flowchart of method 200. At step 210 of method 200, at a certain point in time, for example, T1 is selected. The projection view represents the data of the scanned subject and is obtained from the scan of the subject. For example, the projection view is obtained from the detector of a CT device. The projection view includes the unprocessed data of the scanned subject and represents the trajectory of one cycle of scanning as shown in Figure 2B. At step 210 of method 200, the projection view at time T1 of the trajectory is obtained.

[0033] At step 230, as shown in Figure 2B, two data ranges corresponding to the projection views on both sides at time T1 of the trajectory are selected. These data ranges are the same but exist on both sides. In the exemplary implementation shown in Figure 2B, the time ranges of the selected data ranges 235 and 235' respectively correspond to an angle of 90° from the time point T1 selected at step 210 of process 200, but the data ranges can be changed based on, for example, the speed of scanning the subject and the size and position of the target features of the subject being scanned.

[0034] At step 250, feature map reconstruction is performed by extracting features from the projection data corresponding to ranges 235 and 235'. Thereby, a feature map pair 270 corresponding to the time point selected at step 210 can be obtained.

[0035] Conventionally, overall image reconstruction has been performed within the data range of the projection view of the scanned subject. In the implementation forms of the embodiments disclosed herein, partial reconstruction is performed to obtain a feature map. Step 250 can be executed in any of the two methods disclosed herein. In the first implementation form of the exemplary embodiment, projection-based feature extraction is performed.

[0036] In the first implementation form of step 250, features are extracted and enhanced from the unprocessed projection data of the two selected data ranges, and a feature map is generated. That is, a feature extraction process is applied to the projection data of the two selected data ranges to obtain feature data, and a plurality of pairs of feature maps can be reconstructed based on the feature data. To perform such enhancement, a feature enhancement filter is used. Usually, a high-pass filter or a band-pass filter can be used for feature enhancement. In a specific exemplary embodiment, non-linear transformation can be used to perform the enhancement. In one example, feature extraction and enhancement are performed by implementing a deep neural network.

[0037] In this example, the feature map is generated using the extracted features using any of the various methods described above. It is also possible to generate a feature map of the scanned subject using the projection data with the back-projected features enhanced. Two feature maps are generated for the two selected data ranges of the projection data. A feature map pair including the two feature maps, for example, FM11 and FM12, is generated. The feature map pair for the first two data ranges can have a high temporal resolution, for example, due to the use of short data ranges.

[0038] In the second implementation form of step 250, image-based feature extraction is performed. In such an implementation form, using the back-projected projection data corresponding to two selected data ranges, a partial reconstruction image pair, P11 and P12, is obtained. P11 and P12 are partial angle reconstruction (PAR) images generated from two selected data ranges. Next, using the partial reconstruction image pair P11 and P12, a feature map pair FM11 and FM12 is generated. That is, by applying the feature extraction process to the PAR image, a feature map can be obtained. As described above, each of the feature map pairs FM11 and FM12 is generated by using a feature enhancement filter such as, for example, a high-pass filter or a band-pass filter, a non-linear transformation, or a feature extraction deep neural network.

[0039] The next steps 210 and 250 can be repeated for additional points in the trajectory. That is, as shown in FIG. 2A, after step 250, it is determined whether there are remaining unselected points (step 290). If there are remaining points (step 290 affirmative), the processes of steps 210 and 250 are repeated. For example, a new point T2 is selected on the trajectory as shown in FIG. 2C. Step 250 is repeated for point T2 in order to obtain the corresponding feature map pair, for example, FM21 and FM22. On the other hand, if there are no remaining unselected points (step 290 negative), that is, when the processes of steps 210 and 250 for each point at which the projection data has been collected are completed, the process ends.

[0040] Using the generated feature map pair, a motion field is generated. Accurate calculation of the motion field is an important aspect in 4D medical imaging because the motion field can affect the change in image intensity over time. In the embodiments disclosed herein, the feature map is generated to create a 4D motion field of the scanned subject.

[0041] In a first embodiment of 4D motion estimation, a first step for generating a 4D motion field is 3D image registration. FIG. 3 shows a method 300 for generating a 4D motion field from a pair of feature maps using 3D image registration.

[0042] In step 310, 3D image registration, for example, conventional non-rigid image registration is performed. To perform this step, each of the pairs of feature maps 335, 335' obtained from step 250 corresponding to all time points of the scan trajectory is used as an input. In one implementation, 3D image registration generates a 3D subject shape using the pair of feature maps at each time point.

[0043] In the foregoing embodiment, each result of 3D image registration, for example, the output of step 310, is further used in step 330 to generate a 3D motion field 355 at each time point. For example, MVF1(x, y, z) corresponding to the pair of feature maps FM11 and FM12 at time point T1, MVF2(x, y, z) corresponding to the pair of feature maps FM21 and FM22 at time point T2, and maximum MVFN(x, y, z) corresponding to the nth pair of feature maps FMN1 and FMN2 at time point TN are generated.

[0044] In step 350, the 3D motion field 355, or "MVF1(x, y, z), MVF2(x, y, z), MVF3(x, y, z),..., MVFN(x, y, z)" is collated through motion fitting at each time point to obtain a 4D motion field MVF(x, y, z, t) 375. The 4D motion field 375 provides a detailed description of any motion of the scanned subject.

[0045] In the second embodiment of 4D motion estimation, a 3D Deep Convolutional Neural Network (3D DCNN) is trained to output a 3D motion field based on an input feature map. A Deep Convolutional Neural Network (DCNN) is a deep learning (DL) network that typically includes a large number of hidden layers exceeding five layers, and these hidden layers are used to extract more features and improve the accuracy of prediction. Generally, DL can be adapted to image processing to improve the spatial resolution of images and reduce noise. Compared with conventional methods, DL does not require accurate noise and edge modeling and depends only on the training dataset. Furthermore, DL has the ability to capture inter-layer image features and construct an advanced network between noisy observed images and potential vivid images.

[0046] FIG. 4 is a flowchart of a method 400 that includes two processes, namely, a process 410 for offline training of a 3D DCNN and a process 440 for generating a 4D motion field from a pair of feature maps using the trained 3D DCNN.

[0047] In process 410, offline training of a 3D DL network 435 is performed, and as a result, a 3D DL network is output in step 430. In one example, the offline 3D DL training process 410 uses a number of sets of 3D feature map pairs 415, such as FM11 and FM12 generated by the process shown in FIG. 2A, to train the 3D DL network 435. The 3D DL network 435 basically performs an image registration process. For example, the 3D DL network records a set of feature map pairs 415, such as a pair like FM11 and FM12, and obtains the difference in the position of each feature between the pairs that is the motion.

[0048] In process 440, at step 450, each of the feature map pairs 445, 445' utilizes the trained 3D DL network 435 to generate the corresponding 3D motion field at each time point, such as "MVF1(x,y,z) corresponding to the feature map pair FM11 and FM12 at time point T1, MVF2(x,y,z) corresponding to the feature map pair FM21 and FM22 at time point T2, the maximum MVFN(x,y,z) corresponding to the nth feature map pair FMN1 and FMN2 at time point TN". The 3D motion field 465 at each time point corresponds to the position of each part of the scanned subject, that is, the features within the three-dimensional region.

[0049] At step 460, the 3D motion fields 465 (MVF1(x,y,z), MVF2(x,y,z), MVF3(x,y,z), …, MVFN(x,y,z)) at each time point are collated through motion fitting at each time point to obtain the 4D motion field MVF(x,y,z,t) 485. The 4D motion field 485 provides a detailed description of any movement of the scanned subject.

[0050] Figure 5 shows an alternative implementation of 4D motion field estimation. Method 500 includes two processes, namely, process 510 for training a 4D DCNN offline, and process 540 for generating a 4D motion field from a feature map pair using the 4D DCNN.

[0051] In process 510 of method 500, the 4D offline training of the 4D DL network 535 is performed. At step 530, a set of feature map pairs 515, such as FM11 and FM12, are used as training data, and the 4D DL network 535 is trained in the same manner as the training of the 3D DL network 435. As a result of the training, a 4D DL network is output from step 530.

[0052] In process 540, for each of the feature map pairs 545, 545', a 4D motion field, 4D MVF(x,y,z,t) 565 is generated using the trained 4D DL network 535. The 4D motion field 565 provides a detailed description of any motion of the scanned subject.

[0053] Figures 6A, 6B, 6C, and 6D show various examples of the 3D DL network 435 or the 4D DL network 535.

[0054] Figure 6A shows an example of a general artificial neural network (ANN) having N inputs, K hidden layers, and 3 outputs. Each layer is composed of nodes (also called neurons), and each node performs a weighted sum of the inputs and compares the result of the weighted sum with a threshold to generate an output. The ANN creates a class of functions, where the members of this class are obtained by changing details of the structure such as the threshold, connection weights, or the number and / or connectivity of the nodes. The nodes of the ANN may be called neurons (or neuron nodes), and these neurons may have interconnections between different layers of the ANN system. The simplest ANN has three layers and is called an autoencoder. The 3D DL network 435 or the 4D DL network 535 generally has more than three layers of neurons and has the same number of output neurons as input neurons, where N is the number of pixels of the motion field. Synapses (i.e., the junctions between neurons) store a value called a "weight" (also called a "coefficient" or "weighting coefficient" in the same sense) and process data during calculations. The output of the ANN depends on three types of parameters: (i) the interconnection pattern between different layers of neurons, (ii) the learning process for updating the weights of the interconnections, and (iii) the activation function that converts the weighted inputs of the neurons into their output activations.

[0055] Mathematically, the neural network function is defined as a composition of other functions, which can in turn be defined as a composition of yet other functions. This may be conveniently represented as a network structure with arrows representing the dependencies between variables, as shown in FIG. 6. For example, an ANN may use a non-linear weighted sum, where K (usually called an activation function) is some predefined function such as the hyperbolic tangent.

[0056] In FIG. 6A (and similarly in FIG. 6B), the neurons (i.e., nodes) are depicted as circles around a threshold function. In the non-limiting example shown in FIG. 6A, the inputs are depicted as circles around a linear function, and the arrows indicate directed connections between the neurons. In certain implementations, the 3D DL network 435 or the 4D DL network 535 is a feed-forward network as exemplified in FIGS. 6A and 6B (which may be represented, for example, as a directed acyclic graph).

[0057] The 3D DL network 435 or the 4D DL network 535 functions to achieve specific tasks such as motion artifact estimation of CT images, noise removal of CT images, etc., by searching within the class of functions F to be learned using a set of observations in order to find one that solves a particular task according to some optimal criterion. For example, in certain implementations, this can be achieved by defining a cost function such that there is no solution with a cost lower than that of the optimal solution (i.e., the optimal solution has the lowest cost). The cost function C is a measure of how far a particular solution is from the optimal solution for the problem to be solved (e.g., the error). The learning algorithm repeatedly searches through the solution space to find the function with the minimum cost. In certain embodiments, the cost is minimized over a sample of the data (i.e., the training data).

[0058] Figure 6B shows a non - limiting example where the 3D DL network 435 or the 4D DL network 535 is a convolutional neural network (CNN). A CNN is a type of ANN that has beneficial properties for image processing and thus has particular relevance for applications of motion estimation. A CNN uses a feed - forward ANN where the connectivity pattern between neurons may represent a convolution during image processing. For example, a CNN may be used in image - processing optimization by using multiple layers of small neuron sets that process a part of the input image called the receptive field. Next, the outputs of those sets may be tiled in an overlapping manner to obtain a better depiction of the original image. This processing pattern may be repeated over multiple layers having alternating convolutional layers and pooling layers.

[0059] Figure 6C shows an example of a 4×4 kernel applied to map values from an input layer representing a 2 - D image to the first hidden layer, which is a convolutional layer. The kernel maps each 4×4 pixel region to the corresponding region of the first hidden layer.

[0060] Following the convolutional layer, the CNN may include local and / or global pooling layers that combine the outputs of neuron clusters within the convolutional layer. Further, in certain implementations, the CNN may also include various combinations of convolutional layers and fully - connected layers with point - wise non - linearities applied at the end of each layer or after each layer.

[0061] CNNs have several advantages regarding image processing. To reduce the number of free parameters and improve generalization, convolutional operations over small regions of the input are incorporated. One significant advantage of a particular implementation of a CNN is the sharing and use of weights across each convolutional layer. This means that the same filter (weight bank) is used as a coefficient for each pixel within the layer. This reduces the memory footprint and improves performance. Compared to other image processing methods, CNNs advantageously use relatively little preprocessing. This means that the network serves to learn filters that were manually designed with conventional algorithms. The main advantage of CNNs is that they do not rely on prior knowledge and human effort when designing features.

[0062] Figure 6D shows an implementation of a 3D DL network 435 that utilizes the similarity between adjacent layers of a reconstructed 3D medical image. Signals within adjacent layers are typically highly correlated, while noise is not. That is, generally, 3D volume-based images in CT can capture more volume-based features and thus can usually provide more diagnostic information than single-slice cross-sectional 2D images. Based on this insight, particular implementations of the methods described herein use volume-based deep learning algorithms to improve CT images. This insight and corresponding methods are also applicable to other medical imaging fields such as MRI and PET.

[0063] As shown in FIG. 6D, the slice and adjacent slices (i.e., the slices above and below the central slice) are identified as inputs to three channels for the network. To these three layers, a kernel of "W×W×3" is applied M times, generating M values for the convolutional layer, which are then used for subsequent network layer / layers (e.g., pooling layer). This "W×W×3" can be implemented by considering it as three "W×W" kernels applied respectively as three-channel kernels. These are applied to three slices of volume-based image data, and the result becomes the output for the central layer and is used as the input for subsequent network layers. The value M is the total number of filters for a given slice of the convolutional layer, and W is the kernel size (e.g., in FIG. 6C, W = 4).

[0064] In one embodiment, it can be understood that the foregoing implementations of motion estimation and compensation are applicable to a computed tomography (CT) apparatus or scanner. FIG. 7 shows an implementation of a radiation imaging gantry provided in a CT apparatus or scanner. As shown in FIG. 7, the radiation imaging gantry 9900 is depicted as viewed from the side and further includes an X-ray tube 9901, an annular frame 9902, and a multi-row or two-dimensional array type X-ray detector 9903. The X-ray tube 9901 and the X-ray detector 9903 are placed on the annular frame 9902, diametrically opposite across a subject such as a patient, and the annular frame is rotatably supported about a rotation axis RA. The rotation device 9907 rotates the annular frame 9902 at a high speed, for example, 0.4 seconds / rotation, and at the same time, the subject is moved along the axis RA in the direction into or out of the plane shown. That is, the CT scanner apparatus shown in FIG. 7 can acquire a set of projection data by helical scanning. Of course, the scanning method executed by the CT scanner apparatus is not limited to this. For example, the CT scanner apparatus may acquire a set of projection data by volume scanning. That is, the CT scanner apparatus may rotate the annular frame 9902 without moving the subject along the axis RA to acquire a set of projection data.

[0065] Embodiments of an X-ray computed tomography (CT) apparatus according to the present disclosure will be described below with reference to the accompanying drawings. The X-ray CT apparatus includes, for example, a rotation / rotation type apparatus in which both an X-ray tube and an X-ray detector rotate around a subject to be examined, and a fixed / rotation type apparatus in which a large number of detector elements are arranged in an annular or horizontal shape and only the X-ray tube rotates around the subject to be examined. It should be noted that the present disclosure is applicable to any type. Here, the rotation / rotation type, which is currently the mainstream, will be exemplified.

[0066] The multi-slice X-ray CT apparatus further includes a high-voltage generator 9909 that generates a tube voltage applied to the X-ray tube 9901 through a slip ring 9908 so that the X-ray tube 9901 generates X-rays. The X-ray detector 9903 is disposed on the opposite side of the X-ray tube 9901 across the subject to detect the irradiated X-rays that have propagated through the subject. The X-ray detector 9903 is, for example, a photon-counting detector. The X-ray detector or the photon-counting detector 9903 further includes individual detector elements or units such as a processing circuit.

[0067] The CT apparatus further includes other devices that process the detection signals from the X-ray detector 9903. A data acquisition circuit or a data acquisition system (DAS) 9904 converts the signal output from the X-ray detector 9903 for each channel into a voltage signal, amplifies the signal, and further converts the signal into a digital signal. The X-ray detector 9903 and the DAS 9904 are configured to manage a predetermined total number of projections per rotation (TPPR).

[0068] The above data is transmitted through the contactless data transmitter 9905 to the preprocessing device 9906 housed in a console outside the radiography gantry 9900. The preprocessing device 9906 performs specific corrections such as sensitivity correction on the unprocessed data, or various implementations of motion estimation and compensation in CT scan images.

[0069] In one embodiment, the preprocessing device 9906 executes various implementations of motion estimation and compensation in CT scan images as described in the above embodiments. The memory 9912 stores the result data, which is also called projection data at a stage immediately before the reconstruction process. The memory 9912 is connected to the system controller 9910 through the data / control bus 9911 together with the reconstruction device 9914, the input device 9915, and the display 9916. The system controller 9910 controls the current regulator 9913 that limits the current to a level sufficient to drive the CT system.

[0070] Note that FIG. 7 is merely an example, and various modifications are possible for the specific device configuration. For example, the various processes described in the above embodiments may be executed only by the preprocessing device 9906, or may be distributed and executed by the preprocessing device 9906 and the reconstruction device 9914. For example, the preprocessing device 9906 executes feature map reconstruction using the unprocessed data 115 to generate the feature map 135, and executes motion estimation using the feature map 135 to estimate the motion field 155. Then, the reconstruction device 9914 executes motion compensation reconstruction using the motion field 155 to obtain the motion-corrected image 175.

[0071] The preprocessing device 9906 and the reconstruction device 9914 are realized by a processing circuit. For example, the processing circuit is a processor that realizes the functions corresponding to each program by reading and executing the program from the memory 9912. In other words, the processing circuit in the state of having read the program has the functions corresponding to the read program.

[0072] For example, the processing circuit includes an acquisition function, a generation function, an estimation function, and a reconstruction function. The acquisition function is an example of an acquisition unit. The generation function is an example of a generation unit. The estimation function is an example of an estimation unit. The reconstruction function is an example of a reconstruction unit. Note that such a processing circuit may be provided as a part of the CT apparatus as shown in FIG. 7, or may be provided in a medical image processing apparatus separate from the CT apparatus.

[0073] For example, the processing circuit acquires the unprocessed data 115 by reading and executing a program corresponding to the acquisition function. Further, the processing circuit generates the feature map 135 by reading and executing a program corresponding to the generation function and performing feature map reconstruction using the unprocessed data 115. Further, the processing circuit estimates the motion field 155 by reading and executing a program corresponding to the estimation function and performing motion estimation using the feature map 135. Further, the processing circuit acquires the motion-corrected image 175 by reading and executing a program corresponding to the reconstruction function and performing motion compensation reconstruction using the motion field 155.

[0074] The term "processor" used in the above description means, for example, a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or a circuit such as an application specific integrated circuit (ASIC), a programmable logic device (for example, a simple programmable logic device (SPLD), a complex programmable logic device (CPLD), and a field programmable gate array (FPGA)). Further, instead of storing a program in the storage unit, the program may be directly incorporated into the circuit of the processor. In this case, the processor realizes its function by reading and executing the program incorporated in the circuit.

[0075] The detector is rotated and / or fixed with respect to a subject to be scanned, such as a patient, regardless of the generation of the CT scanner system. In one implementation, the above-described CT system may be an example in which a third-generation geometry system and a fourth-generation geometry system are combined. In the third-generation system, the X-ray tube 9901 and the X-ray detector 9903 are placed on the annular frame 9902 in diametrically opposite positions and rotate around the subject when the annular frame 9902 rotates about the rotation axis RA. In the fourth-generation geometry system, the detector is fixedly installed around the patient, and the X-ray tube 9901 rotates around the patient. In an alternative embodiment, the radiation imaging gantry 9900 has a number of detectors arranged on an annular frame 9902 supported by a C-arm and a stand.

[0076] Memory 9912 may store measurement values indicating the X-ray exposure dose in the X-ray detector device 9903. Further, the memory 9912 can store, for example, a dedicated program for executing various steps of a method for motion estimation and compensation in a CT scan image, such as method 100.

[0077] The post-reconstruction processing performed by the reconstruction device 9914 can include, as necessary, image filtering and image smoothing, volume rendering processing, and image difference processing. The image reconstruction process can execute various CT image reconstruction methods. The reconstruction device 9914 can use the memory to store, for example, projection data, reconstructed images, calibration data and parameters, and computer programs.

[0078] Further, the memory 9912 may be a non-volatile memory such as ROM, EPROM, EEPROM, or FLASH memory. The memory 9912 can be a volatile memory such as static or dynamic RAM, and a processor such as a microcontroller or microprocessor may be provided to manage the interaction between the electronic memory and FPGA or CPLD and the memory.

[0079] In one implementation form, the reconstructed image may be displayed on the display 9916. The display 9916 can be an LCD display, a CRT display, a plasma display, an OLED, an LED, or any other display known in the art. The memory 9912 can be a hard disk drive, a CD-ROM drive, a DVD drive, a FLASH drive, a RAM, a ROM, or any other electronic storage device known in the art.

[0080] While specific implementations have been described, these implementations are presented by way of example only and are not intended to limit the teachings of the present disclosure. In fact, the novel methods, apparatuses, and systems described herein can be embodied in various other forms, and furthermore, various omissions, replacements, and changes in the forms of the methods, apparatuses, and systems described herein may be made without departing from the spirit and scope of the present disclosure.

[0081] According to at least one of the embodiments described above, reduction of motion artifacts can be achieved at low cost.

[0082] Although some embodiments of the present invention have been described, these embodiments are presented by way of example and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, replacements, and changes can be made without departing from the gist of the invention. These embodiments and their modifications are included in the scope and gist of the invention, as well as in the invention described in the claims and its equivalent scope.

Explanation of Reference Numerals

[0083] 9900: Radiation imaging gantry 9906: Pretreatment device 9914: Reconfiguration device

Claims

1. Obtaining a data set obtained by scanning a three-dimensional region of a subject; generating, for each of a plurality of time points during the scan, a pair of feature maps representing features of an image reconstructed from a portion of the data set corresponding to the time point, the feature maps being used to estimate motion at the time point; estimating a four-dimensional motion field indicative of changes in the subject's motion over time in the three-dimensional region based on a plurality of pairs of the feature maps corresponding to the plurality of time points; generating a reconstructed image of the subject based on the four-dimensional motion field and the data set; A medical image processing method comprising:

2. 2. The medical image processing method of claim 1, wherein the data set is a set of projection data obtained from a Computed Tomography (CT) scan of the three-dimensional region.

3. The medical image processing method according to claim 1 , further comprising: applying a feature extraction process to at least a portion of the data set to obtain feature data; and reconstructing the plurality of pairs of feature maps based on the feature data.

4. The medical image processing method according to claim 3 , wherein the feature extraction process includes using a feature enhancement filter, using a non-linear transformation, or using a feature extraction machine learning model.

5. 2. The medical image processing method of claim 1, further comprising: reconstructing a plurality of Partial Angle Reconstruction (PAR) images based on the portion; and applying a feature extraction process to the plurality of PAR images to obtain a plurality of pairs of the feature maps.

6. performing a registration process between each pair of the feature maps to generate a three-dimensional motion field for each time point; The medical image processing method according to claim 1 , further comprising fitting a plurality of three-dimensional motion fields generated for the plurality of time points to obtain the four-dimensional motion field.

7. applying each pair of the plurality of pairs of feature maps to a trained machine learning model for motion estimation to generate a 3D motion field for each time point; The medical image processing method according to claim 1 , further comprising fitting the generated three-dimensional motion fields to the multiple time points to obtain the four-dimensional motion field.

8. The medical image processing method of claim 7 , wherein the trained machine learning model for motion estimation is a 3D deep convolutional neural network.

9. The step of applying the machine learning model includes: training a neural network using training data and a function expressing discrepancies between pairs of data as an error value, the training data including pairs including defect-exhibiting data paired with corresponding defect-minimizing data, the neural network performing, for each of the pairs: applying the neural network to paired defect presentation data to generate network processed data; calculating the error value between the network processed data and the defect minimised data of the pair using the function; updating weighting coefficients of the neural network based on the calculated error value; repeating the applying, calculating, and updating steps using each pair of training data until one or more stopping criteria are met; The medical image processing method of claim 7 , further comprising:

10. The method of claim 1 , further comprising: applying each one of the plurality of pairs of feature maps to a trained machine learning model for motion estimation to obtain the four-dimensional motion field.

11. The medical image processing method of claim 9 , wherein the trained machine learning model for motion estimation is a 4D deep convolutional neural network.

12. The medical imaging method of claim 2 , further comprising acquiring the set of projection data using a computed tomography (CT) scanner device.

13. The medical image processing method according to claim 12 , wherein the set of projection data is acquired by helical scanning or volume scanning.

14. an acquisition unit for acquiring a data set obtained by scanning a three-dimensional region of a subject; a generation unit that generates, for each of a plurality of time points during the scan, a pair of feature maps representing features of an image reconstructed from a portion of the data set corresponding to the time point, the feature maps being used to estimate motion at the time point; an estimation unit that estimates a four-dimensional motion field that indicates a change in a motion of the subject over time in the three-dimensional region based on a plurality of pairs of the feature maps corresponding to the plurality of time points; a reconstruction unit for generating a reconstructed image of the subject based on the four-dimensional motion field and the data set; A medical image processing device comprising: