Computer tomography reconstruction method, program product and device based on deep learning
Through the computer tomography reconstruction method based on deep learning, the measurement projected images are processed and network input is updated, which solves the problem of inter-layer aliasing artifacts and achieves better reconstruction image quality.
Patent Information
- Application Number
- CN202411494978.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-24
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-10-24
AI Technical Summary
When computer tomography technology image large-size flat objects, there are severe interlayer aliasing artifacts, resulting in blurred reconstruction image structure.
Using a computer tomography reconstruction method based on deep learning, the initial reconstruction image is calculated by obtaining the measured projected image, and inputting it into the pre-trained generative adversarial network for processing, outputting the processing image, converting it into a projected image, calculating the residual projection map, performing conversion processing, obtaining an incremental image, updating the current network input image until the loop termination condition is met, and the final reconstruction image is obtained.
Effectively reduce the impact of inter-layer aliasing artifacts, improve the quality and data integrity of the reconstruction image.
Smart Images

Figure CN119478086B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer tomography technology, and in particular to a computer tomography reconstruction method based on deep learning, a computer program product and a computing device. Background Art
[0002] Computed Laminography (CL) is a nondestructive testing technology. In practical applications, due to the limitations of detection space and radiation energy, the imaging effect of Computed Tomography (CT) is not ideal when imaging large-sized flat objects such as integrated circuits and printed circuit boards (PCBs). Since CL technology has different scanning characteristics from CT, it has become a powerful tool to replace CT technology for nondestructive testing of plate-like samples. However, due to the limitations of CL imaging technology itself, its scanning angle is limited, resulting in serious inter-layer aliasing artifacts in its reconstructed image, which manifests as structural blur and other phenomena. Summary of the invention
[0003] In order to solve the existing technical problems, the present invention provides a computer tomography reconstruction method, a computer program product and a computing device based on deep learning, which can reduce the influence of inter-layer aliasing artifacts and make the final reconstruction effect better.
[0004] In a first aspect, a computer tomography reconstruction method based on deep learning is provided, comprising: obtaining a measurement projection map detected by a detection device, and calculating an initial reconstructed image based on the measurement projection map; forming a current network input image of a pre-trained generative adversarial network based on the initial reconstructed image; using the current network input image as an input of the pre-trained generative adversarial network, and outputting a processed image; converting the processed image into a projection image, and obtaining a residual projection map based on the projection image and the measurement projection map; converting the residual projection map to obtain an incremental image, and obtaining an intermediate image based on the current network input image and the incremental image, updating the current network input image according to the intermediate image, continuing to use the updated current network input image as an input of the pre-trained generative adversarial network, until a loop termination condition is met, and using the intermediate image after the loop termination condition as the final reconstructed image.
[0005] In a second aspect, a computer program product is provided, comprising a computer program, characterized in that when the computer program is executed by a processor, the computer tomography reconstruction method based on deep learning as described in any embodiment of the present application is implemented.
[0006] In a third aspect, a computing device is provided, characterized in that it includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the deep learning-based computer tomography reconstruction method as described in any embodiment of the present application.
[0007] In the embodiment of the present application, a measured projection image is reconstructed by a traditional image reconstruction algorithm to obtain an initial reconstructed image, a current network input image of a pre-trained generative adversarial network is obtained based on the initial reconstructed image, and then a processed image is output by the pre-trained generative adversarial network, the processed image is converted into a projection image, and a residual projection image is obtained based on the projection image and the measured projection image, the residual projection image is converted and processed by a traditional image reconstruction algorithm to obtain an incremental image, the current network input image is compensated based on the incremental image to obtain an intermediate image, and then the intermediate image is used as the updated current network input image, and the updated current network input image is continued to be used as the input of the pre-trained generative adversarial network, and the incremental images are obtained in a continuous cycle, the current network input image is compensated and optimized by the incremental images, and the final reconstructed image is obtained after the loop termination condition is met, so that the final reconstructed image can reduce the influence of inter-layer aliasing artifacts and make the data of the final reconstructed image more complete. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 A diagram of an application environment of a computer tomography reconstruction method based on deep learning in one embodiment;
[0009] Figure 2 is a flowchart of a computer tomography reconstruction method based on deep learning in one embodiment;
[0010] Figure 3 is an overall block diagram of a computer tomography reconstruction method based on deep learning in one embodiment;
[0011] Figure 4 A network diagram of a generative adversarial network in one embodiment;
[0012] Figure 5 is a schematic diagram of an image domain generation network in one embodiment;
[0013] Figure 6 is a schematic diagram of a multi-head attention processing block in one embodiment;
[0014] Figure 7 is a schematic diagram of a computer tomography reconstruction device based on deep learning in one embodiment;
[0015] Figure 8 is a schematic diagram of a computing device in one embodiment. DETAILED DESCRIPTION
[0016] The technical solution of the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those generally understood by those skilled in the art of the technical field of the present invention. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the scope of protection of the present invention. The term "and / or" used herein includes any and all combinations of one or more related listed items.
[0018] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it should be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0019] See also Figure 1 , is an application environment diagram of a computer tomography reconstruction method based on deep learning in one embodiment. The computer tomography reconstruction method based on deep learning is applied to a computing device 10, which can obtain a measurement projection map, and the measurement projection map can be acquired by a detection device. The ray source in the detection device rotates around the object to be measured, and obtains the measurement projection map of the object to be measured from different angles. These projection data are then used to reconstruct a cross-sectional image of the scanned area. The acquisition process of the measurement projection map involves the attenuation of the ray source. Tissues of different densities absorb the ray source to different degrees, and this difference is used to distinguish different tissue structures. During the scan, the ray tube emits a narrow beam of rays, and the detector opposite the tube measures the ray attenuation. The data collected by the detection device are ray projection and contour data. The computing device 10 can reconstruct the measurement projection map based on the computer tomography reconstruction method of deep learning to obtain a reconstructed image of the object to be measured.
[0020] See also Figure 2 , is a flow chart of a deep learning-based computer tomography reconstruction method provided in an embodiment of the present application. The deep learning-based computer tomography reconstruction method is applied to a computing device, and the deep learning-based computer tomography reconstruction method includes the following steps:
[0021] S11, obtaining a measurement projection image detected by a detection device, and calculating an initial reconstructed image based on the measurement projection image.
[0022] In this embodiment, the measured projection image is reconstructed using a traditional image reconstruction algorithm, which includes but is not limited to a filtered back projection (FBP) algorithm, a super-Laplacian overlapping group sparse prior reconstruction algorithm, and the like.
[0023] In an optional embodiment, based on the measured projection image, calculating the initial reconstructed image includes:
[0024] Construct an optimization objective function, the parameters of which include the reconstructed image and the projection matrix
[0025] Initialize the parameters in the objective function, iteratively optimize the optimization objective function using an iterative algorithm, and when an iterative optimization termination condition is met, use the reconstructed image after the iterative optimization termination condition as the initial reconstructed image.
[0026] The optimization objective function is:
[0027]
[0028] Where A is the projection matrix, p is the measured projection image, μ is the fidelity term coefficient, is the regularization term, λ is the regularization term coefficient, and f is the reconstructed image.
[0029] The iterative optimization termination condition includes but is not limited to the number of iterations being greater than a preset number.
[0030] S12. Forming a current network input image of a pre-trained generative adversarial network based on the initial reconstructed image.
[0031] In this embodiment, the generative adversarial network is pre-trained based on a training sample data set. The generative adversarial network (GAN) consists of a generative network and a discriminative network. The task of the generative network is to generate a CL image with inter-layer aliasing artifacts removed based on the input data. The task of the discriminative network is to distinguish between images generated by the generative network and CL images without inter-layer aliasing artifacts. The discriminative network attempts to distinguish whether the image generated by the generative network looks real enough. The generative network and the discriminative network confront each other, the generative network generates realistic samples as much as possible, and the discriminative network tries to determine whether the image generated by the generative network is real or fake.
[0032] In this embodiment, the initial reconstructed image can be directly used as the current network input image, or the initial reconstructed image can be preprocessed and the preprocessed image can be used as the current network input image. The preprocessing operation includes but is not limited to adjusting the image size, denoising, etc.
[0033] S13. Use the current network input image as the input of the pre-trained generative adversarial network and output the processed image.
[0034] In this embodiment, since the generative adversarial network is a pre-trained network, after the current network input image is input into the pre-trained generative adversarial network, the current network input image can be processed by the generative network in the generative adversarial network to obtain a processed image. The processed image has higher resolution and fewer artifacts than the current network input image, and the probability of the processed image being a real sample can be output by the discriminative network in the generative adversarial network.
[0035] S14, converting the processed image into a projection image, and obtaining a residual projection image based on the projection image and the measured projection image.
[0036] In this embodiment, a Radon transform is performed on the processed image to convert the two-dimensional processed image into a one-dimensional projection image, wherein the Radon transform is used to convert the image into a series of projection data at different angles.
[0037] In an optional embodiment, a difference is made between the projection image and the measured projection map to calculate a residual projection map, where the residual projection map represents the difference between the calculated projection image and the measured projection map.
[0038] S15. Convert the residual projection image to obtain an incremental image, and obtain an intermediate image based on the current network input image and the incremental image. Update the current network input image according to the intermediate image, and continue to use the updated current network input image as the input of the pre-trained generative adversarial network until the loop termination condition is met, and use the intermediate image after the loop termination condition as the final reconstructed image.
[0039] In this embodiment, the residual projection image is converted and processed using a traditional image reconstruction algorithm to obtain an incremental image, and the current network input image is compensated based on the incremental image to obtain an intermediate image, and then the intermediate image is used as the updated current network input image, and the updated current network input image is continued as the input of the pre-trained generative adversarial network, so that the incremental image is obtained in a continuous cycle, and the current network input image is compensated by the incremental image. After the loop termination condition is met, the final reconstructed image is obtained, so that the final reconstructed image can reduce the influence of inter-layer aliasing artifacts, and the data of the final reconstructed image is more complete. The loop termination condition includes but is not limited to the number of iterations being greater than the preset number.
[0040] like Figure 3 As shown, Figure 3It is an overall block diagram of a computer tomography reconstruction method based on deep learning in one embodiment; the measurement projection map actually measured by the detection device is used to obtain an initial reconstructed image through step S11, and the current network input image of the first cycle of the generative adversarial network is obtained through step S12. By using the initial reconstructed image as the current network input image of the first cycle of the generative adversarial network, the final reconstructed image can be quickly obtained; through step S13, a processed image is obtained, and through step S14, a residual projection map can be obtained based on the processed image and the measurement projection map; through step S15, an incremental image can be obtained based on the residual projection map, and an intermediate image can be obtained based on the incremental image, and then the intermediate image is used as the input of the generative adversarial network, and S13 to S15 are repeatedly executed until the loop termination condition is met to obtain the final reconstructed image.
[0041] In the above embodiment, the measured projection image is reconstructed by a traditional image reconstruction algorithm to obtain an initial reconstructed image, and the current network input image of the pre-trained generative adversarial network is obtained based on the initial reconstructed image. Then, the pre-trained generative adversarial network outputs a processed image, and the processed image is converted into a projection image. Based on the projection image and the measured projection image, a residual projection image is obtained. The residual projection image is converted and processed by a traditional image reconstruction algorithm to obtain an incremental image. The current network input image is compensated based on the incremental image to obtain an intermediate image. The intermediate image is then used as the updated current network input image, and the updated current network input image is continued to be used as the input of the pre-trained generative adversarial network. In this way, the incremental image is obtained in a continuous cycle, and the current network input image is compensated and optimized by the incremental image. After the loop termination condition is met, the final reconstructed image is obtained, so that the final reconstructed image can reduce the influence of inter-layer aliasing artifacts and improve the reconstruction quality of the final reconstructed image.
[0042] In some embodiments, the generative adversarial network includes an image domain generative network, a wavelet domain generative network, and a discriminative network, and the step of taking the current network input image as the input of the pre-trained generative adversarial network and outputting the processed image includes:
[0043] Processing the current network input image through the image domain generation network to obtain an image domain processed image;
[0044] The image-domain processed image is used as the input of the wavelet-domain generation network, the first image is decomposed into a low-frequency component and a high-frequency component by the wavelet-domain generation network, low-frequency features corresponding to the low-frequency components and high-frequency features corresponding to the high-frequency components are extracted respectively, and an inverse wavelet transform operation is performed based on the low-frequency features and the high-frequency features to obtain the processed image;
[0045] The processed image is used as an input of the discriminant network to obtain a discriminant probability corresponding to the processed image, wherein the discriminant probability represents a probability that the processed image is an artifact-free reconstructed image.
[0046] In this embodiment, if Figure 4 As shown, Figure 4 The network diagram of the generative adversarial network in one embodiment is shown in FIG. 1 ; the generative adversarial network includes an image domain generative network, a wavelet domain generative network and a discriminant network. The image domain generative network is used to improve the resolution of the image and preliminarily remove artifacts, and the wavelet domain generative network is used to further remove artifacts of the reconstructed image. The wavelet domain generative network can estimate the distribution of artifacts more accurately and can more effectively separate artifacts from real image content. The wavelet transform can decompose the image into multi-frequency components like other frequency decomposition methods, so that the data distribution in each wavelet component can be processed and it helps to separate artifacts from the intrinsic content of the image.
[0047] The current network input image is used as the input of the image domain generation network, and the output of the image domain generation network, that is, the image domain processed image, is used as the input of the wavelet domain generation network. The wavelet domain generation network will output the processed image, which is the image generated by the generation network. The processed image is used as the input of the discriminant network, and the probability of the processed image being a real image is determined by the discriminant network.
[0048] In an optional implementation, after wavelet transformation, a low-frequency component and three high-frequency components can be obtained, and the three high-frequency components correspond to different frequencies. Each component is processed using a wavelet domain generation network.
[0049] In the above embodiment, the input image is processed by the image domain generation network, so that the resolution of the image can be improved and the artifacts can be reduced. The output image of the image domain generation network is further processed by the wavelet domain generation network after wavelet decomposition, so that the distribution of artifacts can be estimated more accurately, and the artifacts can be separated from the real image content more effectively, thereby improving the quality of the reconstructed image.
[0050] In some embodiments, the image domain generation network includes an encoder, a multi-head attention processing block, a residual processing block and a decoder, and the image domain generation network is used to process the current network input image to obtain an image domain processed image, including:
[0051] Based on the encoder, extracting a first image domain feature map of the current network input image;
[0052] Based on the first image domain feature map, an input of the multi-head attention processing block is formed, a plurality of linear transformations are performed on the first image domain feature map to obtain a transformation feature map corresponding to each linear transformation, the transformation feature map corresponding to each linear transformation is merged to obtain a second image domain feature map, and a fusion feature map is obtained based on the first image domain feature map and the second image domain feature map;
[0053] Based on the fused feature map, an input of the residual processing block is formed, a third image domain feature map is obtained through processing by the residual processing block, and a decoded input image of the decoder is formed based on the third image domain feature map;
[0054] The decoded input image is input into the decoder, and the decoded input image is decoded to obtain the image domain processed image.
[0055] In this embodiment, if Figure 5 As shown, Figure 5 It is a schematic diagram of an image domain generation network in an embodiment; the image domain generation network includes but is not limited to an encoder, an average pooling layer, a multi-head attention processing module, a bilinear interpolation, a residual processing block, and a decoder. Feature extraction is performed by the encoder. In an optional implementation, the encoder is composed of multiple layers of 3D convolutional layers, which are responsible for gradually extracting features and reducing the spatial dimensions of the image, providing compressed feature representation for the generation process. The encoder uses 10 convolutional blocks, each of which uses a convolutional layer, a batch normalization layer, and an activation layer. The encoder outputs the first image domain feature map, and the average pooling layer averages the first image domain feature map, and the average pooling result is used as the input of the multi-head attention processing block. By using the multi-head attention processing block to capture the dependencies between different regions in the image, the accuracy of reconstruction is improved. The multi-head attention processing block passes its input through multiple sets of linear transformations to obtain a multi-head attention space. These multi-heads can process information in parallel to capture different correlations, and then merge the information output by these multi-heads in dimension to obtain the output of the multi-head attention processing block, that is, to obtain the second image domain feature map. The final output is the weighted sum of the outputs of all multi-heads in the same dimension. Then, the bilinear interpolation method is used to double the spatial dimension of the second image domain feature map, thereby restoring the spatial resolution of the image and obtaining the interpolated feature map. The interpolated feature map is weighted with the first image domain feature map to obtain a fused feature map. Figure 6 As shown, it is a schematic diagram of a multi-head attention processing block in one embodiment. After processing the average pooling result through three linear transformations, three transformed feature maps are obtained respectively. The three transformed feature maps are merged to obtain a merged feature map, that is, a second image domain feature map is obtained.
[0056] The residual processing block is used to extract features from the fused feature map and enhance the network's expressiveness. In an optional implementation, the residual processing block structure includes 8 residual blocks, and two three-dimensional convolutional layers are defined in each residual block. The first convolutional layer is set to 64 output channels, a 5x5 convolution kernel, and a stride of 1. After being processed by the residual processing block, a decoded input image is obtained, and the decoder processes the decoded input image to restore the low-dimensional feature map to a high-dimensional image space. The decoder consists of 10 deconvolution blocks, each of which uses a deconvolution layer, a batch normalization layer, and a ReLU layer. The deconvolution block is designed to gradually restore the spatial resolution of the feature map to generate a target image, that is, to obtain an image domain processed image.
[0057] The decoder gradually restores the low-dimensional feature map to a high-resolution image through a series of transposed convolutional layers and channel number adjustments, and completes the image normalization through the activation function to achieve the purpose of image generation. Batch normalization is used after the convolutional layer to standardize the activation value to help improve training efficiency and model performance. After batch normalization, PReLU (Parametric ReLU) is used as the activation function to increase the nonlinear expression ability of the model. The negative semi-axis slope of the PReLU activation function is learnable, allowing the model to adjust the activation characteristics more flexibly.
[0058] In an optional implementation, the fused feature map is residually connected with the third image domain feature map to obtain the decoded input image.
[0059] In this embodiment, a residual connection is introduced, that is, the third image domain feature map and the fusion feature map are weighted to obtain a decoded input image, which not only helps the depth of feature learning, but also effectively solves the gradient disappearance problem. During the training process, the use of this network also improves the learning ability and training effect of the model.
[0060] In an optional implementation, the network structure of the wavelet domain generation network is similar to the network structure of the image domain generation network, and the network structure of the wavelet domain generation network is not described in detail here.
[0061] In the above embodiment, in the image domain generation network, by performing multiple linear transformations on its input through a multi-head attention processing block, relevant information in the input image can be captured, and more relevant features can be extracted, thereby improving the resolution of the image. By fusing the first image domain feature map and the second image domain feature map, the importance of the features can be dynamically adjusted, and the processing capability of complex features is improved. The residual block can reduce gradient vanishing, and can also extract richer feature information, thereby improving the quality of the reconstructed image.
[0062] In some embodiments, the method further comprises:
[0063] Acquire a training sample data set, wherein each training sample in the training sample data set includes a sample reconstructed image and a label image corresponding to the sample reconstructed image, wherein the label image is a reconstructed image without artifacts;
[0064] Construct an initial generative adversarial network, which includes an initial image domain generative network, an initial wavelet domain generative network and an initial discriminative network;
[0065] In the iterative process, the first iterative process is first performed, and the first iterative process is used to train the initial image domain generation network and the initial wavelet domain generation network, obtain training samples from the training sample data set, and form input images of the image domain generation network in training based on the training samples, obtain a sample image domain processed image through processing by the image domain generation network in training, and calculate a first loss value based on the image domain loss function, the sample image domain processed image and the label image;
[0066] The sample processed image is used as an input image of a wavelet domain generation network in training, the wavelet domain generation network in training decomposes the sample processed image to obtain a sample low-frequency component and a sample high-frequency component, respectively extracts a sample low-frequency feature corresponding to the sample low-frequency component and a sample high-frequency feature corresponding to the sample high-frequency component, performs an inverse wavelet transform operation based on the sample low-frequency feature and the sample high-frequency feature to obtain a sample processed image, and calculates a second loss value based on a wavelet domain loss function, the sample processed image and the label image;
[0067] Calculating a total loss of the generated network based on the first loss value and the second loss value, and continuing to execute the first iterative process based on the total loss of the generated network until the first iterative process meets a first loop termination condition;
[0068] After the first iterative process satisfies the first loop termination condition, a second iterative process is executed, the second iterative process is used to train the initial discriminant network, the sample processed image obtained in the execution of the first iterative process is input into the discriminant network in training, a discriminant result is obtained through the discriminant processing of the discriminant network in training, and a discriminant loss value is calculated based on a discriminant loss function, the discriminant result and a label image corresponding to the sample processed image, and the second iterative process is continued to be executed based on the discriminant loss value until the second iterative process satisfies the second loop termination condition;
[0069] When the iteration process does not meet the iteration termination condition, the first iteration process and the second iteration process are performed alternately until the iteration termination condition is met.
[0070] In this embodiment, during the training process, the initial image domain generation network, the initial wavelet domain generation network and the initial discriminant network are untrained network structures. The image domain generation network and the wavelet domain generation network are similar to the network structures described in the above embodiments and are not described in detail here. The discriminant network includes but is not limited to convolutional layers, flattening layers, activation layers, and fully connected layers. The discriminant network and the generative network are trained in an alternating manner. During the execution of the first iteration, the parameters of the discriminant network are fixed. After the generative network is trained, during the execution of the second iteration, the parameters of the generative network are fixed and the discriminant network is trained. The two are trained alternately, and the final network parameters are obtained by back propagation of the loss function, and finally a trained generative adversarial network is obtained. The iteration termination conditions include but are not limited to the number of iterations being greater than the preset number of iterations.
[0071] During the training process, the generative network takes the CL training samples as input and attempts to output a complete CL reconstructed image and remove inter-layer aliasing artifacts. The discriminative network is used to evaluate the quality of the reconstructed image output by the generative network. The discriminative network distinguishes the generated CL reconstructed image from the real complete CL image, thereby guiding the generative network to improve its output. The generative network and the discriminative network compete with each other through adversarial training. The generative network attempts to generate a complete CL image that is as realistic as possible to deceive the discriminative network, while the discriminative network continuously improves its ability to distinguish between real and generated images. Finally, through continuous training, the generative adversarial network can learn how to reconstruct CL image data and remove inter-layer aliasing artifacts.
[0072] In an optional implementation,
[0073] The image domain loss function is:
[0074] TotalLoss=0.5×PixelLoss+0.5×SmoothLoss
[0075] Among them, TotalLoss is the image domain loss, PixellLoss is the pixel loss, and SmoothLoss is the pixel loss;
[0076] The expression of pixel loss PixelLoss is:
[0077]
[0078] where Y ij represents the value of the label image at the pixel position (i, j), Gx ij represents the value of the sample image in the processed image at the pixel position (i, j),
[0079] The expression of smooth loss SmoothLoss is as follows:
[0080]
[0081] Among them, Horizontal LOSS Represents the loss in the horizontal direction, Vertical LOSS Represents the loss in the vertical direction, horizontal_normal ij Indicates the value of the pixel in the i-th row and j-th column of the sample image domain processed image, horizontal_one_right ij represents the value of the pixel at the i-th row and j+1-th column in the sample image domain processing image, vertical_normal ij Represents the value of the pixel in the i-th column and j-th row in the sample image domain processing image, vertical-one_right ij Represents the value of the pixel in the i+1th column and jth row in the sample image domain processed image.
[0082] In this embodiment, pixel loss is calculated, and the square and average of each pixel difference are taken to measure the accuracy of image reconstruction. Pixel loss directly promotes the generated image to be as close as possible to the target image, focusing on the details and accuracy of the generated image. The goal of smooth loss is to smooth the image by reducing the difference between adjacent pixels to avoid obvious artifacts or unnatural textures in the generated image. It helps ensure that the generated image is spatially coherent, making the image look more natural and smooth.
[0083] In an optional implementation, the discriminant loss function is:
[0084]
[0085] D y Denotes the output prediction value of the discriminant network for the label image, D Gx Represents the output prediction value of the discriminant network for the sample processed image.
[0086] In this embodiment, ‖D y ―1‖ 2 Calculated D y The mean square error between and 1 (L2 loss). Minimizing this loss term aims to make D y Approaching 1, thus optimizing the accuracy of the discriminator in identifying real samples. Gx represents the output prediction value of the discriminator for the generated sample. The target value is 0, indicating the generated sample. Therefore, the term ‖D Gx ―0‖ 2 Calculated D Gx The mean square error between 0 and 0 (L2 loss). Minimizing this loss term aims to make D GxApproaching 0, thus optimizing the accuracy of the discriminator in identifying generated samples.
[0087] In the above embodiment, by training the generative network and the discriminative network with a training data set, the image domain generative network and the wavelet domain generative network in the generative network can learn various features in the CL image. After the training is completed, the image domain generative network can improve the image resolution and extract more detailed features, and the wavelet domain generative network can better reduce the artifacts in the CL image, thereby improving the image reconstruction quality.
[0088] On the other hand, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the deep learning-based computer tomography reconstruction method described in any embodiment of the present application.
[0089] Among them, in the computer program product, an optional implementation form of the program module architecture of the computer program that implements each step of the target identification method can be a computer tomography reconstruction device based on deep learning.
[0090] See also Figure 7 , an embodiment of the present application provides a computer tomography reconstruction device based on deep learning, comprising:
[0091] An acquisition module 70 is used to acquire a measurement projection image detected by a detection device, and calculate an initial reconstructed image based on the measurement projection image;
[0092] The acquisition module 70 is also used to form a current network input image of a pre-trained generative adversarial network based on the initial reconstructed image;
[0093] A computing module 71 is used to use the current network input image as the input of the pre-trained generative adversarial network and output a processed image;
[0094] The calculation module 71 is also used to convert the processed image into a projection image, and obtain a residual projection image based on the projection image and the measured projection image;
[0095] The calculation module 71 is also used to transform the residual projection image to obtain an incremental image, and obtain an intermediate image based on the current network input image and the incremental image, update the current network input image according to the intermediate image, and continue to use the updated current network input image as the input of the pre-trained generative adversarial network until the loop termination condition is met, and use the intermediate image after the loop termination condition as the final reconstructed image.
[0096] Optionally, the calculation module 71 is further used to: process the current network input image through the image domain generation network to obtain an image domain processed image;
[0097] The image-domain processed image is used as the input of the wavelet-domain generation network, and the image-domain processed image is decomposed into a low-frequency component and a high-frequency component through the wavelet-domain generation network, and low-frequency features corresponding to the low-frequency components and high-frequency features corresponding to the high-frequency components are respectively extracted, and an inverse wavelet transform operation is performed based on the low-frequency features and the high-frequency features to obtain the processed image;
[0098] The processed image is used as an input of the discriminant network to obtain a discriminant probability corresponding to the processed image, wherein the discriminant probability represents a probability that the processed image is an artifact-free reconstructed image.
[0099] Optionally, the image domain generation network includes an encoder, a multi-head attention processing block, a residual processing block and a decoder, and the calculation module 71 is further used for:
[0100] Based on the encoder, extracting a first image domain feature map of the current network input image;
[0101] Based on the first image domain feature map, an input of the multi-head attention processing block is formed, a plurality of linear transformations are performed on the first image domain feature map to obtain a transformation feature map corresponding to each linear transformation, the transformation feature map corresponding to each linear transformation is merged to obtain a second image domain feature map, and a fusion feature map is obtained based on the first image domain feature map and the second image domain feature map;
[0102] Based on the fused feature map, an input of the residual processing block is formed, a third image domain feature map is obtained through processing by the residual processing block, and a decoded input image of the decoder is formed based on the third image domain feature map;
[0103] The decoded input image is input into the decoder, and the decoded input image is decoded to obtain the image domain processed image.
[0104] Optionally, the calculation module 71 is further used for:
[0105] The fused feature map is residually connected with the third image domain feature map to obtain the decoded input image.
[0106] Optionally, the calculation module 71 is further used for:
[0107] Acquire a training sample data set, wherein each training sample in the training sample data set includes a sample reconstructed image and a label image corresponding to the sample reconstructed image, wherein the label image is a reconstructed image without artifacts;
[0108] Construct an initial generative adversarial network, which includes an initial image domain generative network, an initial wavelet domain generative network and an initial discriminative network;
[0109] In the iterative process, the first iterative process is first performed, and the first iterative process is used to train the initial image domain generation network and the initial wavelet domain generation network, obtain training samples from the training sample data set, and form input images of the image domain generation network in training based on the training samples, obtain a sample image domain processed image through processing by the image domain generation network in training, and calculate a first loss value based on the image domain loss function, the sample image domain processed image and the label image;
[0110] The sample processed image is used as an input image of a wavelet domain generation network in training, the wavelet domain generation network in training decomposes the sample processed image to obtain a sample low-frequency component and a sample high-frequency component, respectively extracts a sample low-frequency feature corresponding to the sample low-frequency component and a sample high-frequency feature corresponding to the sample high-frequency component, performs an inverse wavelet transform operation based on the sample low-frequency feature and the sample high-frequency feature to obtain a sample processed image, and calculates a second loss value based on a wavelet domain loss function, the sample processed image and the label image;
[0111] Calculating a total loss of the generated network based on the first loss value and the second loss value, and continuing to execute the first iterative process based on the total loss of the generated network until the first iterative process meets a first loop termination condition;
[0112] After the first iterative process satisfies the first loop termination condition, a second iterative process is executed, the second iterative process is used to train the initial discriminant network, the sample processed image obtained in the execution of the first iterative process is input into the discriminant network in training, a discriminant result is obtained through the discriminant processing of the discriminant network in training, and a discriminant loss value is calculated based on a discriminant loss function, the discriminant result and a label image corresponding to the sample processed image, and the second iterative process is continued to be executed based on the discriminant loss value until the second iterative process satisfies the second loop termination condition;
[0113] When the iteration process does not meet the iteration termination condition, the first iteration process and the second iteration process are performed alternately until the iteration termination condition is met.
[0114] Optionally, the image domain loss function is:
[0115] TotalLoss=0.5×PixelLoss+0.5×SmoothLoss
[0116] Among them, TotalLoss is the image domain loss, PixellLoss is the pixel loss, and SmoothLoss is the pixel loss;
[0117] The expression of pixel loss PixelLoss is:
[0118]
[0119] where Y ij represents the value of the label image at the pixel position (i, j), Gx ij represents the value of the sample image in the processed image at the pixel position (i, j),
[0120] The expression of smooth loss SmoothLoss is as follows:
[0121]
[0122] Among them, Horizontal LOSS Represents the loss in the horizontal direction, Vertical LOSS Represents the loss in the vertical direction, horizontal_normal ij Indicates the value of the pixel in the i-th row and j-th column of the sample image domain processed image, horizontal_one_right ij represents the value of the pixel at the i-th row and j+1-th column in the sample image domain processing image, vertical_normal ij Represents the value of the pixel in the i-th column and j-th row in the sample image domain processing image, vertical-one_right ij Represents the value of the pixel in the i+1th column and jth row in the sample image domain processed image.
[0123] Optionally, the discriminant loss function is:
[0124]
[0125] D y Denotes the output prediction value of the discriminant network for the label image, D Gx Represents the output prediction value of the discriminant network for the sample processed image.
[0126] Optionally, the calculation module 71 is further used for:
[0127] An image obtained by adding the network input image and the incremental image is used as the intermediate image.
[0128] See also Figure 8In another aspect of the embodiment of the present application, a computing device 8 is provided, including a memory 3011 and a processor 3012, wherein the memory 3011 stores a computer program, and when the computer program is executed by the processor, the processor 3012 executes the steps of the deep learning-based computer tomography reconstruction method provided in any of the above embodiments of the present application. The computing device 10 is (for example, a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (for example, a smart phone, a wireless phone, etc.), a wearable device (for example, a pair of smart glasses or a smart watch), or a similar device.
[0129] The processor 3012 is the control center, which uses various interfaces and lines to connect various parts of the entire computer device, and executes various functions of the computer device and processes data by running or executing software programs and / or modules stored in the memory 3011, and calling data stored in the memory 3011. Optionally, the processor 3012 may include one or more processing cores; preferably, the processor 3012 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user pages and application programs, etc., and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 3012.
[0130] The memory 3011 can be used to store software programs and modules. The processor 3012 executes various functional applications and data processing by running the software programs and modules stored in the memory 3011. The memory 3011 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 3011 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory 3011 may also include a memory controller to provide the processor 3012 with access to the memory 3011.
[0131] On the other hand, an embodiment of the present application further provides a storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the deep learning-based computer tomography reconstruction method provided in any of the above embodiments of the present application.
[0132] Those skilled in the art can understand that all or part of the processes in the methods provided in the above embodiments can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0133] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. The protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A computer tomography reconstruction method based on deep learning, characterized in that: include: Acquire a measurement projection image detected by a detection device, and calculate an initial reconstructed image based on the measurement projection image; Forming a current network input image of a pre-trained generative adversarial network based on the initial reconstructed image; Take the current network input image as the input of the pre-trained generative adversarial network and output the processed image; Converting the processed image into a projection image, and obtaining a residual projection image based on the projection image and the measured projection image; The residual projection image is transformed to obtain an incremental image, and an intermediate image is obtained based on the current network input image and the incremental image. The current network input image is updated according to the intermediate image, and the updated current network input image continues to be used as the input of the pre-trained generative adversarial network until the loop termination condition is met, and the intermediate image after the loop termination condition is used as the final reconstructed image.
2. The deep learning-based computer tomography reconstruction method according to claim 1, characterized in that: The generative adversarial network includes an image domain generative network, a wavelet domain generative network and a discriminant network. The current network input image is used as the input of the pre-trained generative adversarial network to output a processed image, including: Processing the current network input image through the image domain generation network to obtain an image domain processed image; The image-domain processed image is used as the input of the wavelet-domain generation network, and the image-domain processed image is decomposed into a low-frequency component and a high-frequency component through the wavelet-domain generation network, and low-frequency features corresponding to the low-frequency components and high-frequency features corresponding to the high-frequency components are respectively extracted, and an inverse wavelet transform operation is performed based on the low-frequency features and the high-frequency features to obtain the processed image; The processed image is used as an input of the discriminant network to obtain a discriminant probability corresponding to the processed image, wherein the discriminant probability represents a probability that the processed image is an artifact-free reconstructed image.
3. The deep learning-based computer tomography reconstruction method according to claim 2, characterized in that: The image domain generation network includes an encoder, a multi-head attention processing block, a residual processing block and a decoder. The image domain generation network is used to process the current network input image to obtain an image domain processed image, including: Based on the encoder, extracting a first image domain feature map of the current network input image; Based on the first image domain feature map, an input of the multi-head attention processing block is formed, a plurality of linear transformations are performed on the first image domain feature map to obtain a transformation feature map corresponding to each linear transformation, the transformation feature map corresponding to each linear transformation is merged to obtain a second image domain feature map, and a fusion feature map is obtained based on the first image domain feature map and the second image domain feature map; Based on the fused feature map, an input of the residual processing block is formed, a third image domain feature map is obtained through processing by the residual processing block, and a decoded input image of the decoder is formed based on the third image domain feature map; The decoded input image is input into the decoder, and the decoded input image is decoded to obtain the image domain processed image.
4. The deep learning-based computer tomography reconstruction method according to claim 3, characterized in that: The step of forming an input of the residual processing block based on the fused feature map, obtaining a third image domain feature map through processing by the residual processing block, and forming a decoded input image of the decoder based on the third image domain feature map comprises: The fused feature map is residually connected with the third image domain feature map to obtain the decoded input image.
5. The deep learning-based computer tomography reconstruction method according to claim 1, characterized in that: The method further comprises: Acquire a training sample data set, wherein each training sample in the training sample data set includes a sample reconstructed image and a label image corresponding to the sample reconstructed image, wherein the label image is a reconstructed image without artifacts; Construct an initial generative adversarial network, which includes an initial image domain generative network, an initial wavelet domain generative network and an initial discriminative network; In the iterative process, the first iterative process is first performed, and the first iterative process is used to train the initial image domain generation network and the initial wavelet domain generation network, obtain training samples from the training sample data set, and form input images of the image domain generation network in training based on the training samples, obtain a sample image domain processed image through processing by the image domain generation network in training, and calculate a first loss value based on the image domain loss function, the sample image domain processed image and the label image; The sample image domain processed image is used as an input image of the wavelet domain generation network in training, the wavelet domain generation network in training decomposes the sample image domain processed image to obtain a sample low-frequency component and a sample high-frequency component, respectively extracts a sample low-frequency feature corresponding to the sample low-frequency component and a sample high-frequency feature corresponding to the sample high-frequency component, performs an inverse wavelet transform operation based on the sample low-frequency feature and the sample high-frequency feature to obtain a sample processed image, and calculates a second loss value based on a wavelet domain loss function, the sample processed image and the label image; Calculating a total loss of the generated network based on the first loss value and the second loss value, and continuing to execute the first iterative process based on the total loss of the generated network until the first iterative process meets a first loop termination condition; After the first iterative process satisfies the first loop termination condition, a second iterative process is executed, the second iterative process is used to train the initial discriminant network, the sample processed image obtained in the execution of the first iterative process is input into the discriminant network in training, a discriminant result is obtained through the discriminant processing of the discriminant network in training, and a discriminant loss value is calculated based on a discriminant loss function, the discriminant result and a label image corresponding to the sample processed image, and the second iterative process is continued to be executed based on the discriminant loss value until the second iterative process satisfies the second loop termination condition; When the iteration process does not meet the iteration termination condition, the first iteration process and the second iteration process are performed alternately until the iteration termination condition is met.
6. The deep learning-based computer tomography reconstruction method according to claim 5, characterized in that: The image domain loss function is: TotalLoss=0.5×PixelLoss+0.5×SmoothLoss Among them, TotalLoss is the image domain loss, PixellLoss is the pixel loss, and SmoothLoss is the smooth loss; The expression of pixel loss PixelLoss is: where Y ij represents the value of the label image at the pixel position (i, j), Gx ij represents the value of the sample image in the processed image at the pixel position (i, j), The expression of smooth loss SmoothLoss is as follows: Among them, Horizontal LOSS Represents the loss in the horizontal direction, Vertical LOSS Represents the loss in the vertical direction, horizontal_normal ij Indicates the value of the pixel in the i-th row and j-th column in the sample image domain processing image, horizontal_one_right ij represents the value of the pixel at the i-th row and j+1-th column in the sample image domain processing image, vertical_normal ij Represents the value of the pixel in the i-th column and j-th row in the sample image domain processing image, vertical-one_right ij Represents the value of the pixel in the i+1th column and jth row in the sample image domain processed image.
7. The deep learning-based computer tomography reconstruction method according to claim 5, characterized in that: The discriminant loss function is: D y Denotes the output prediction value of the discriminant network for the label image, D Gx Represents the output prediction value of the discriminant network for the sample processed image.
8. The deep learning-based computer tomography reconstruction method according to any one of claims 1 to 7, characterized in that: The obtaining of the intermediate image based on the network input image and the incremental image comprises: An image obtained by adding the network input image and the incremental image is used as the intermediate image.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the computer tomography reconstruction method based on deep learning as described in any one of claims 1 to 8 is implemented.
10. A computing device, characterized in that It comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the deep learning-based computer tomography reconstruction method as claimed in any one of claims 1 to 8.
Citation Information
Patent Citations
Low-dose cone beam CT image reconstruction method based on deep learning
CN112348936A
3D-CNN processing for CT image denoising
CN115777114A