CBCT Artifact Correction System and Method Based on a Multi-Stage Reconstruction Network

Through the CBCT artifact correction system based on multi-stage reconstruction network, the problem of artifact processing in sparse angle CBCT is solved by using the information of the projection domain and the image domain, and the image quality and diagnostic accuracy are significantly improved.

CN119273799BActive Publication Date: 2025-05-27SHANDONG NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411814401.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-05-27
Estimated Expiration
2044-12-11

AI Technical Summary

Technical Problem

The prior art is difficult to effectively process artifacts in sparse angle CBCT imaging, resulting in a decrease in image quality and limited diagnostic accuracy.

Method used

The CBCT artifact correction system based on multi-stage reconstruction network is adopted, and the multi-stage hybrid attention reconstruction network is used to make full use of the supplementary information of the projection domain and the image domain to correct the artifacts in sparse angle CBCT through image preprocessing, model training and image generation modules.

Benefits of technology

The quality of sparse angle CBCT is significantly improved, stripe artifacts and noise is reduced, and image clarity and clinical usability is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119273799B_ABST
    Figure CN119273799B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of medical image processing, and provides a CBCT artifact correction system and method based on a multi-stage reconstruction network. The system includes: an image preprocessing module, a model training module, and an image generation module. The image preprocessing module is used to reconstruct the projection data collected by CBCT scanning to obtain a reference CBCT and a sparse-angle CBCT; the model training module is used to train a multi-stage hybrid attention reconstruction network based on the sparse-angle CBCT and the reference CBCT to obtain a sparse-angle CBCT artifact correction model; the image generation module is used to input the actual sparse-angle CBCT into the sparse-angle CBCT artifact correction model to generate a corrected CBCT. The present invention effectively reduces the common streak artifacts and noise in sparse-angle reconstruction through multi-stage fusion in the image domain and the projection domain, and improves the clarity and clinical usability of CBCT.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and particularly to a CBCT artifact correction system and method based on a multi-stage reconstruction network. Background Art

[0002] Computed tomography (CT) is an important medical imaging technology that can generate detailed internal structure images. Traditional CT obtains multiple projection data by rotating an X-ray beam to reconstruct cross-sectional images of the body interior. Although CT is very effective in diagnosis, it has a high radiation dose, slow imaging speed, and is sensitive to motion artifacts. Cone-beam computed tomography (CBCT) is an improved form of CT that uses a cone-shaped X-ray beam to obtain data for an entire volume in one rotation. CBCT not only has a fast imaging speed and low radiation dose, but also can generate high-resolution three-dimensional images, so it is widely used in fields such as dentistry and oncology.

[0003] In order to reduce the radiation dose of CBCT and improve imaging efficiency, those skilled in the art have begun to focus on sparse-angle CBCT imaging technology. However, due to insufficient acquisition angles, sparse-angle CBCT imaging often produces artifacts, seriously affecting the image quality and diagnostic accuracy. Therefore, developing a method that can effectively process sparse-view projection data and reconstruct high-quality CT images is a task with extremely high clinical value but is extremely challenging.

[0004] In recent years, deep learning technology has made remarkable progress and achievements in the field of sparse-angle reconstruction. In the prior art, deep learning-based methods mainly fall into the following three categories: One is single-domain image-to-image. Single-domain image-to-image mainly relies on the features of existing images for reconstruction, which is easily limited by the quality of the input image. When there are obvious artifacts in the input low-quality image, the reconstruction result may deteriorate further. In addition, this method lacks full utilization of projection data, which may lead to information loss. The second is single-domain projection-to-image. Although single-domain projection-to-image can directly perform image reconstruction from projection data, under sparse-view conditions, the projection data itself has insufficient information, which may lead to problems such as blurred and inaccurate reconstructed images. At the same time, this method has poor robustness in dealing with noise and artifacts and often has difficulty generating high-quality images. The third is dual-domain projection-to-image. Most existing dual-domain projection-to-image methods use interpolation methods to supplement the missing projections of sparse-angle images, which will introduce noise and artifacts during the reconstruction process and cannot provide high-quality input images for subsequent image-domain recovery networks.

[0005] In summary, there is an urgent need to provide a method that can more efficiently utilize dual-domain information, thereby significantly improving the quality of sparse-angle CBCT. Summary of the Invention

[0006] Based on this, the present invention proposes a CBCT artifact correction system and method based on a multi-stage reconstruction network to solve the technical defects existing in the prior art. The system and method can make full use of the complementary information in the projection domain and the image domain, significantly improve the quality of sparse-angle CBCT, simplify the clinical diagnosis process, help doctors diagnose diseases more accurately, and provide better treatment options for patients.

[0007] In a first aspect, the present invention provides a CBCT artifact correction system based on a multi-stage reconstruction network. The system includes: an image preprocessing module, a model training module, and an image generation module. Among them,

[0008] The image preprocessing module is used to preprocess the projection data collected by CBCT scanning to obtain a reference CBCT and a sparse-angle CBCT;

[0009] The model training module is used to train a multi-stage hybrid attention reconstruction network based on the sparse-angle CBCT and the reference CBCT to obtain a sparse-angle CBCT artifact correction model;

[0010] The image generation module is used to input the actual sparse-angle CBCT into the sparse-angle CBCT artifact correction model to generate a corrected CBCT.

[0011] In a possible implementation manner, the model training module includes an image domain restoration module, a projection domain enhancement module, and an image domain denoising module connected in sequence;

[0012] The image domain restoration module includes a first multi-input hybrid attention convolutional neural network; the projection domain enhancement module includes a forward projection module and a second multi-input hybrid attention convolutional neural network connected in series; the image domain denoising module includes an FDK reconstruction module and a third multi-input hybrid attention convolutional neural network connected in series.

[0013] In a possible implementation manner, the first multi-input hybrid attention convolutional neural network, the second multi-input hybrid attention convolutional neural network, and the third multi-input hybrid attention convolutional neural network all include a multi-input encoder, a shallow semantic information fusion attention module, and a single-output decoder.

[0014] In a possible implementation manner, the multi-input encoder includes a first convolutional block, a second convolutional block, a third convolutional block, a first residual block, a second residual block, and a third residual block;

[0015] The first convolutional block is serially connected to the first residual block, the second convolutional block is serially connected to the second residual block, and the third convolutional block is serially connected to the third residual block; one output end of the first residual block is connected to the input end of the second residual block, and the other output end is connected to the shallow semantic information fusion attention module; one output end of the second residual block is connected to the input end of the third residual block, and the other output end is connected to the shallow semantic information fusion attention module; the output end of the third residual block is connected to the shallow semantic information fusion attention module;

[0016] The first convolutional block, the second convolutional block, and the third convolutional block are all composed of a cascaded combination of 3×3 convolution and 1×1 convolution; the first residual block, the second residual block, and the third residual block are all composed of two 3×3 convolutional blocks connected by residual connections and a Relu activation function.

[0017] In a possible implementation, the shallow semantic information fusion attention module includes a first resize module, a second resize module, a third resize module, a first hybrid attention module, a second hybrid attention module, and a third hybrid attention module;

[0018] The first hybrid attention module includes a first information fusion attention module, a first channel attention module, and a first regulation module. The first information fusion attention module has three input ends, and the input ends of the first channel attention module and the first regulation module are connected to the third input end of the first information fusion attention module;

[0019] The second hybrid attention module includes a second information fusion attention module, a second channel attention module, and a second regulation module. The second information fusion attention module has three input ends, and the input ends of the second channel attention module and the second regulation module are connected to the third input end of the second information fusion attention module;

[0020] The third hybrid attention module includes a third information fusion attention module, a third channel attention module, and a third regulation module. The third information fusion attention module has three input ends, and the input ends of the third channel attention module and the third regulation module are connected to the third input end of the third information fusion attention module;

[0021] The output end of the first resize module is respectively connected to the third input end of the first information fusion attention module, the first input end of the second information fusion attention module, and the first input end of the third information fusion attention module; the output end of the second resize module is respectively connected to the first input end of the first information fusion attention module, the third input end of the second information fusion attention module, and the second input end of the third information fusion attention module; the output end of the third resize module is respectively connected to the second input end of the first information fusion attention module, the second input end of the second information fusion attention module, and the third input end of the third information fusion attention module.

[0022] In a possible implementation manner, the output feature of the shallow semantic information fusion attention module is:

[0023] ,

[0024] ,

[0025] ,

[0026] Among them, represents the output feature of the first hybrid attention module, represents the output feature of the second hybrid attention module, represents the output feature of the third hybrid attention module; represents the output feature of the first information fusion attention module; represents the output feature of the second information fusion attention module; represents the output feature of the third information fusion attention module; represents the output feature of the first channel attention module; represents the output feature of the second channel attention module; represents the output feature of the third channel attention module; represents the output feature of the first regulation module; represents the output feature of the second regulation module; represents the output feature of the third regulation module;

[0027] ,

[0028] ,

[0029] ,

[0030] ,

[0031] ,

[0032] ,

[0033] ,

[0034] ,

[0035] ,

[0036] SoftMax represents the SoftMax activation function; represents the output of the first resize module; represents the output of the second resize module; represents the output of the third resize module; and represent upsampling and downsampling respectively; Conv represents convolution; MaxPool represents the max pooling function; AvgPool represents the average pooling function; ReLU represents the ReLU activation function.

[0037] In one possible implementation, the single-output decoder includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a first residual group, a second residual group, a third residual group, and a multi-scale supervision mechanism;

[0038] The first residual group is serially connected to the first convolutional layer, the second residual group is serially connected to the second convolutional layer, the third residual group is serially connected to the third convolutional layer, and the output end of the first convolutional layer is connected to the input end of the second residual group, and the output end of the second convolutional layer is connected to the input end of the third residual group.

[0039] In a second aspect, the present invention provides a CBCT artifact correction method for a CBCT artifact correction system based on the multi-stage reconstruction network described in the first aspect, and the method includes:

[0040] Preprocess the projection data collected by CBCT scanning to obtain a reference CBCT and a sparse-angle CBCT;

[0041] Train a multi-stage hybrid attention reconstruction network based on the sparse-angle CBCT and the reference CBCT to obtain a sparse-angle CBCT artifact correction model;

[0042] Input the actual sparse-angle CBCT into the sparse-angle CBCT artifact correction model to generate a corrected CBCT.

[0043] In a possible implementation, training the multi-stage hybrid attention reconstruction network based on the sparse-angle CBCT and the reference CBCT to obtain a sparse-angle CBCT artifact correction model specifically includes:

[0044] Input the sparse-angle CBCT into the image domain restoration module, and train the first multi-input hybrid attention convolutional neural network to obtain a trained first multi-input hybrid attention convolutional neural network;

[0045] Based on the trained first multi-input hybrid attention convolutional neural network, the image domain restoration module preliminarily removes the artifacts of the sparse-angle CBCT and restores a clean sparse-angle CBCT;

[0046] Input the clean sparse-angle CBCT into the projection domain enhancement module, and use the forward projection module to perform forward projection on the clean sparse-angle CBCT to obtain a projection image;

[0047] Use the projection image to train the second multi-input hybrid attention convolutional neural network to obtain a trained second multi-input hybrid attention convolutional neural network;

[0048] Based on the trained second multi-input hybrid attention convolutional neural network, perform projection domain enhancement on the projection image to obtain a projection image with enhanced projection domain;

[0049] Input the projection image with enhanced projection domain into the image domain denoising module, and use the FDK reconstruction module to reconstruct the projection image with enhanced projection domain to obtain a reconstructed CBCT;

[0050] Use the reconstructed CBCT to train the third multi-input hybrid attention convolutional neural network to obtain a trained third multi-input hybrid attention convolutional neural network, thereby obtaining a sparse-angle CBCT artifact correction model.

[0051] In a possible implementation, the training of the first multi-input hybrid attention convolutional neural network, the training of the second multi-input hybrid attention convolutional neural network, and the training of the third multi-input hybrid attention convolutional neural network specifically include:

[0052] Input the input image into the multi-input encoder, and the multi-input encoder extracts features from the input image to generate a first low-size feature map, a second low-size feature map, and a third low-size feature map;

[0053] Input the first low-size feature map into the first convolutional block, and then obtain a first intermediate feature map after passing through the first residual block;

[0054] Input the second lowest - sized feature map into the second convolutional block, and then input the output feature of the second convolutional block and the first intermediate feature map into the second residual block to obtain the second intermediate feature map;

[0055] Input the third lowest - sized feature map into the third convolutional block, and then input the output feature of the third convolutional block and the second intermediate feature map into the third residual block to obtain the third intermediate feature map;

[0056] Input the first intermediate feature map, the second intermediate feature map, and the third intermediate feature map into the shallow - layer semantic information fusion attention module respectively to obtain the first enhanced intermediate feature map, the second enhanced intermediate feature map, and the third enhanced intermediate feature map;

[0057] Input the third enhanced intermediate feature map into the first residual group for refinement, and then after passing through the first convolutional layer, obtain the first large - sized feature map;

[0058] Input the first large - sized feature map and the second enhanced intermediate feature map into the second residual group for refinement, and then after passing through the second convolutional layer, obtain the second large - sized feature map;

[0059] Input the second large - sized feature map and the first enhanced intermediate feature map into the third residual group for refinement, and then after passing through the third convolutional layer, obtain the output image;

[0060] The multi - scale supervision mechanism performs step - by - step supervision on the output features of each stage based on the reference CBCT. When the multi - scale supervision mechanism stabilizes at a lower value after multiple iterations and no longer decreases significantly, it indicates that the network training is completed.

[0061] The present invention constructs a multi - stage hybrid attention reconstruction network to better handle the problem of information loss in sparse - angle CBCT reconstruction. This network not only performs reconstruction in the image domain but also introduces a collaborative optimization mechanism between the image domain and the projection domain in multiple stages, fusing superficial and semantic information, thus achieving remarkable effects in restoring the global structure and detailed texture. Specifically, the multi - stage hybrid attention reconstruction network of the present invention is divided into three main modules: The image - domain restoration module uses sparse projection data to obtain a preliminary reconstructed image through a convolutional neural network; the projection - domain enhancement module uses a multi - scale input strategy and a hybrid attention mechanism to extract supplementary information from the projection domain to further improve the image quality; the image - domain denoising module combines a multi - scale frequency - domain supervision mechanism to effectively remove residual noise and artifacts, enhancing the contrast and clarity of the reconstructed image. Compared with the prior art, the beneficial effects of the present invention are:

[0062] (1) Through the multi - stage fusion of the image domain and the projection domain, the present invention effectively reduces the common streak artifacts and noise in sparse - angle reconstruction, improving the clarity and clinical usability of CBCT.

[0063] (2) The present invention adopts a multi-scale input and frequency-domain supervision strategy, enabling the optimization of both low-frequency and high-frequency information, and further enhancing the structural details and texture expressiveness.

[0064] (3) The present invention integrates a shallow semantic information fusion attention module, which can efficiently screen and fuse key information between the image domain and the projection domain, effectively improving the reconstruction accuracy of sparse data and ensuring the reliability and integrity of the image data. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0066] Figure 1 It is a structural block diagram of a CBCT artifact correction system based on a multi-stage reconstruction network provided by an embodiment of the present invention;

[0067] Figure 2 It is a schematic diagram of the structure and working process of a multi-input hybrid attention convolutional neural network provided by an embodiment of the present invention;

[0068] Figure 3 It is a schematic diagram of the structure and working process of a shallow semantic information fusion attention module provided by an embodiment of the present invention;

[0069] Figure 4 It is a schematic diagram of the structure and working process of a hybrid attention module provided by an embodiment of the present invention;

[0070] Figure 5 It is a schematic diagram of the process of a CBCT artifact correction method based on a multi-stage reconstruction network provided by an embodiment of the present invention;

[0071] Figure 6 It is a comparison diagram of sparse-angle CBCT / extremely sparse-angle CBCT, reference CBCT, and corrected CBCT provided by an embodiment of the present invention;

[0072] Figure 7 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0073] To better understand the technical solutions of the present invention, the following will describe the embodiments of the present invention in detail with reference to the drawings.

[0074] It should be clear that the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0075] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the embodiments of the present invention and the claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0076] It should be understood that the term " / and" used herein is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A / and B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0077] See Figure 1 , which is a structural block diagram of the CBCT artifact correction system based on the multi-stage reconstruction network provided by the embodiments of the present invention. As Figure 1 shown, the system includes: an image preprocessing module 101, a model training module 102, and an image generation module 103. Among them,

[0078] The image preprocessing module 101 is used to preprocess the projection data collected by CBCT scanning to obtain a reference CBCT and a sparse-angle CBCT. Specifically:

[0079] The projection data collected by CBCT scanning are reconstructed at 600 and 150 angles respectively to obtain a reference CBCT and a sparse-angle CBCT; the sparse-angle CBCT and the reference CBCT are paired one by one to obtain an image data pair.

[0080] Optionally, the sparse-angle CBCT can be replaced by an extremely sparse-angle CBCT, and the extremely sparse-angle CBCT is obtained by reconstructing the projection data collected by CBCT scanning at 40 angles.

[0081] The model training module 102 is used to train a multi-stage hybrid attention reconstruction network based on the sparse-angle CBCT and the reference CBCT to obtain a sparse-angle CBCT artifact correction model.

[0082] The image generation module 103 is used to input the actual sparse-angle CBCT into the sparse-angle CBCT artifact correction model to generate a corrected CBCT.

[0083] Specifically, the model training module 102 includes an image domain restoration module, a projection domain enhancement module and an image domain denoising module which are connected in sequence.

[0084] The image domain restoration module is used to preliminarily remove artifacts of the sparse angle CBCT and restore a clean sparse angle CBCT.

[0085] The image domain restoration module includes a first multi-input mixed attention convolutional neural network.

[0086] The projection domain enhancement module is used to forward project the clean sparse angle CBCT to obtain a projection image; fit the projection image and the sparse angle CBCT to eliminate the difference in details between the projection image and the sparse angle CBCT as much as possible;

[0087] The projection domain enhancement module includes a forward projection module and a second multi-input mixed attention convolutional neural network connected in series.

[0088] The forward projection module is used to forward project the clean sparse angle CBCT through a forward projection algorithm to obtain a projection image.

[0089] The image domain denoising module is used to remove errors generated during image domain restoration, projection domain enhancement and reconstruction.

[0090] The image domain denoising module includes a serially connected FDK reconstruction module and a third multi-input mixed attention convolutional neural network.

[0091] The FDK reconstruction module is used to reconstruct the output image of the image domain denoising module into a reconstructed CBCT by using the FDK algorithm.

[0092] Furthermore, the first multi-input mixed attention convolutional neural network, the second multi-input mixed attention convolutional neural network and the third multi-input mixed attention convolutional neural network have the same network structure, and are all used to combine information from different input sources, using a mixed attention mechanism to focus on important features and suppress irrelevant noise, ultimately improving the effect of CBCT reconstruction.

[0093] See also Figure 2 , which is a schematic diagram of the structure and workflow of a multi-input hybrid attention convolutional neural network provided by an embodiment of the present invention. Figure 2 As shown, the first multi-input mixed attention convolutional neural network, the second multi-input mixed attention convolutional neural network and the third multi-input mixed attention convolutional neural network all include a multi-input encoder, a shallow semantic information fusion attention module and a single output decoder.

[0094] The multi-input encoder is used to process multiple downsampled CBCTs (or projection images). By processing these images in parallel, it extracts richer feature representations.

[0095] Preferably, in this embodiment, the multi-input encoder performs three-layer downsampling on the input image, extracts feature images at three different scales, and inputs the feature images at the three different scales into the input ends corresponding to the scales of the shallow semantic information fusion attention module respectively to obtain multi-scale context feature representations.

[0096] The multi-input encoder includes 3 convolutional blocks and 3 residual blocks. Among them,

[0097] The convolutional block is composed of a cascaded combination of a 3×3 convolution and a 1×1 convolution, and is used to extract the shallow features of the input image and input the combined multi-scale features into the input end corresponding to the scale of the shallow semantic information fusion attention module;

[0098] The residual block is composed of two 3×3 convolutional blocks connected by a residual connection and a Relu activation function, and is used to prevent gradient disappearance during network training and integrate shallow features and multi-scale features.

[0099] Specifically, in this embodiment, the multi-input encoder includes a first convolutional block, a second convolutional block, a third convolutional block, a first residual block, a second residual block, and a third residual block. The first convolutional block is serially connected to the first residual block, the second convolutional block is serially connected to the second residual block, and the third convolutional block is serially connected to the third residual block; one output end of the first residual block is connected to the input end of the second residual block, and the other output end is connected to the shallow semantic information fusion attention module; one output end of the second residual block is connected to the input end of the third residual block, and the other output end is connected to the shallow semantic information fusion attention module; the output end of the third residual block is connected to the shallow semantic information fusion attention module.

[0100] See Figure 3 , which is the structural and working flow chart of the shallow semantic information fusion attention module provided by the embodiment of the present invention; see Figure 4 , which is the structural and working flow chart of the hybrid attention module provided by the embodiment of the present invention. Combining Figure 3 - Figure 4 As shown, the shallow semantic information fusion attention module includes a first resize module, a second resize module, a third resize module, a first hybrid attention module, a second hybrid attention module, and a third hybrid attention module.

[0101] The first hybrid attention module includes a first information fusion attention module, a first channel attention module, and a first regulation module. The first information fusion attention module has three input ends. The input ends of the first channel attention module and the first regulation module are connected to the third input end of the first information fusion attention module, such that the first channel attention module, the first regulation module, and the third input end of the first information fusion attention module are connected in parallel.

[0102] The second hybrid attention module includes a second information fusion attention module, a second channel attention module, and a second regulation module. The second information fusion attention module has three input ends. The input ends of the second channel attention module and the second regulation module are connected to the third input end of the second information fusion attention module, such that the second channel attention module, the second regulation module, and the third input end of the second information fusion attention module are connected in parallel.

[0103] The third hybrid attention module includes a third information fusion attention module, a third channel attention module, and a third regulation module. The third information fusion attention module has three input ends. The input ends of the third channel attention module and the third regulation module are connected to the third input end of the third information fusion attention module, such that the third channel attention module, the third regulation module, and the third input end of the third information fusion attention module are connected in parallel.

[0104] The information fusion attention module is used to explore the anatomical structure connection between the shallow surface information and the deep semantic information from the feature space, and establish the long-distance dependence relationship between the two, which helps to capture fine anatomical details and structures.

[0105] The channel attention module is used to compress the information of the feature channels, so as to cover the information of the entire image, which is beneficial to capturing the global features of the image.

[0106] The input ends of the first resize module, the second resize module, and the third resize module are respectively connected to the output ends of the first residual block, the second residual block, and the third residual block.

[0107] The output end of the first resize module is respectively connected to the third input end of the first information fusion attention module, the first input end of the second information fusion attention module, and the first input end of the third information fusion attention module; the output end of the second resize module is respectively connected to the first input end of the first information fusion attention module, the third input end of the second information fusion attention module, and the second input end of the third information fusion attention module; the output end of the third resize module is respectively connected to the second input end of the first information fusion attention module, the second input end of the second information fusion attention module, and the third input end of the third information fusion attention module.

[0108] The output features of the first information fusion attention module are:

[0109] ,

[0110] The output features of the second information fusion attention module are:

[0111] ,

[0112] The output features of the third information fusion attention module are:

[0113] ,

[0114] Among them, SoftMax represents the SoftMax activation function; represents the output of the m th resize module, m = 1, 2, 3; and respectively represent upsampling and downsampling.

[0115] The output features of the first channel attention module are:

[0116] ,

[0117] The output features of the second channel attention module are:

[0118] ,

[0119] The output features of the third channel attention module are:

[0120] ,

[0121] Among them, represents the output of the m th resize module, m = 1, 2, 3; Conv represents convolution; MaxPool represents the maximum pooling function; AvgPool represents the average pooling function.

[0122] It should be particularly noted that the channel attention module in this embodiment does not serially process the results of the maximum pooling operation and the average pooling operation, but multiplies the two using parallel processing to enhance the information expression specificity between different channels and reduce the influence of irrelevant global information on network training.

[0123] The first regulation module, the second regulation module, and the third regulation module are all composed of a cascaded 3×3 convolution and a ReLU activation function.

[0124] The regulation module is used to initialize a learnable parameter, enabling the shallow semantic information fusion attention module to have the ability to adaptively perceive local or global features, so as to avoid conflicts between the information fusion attention module and the channel attention module during feature extraction.

[0125] The output feature of the first regulation module is:

[0126] ,

[0127] The output feature of the second regulation module is:

[0128] ,

[0129] The output feature of the third regulation module is:

[0130] ,

[0131] Among them, represents the output of the m th resize module, m = 1, 2, 3; Conv represents convolution; ReLU represents the ReLU activation function.

[0132] The output feature of the first hybrid attention module is:

[0133] ,

[0134] The output features of the second hybrid attention module are as follows:

[0135] ,

[0136] The output features of the third hybrid attention module are as follows:

[0137] ,

[0138] The single-output decoder is used to integrate scale features and generate a corrected CBCT.

[0139] The single-output decoder includes three convolutional layers, three residual groups, and a multi-scale supervision mechanism.

[0140] Specifically, in this embodiment, the single-output decoder includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a first residual group, a second residual group, and a third residual group. The first residual group is serially connected to the first convolutional layer, the second residual group is serially connected to the second convolutional layer, and the third residual group is serially connected to the third convolutional layer. Moreover, the output end of the first convolutional layer is connected to the input end of the second residual group, and the output end of the second convolutional layer is connected to the input end of the third residual group.

[0141] The convolutional layer is used to initially integrate scale features and perform an upsampling operation on the low-scale feature map.

[0142] The residual group is used to further extract and enhance features, thereby improving the representation ability of the image.

[0143] The residual group is formed by connecting four residual blocks in series. Each residual block contains two convolutional layers and a skip connection, which is used to retain the original information of the input features and reduce the risk of gradient disappearance. By connecting four residual blocks in series, the single-output decoder can gradually refine features, capture image information at different scales, and achieve a more refined artifact correction effect.

[0144] The multi-scale supervision mechanism is used to supervise the output of the single-output decoder at different scales, enabling the single-output decoder to consider both global and local features during the training process, thereby generating a high-quality corrected CBCT. This multi-scale supervision strategy improves image detail restoration and artifact removal while enhancing the robustness and generalization ability of the model.

[0145] Based on the above embodiments, this embodiment provides a CBCT artifact correction method based on a multi-stage reconstruction network. Refer to Figure 5 , which is a schematic flowchart of the CBCT artifact correction method based on the multi-stage reconstruction network provided by the embodiments of the present invention. As shown in Figure 5As shown, the method includes the following steps:

[0146] Step S1: Preprocess the projection data collected by CBCT scanning to obtain a reference CBCT and a sparse-angle CBCT. Specifically:

[0147] Reconstruct the projection data collected by CBCT scanning at 600 and 150 angles respectively to obtain a reference CBCT and a sparse-angle CBCT; pair the sparse-angle CBCT and the reference CBCT one by one to obtain image data pairs.

[0148] Step S2: Train a multi-stage hybrid attention reconstruction network based on the sparse-angle CBCT and the reference CBCT to obtain a sparse-angle CBCT artifact correction model. Specifically:

[0149] Step S21: Input the sparse-angle CBCT into the image domain restoration module to train the first multi-input hybrid attention convolutional neural network, and obtain a trained first multi-input hybrid attention convolutional neural network;

[0150] Step S22: Based on the trained first multi-input hybrid attention convolutional neural network, the image domain restoration module preliminarily removes the artifacts of the sparse-angle CBCT and restores a clean sparse-angle CBCT;

[0151] Step S23: Input the clean sparse-angle CBCT into the projection domain enhancement module, and use the forward projection module to perform forward projection on the clean sparse-angle CBCT to obtain a projection image;

[0152] Step S24: Use the projection image to train the second multi-input hybrid attention convolutional neural network to obtain a trained second multi-input hybrid attention convolutional neural network;

[0153] Step S25: Perform projection domain enhancement on the projection image based on the trained second multi-input hybrid attention convolutional neural network to obtain a projection domain enhanced projection image;

[0154] Step S26: Input the projection domain enhanced projection image into the image domain denoising module, and reconstruct the projection domain enhanced projection image through the FDK reconstruction module to obtain a reconstructed CBCT;

[0155] Step S27: Use the reconstructed CBCT to train the third multi-input hybrid attention convolutional neural network to obtain a trained third multi-input hybrid attention convolutional neural network, and obtain a sparse-angle CBCT artifact correction model.

[0156] Turn to Figure 2, the training of the first multi-input hybrid attention convolutional neural network, the training of the second multi-input hybrid attention convolutional neural network, and the training of the third multi-input hybrid attention convolutional neural network are specifically as follows:

[0157] Input the input image into the multi-input encoder, and the multi-input encoder extracts features from the input image to generate the first low-size feature map, the second low-size feature map, and the third low-size feature map;

[0158] Input the first low-size feature map into the first convolutional block, and then obtain the first intermediate feature map after passing through the first residual block;

[0159] Input the second low-size feature map into the second convolutional block, and then input the output feature of the second convolutional block and the first intermediate feature map into the second residual block to obtain the second intermediate feature map;

[0160] Input the third low-size feature map into the third convolutional block, and then input the output feature of the third convolutional block and the second intermediate feature map into the third residual block to obtain the third intermediate feature map;

[0161] Input the first intermediate feature map, the second intermediate feature map, and the third intermediate feature map into the shallow semantic information fusion attention module respectively to obtain the first enhanced intermediate feature map, the second enhanced intermediate feature map, and the third enhanced intermediate feature map;

[0162] Input the third enhanced intermediate feature map into the first residual group for refinement, and then obtain the first large-size feature map after passing through the first convolutional layer;

[0163] Input the first large-size feature map and the second enhanced intermediate feature map into the second residual group for refinement, and then obtain the second large-size feature map after passing through the second convolutional layer;

[0164] Input the second large-size feature map and the first enhanced intermediate feature map into the third residual group for refinement, and then obtain the output image after passing through the third convolutional layer.

[0165] During the training process, the multi-scale supervision mechanism performs step-by-step supervision on the output of each stage based on the reference CBCT, enabling the model to effectively learn the ability of artifact correction at each level. When the multi-scale supervision mechanism stabilizes at a lower value after multiple iterations and no longer decreases significantly, it indicates that the network training is completed.

[0166] It should be specifically noted that when training the first multi-input hybrid attention convolutional neural network, the input image is sparse-angle CBCT; when training the second multi-input hybrid attention convolutional neural network, the input image is a projection image; when training the third multi-input hybrid attention convolutional neural network, the input image is reconstructed CBCT.

[0167] S3: Input the actual sparse-angle CBCT into the sparse-angle CBCT artifact correction model to generate corrected CBCT.

[0168] See Figure 6 , which is a comparison diagram of sparse-angle CBCT / extremely sparse CBCT, reference CBCT, and corrected CBCT provided by an embodiment of the present invention. As Figure 6 shown, it can be clearly seen that the images generated by the system and method provided by the embodiments of the present disclosure are very similar to the reference CBCT, which proves the effectiveness of the system and method provided by the embodiments of the present disclosure.

[0169] Corresponding to the above embodiments, an embodiment of the present invention further provides an electronic device.

[0170] See Figure 7 , which is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 7 shown, the electronic device 700 may include: a processor 701, a memory 702, and a communication unit 703. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the electronic device shown in the figure does not constitute a limitation on the embodiments of the present invention. It can be a bus structure, a star structure, and may also include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0171] Among them, the communication unit 703 is used to establish a communication channel so that the electronic device can communicate with other devices.

[0172] The processor 701 is the control center of the electronic device. It connects various parts of the entire electronic device using various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 702, and by invoking the data stored in the memory, it executes various functions of the electronic device and / or processes data. The processor may be composed of an integrated circuit (IC). For example, it may be composed of a single packaged IC, or it may be composed of multiple packaged ICs with the same or different functions connected together. For example, the processor 701 may include only a central processing unit (CPU). In the embodiments of the present invention, the CPU may be a single arithmetic core or may include multiple arithmetic cores.

[0173] The memory 702 is used to store the execution instructions of the processor 701. The memory 702 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.

[0174] When the execution instructions in the memory 702 are executed by the processor 701, the electronic device 700 is enabled to execute some or all of the steps in the above method embodiments.

[0175] Corresponding to the above embodiments, an embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium can store a program. When the program runs, it can control the device where the computer-readable storage medium is located to execute some or all of the steps in the above method embodiments. Specifically, the computer-readable storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.

[0176] Corresponding to the above embodiments, an embodiment of the present invention further provides a computer program product. The computer program product contains executable instructions. When the executable instructions are executed on a computer, the computer is enabled to execute some or all of the steps in the above method embodiments.

[0177] In the embodiments of the present invention, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent the cases of A existing alone, A and B existing simultaneously, and B existing alone. Where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one of the following" and its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, and c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0178] Those of ordinary skill in the art can realize that the various units and algorithm steps described in the embodiments disclosed herein can be implemented by a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0179] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0180] In several embodiments provided by the present invention, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM for short), random access memories (RAM for short), magnetic disks, or optical discs that can store program codes.

[0181] The above is only the specific implementation manner of the present invention. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention and should be covered by the protection scope of the present invention.

Claims

1. A CBCT artifact correction system based on a multi-stage reconstruction network, characterized in that: include: Image preprocessing module, model training module, image generation module; An image preprocessing module, used for preprocessing the projection data acquired by CBCT scanning to obtain reference CBCT and sparse angle CBCT; A model training module is used to train a multi-stage hybrid attention reconstruction network based on sparse angle CBCT and reference CBCT to obtain a sparse angle CBCT artifact correction model; An image generation module, used for inputting the actual sparse angle CBCT into the sparse angle CBCT artifact correction model to generate a corrected CBCT; The model training module includes an image domain restoration module, a projection domain enhancement module and an image domain denoising module which are connected in sequence; The image domain restoration module includes a first multi-input hybrid attention convolutional neural network, which is used to preliminarily remove artifacts of the sparse angle CBCT and restore a clean sparse angle CBCT; The projection domain enhancement module includes a serially connected forward projection module and a second multi-input mixed attention convolutional neural network, which is used for forward projection of the clean sparse angle CBCT by the forward projection module through a forward projection algorithm to obtain a projection image, and performs projection domain enhancement on the projection image to obtain a projection domain enhanced projection image; The image domain denoising module includes a serially connected FDK reconstruction module and a third multi-input mixed attention convolutional neural network, which is used to reconstruct the projection image enhanced in the projection domain to obtain a reconstructed CBCT; based on the reconstructed CBCT, the third multi-input mixed attention convolutional neural network is trained to obtain a sparse angle CBCT artifact correction model; The first multi-input mixed attention convolutional neural network, the second multi-input mixed attention convolutional neural network and the third multi-input mixed attention convolutional neural network each include a multi-input encoder, a shallow semantic information fusion attention module and a single output decoder.

2. The CBCT artifact correction system based on the multi-stage reconstruction network according to claim 1, characterized in that: The multi-input encoder includes a first convolution block, a second convolution block, a third convolution block, a first residual block, a second residual block and a third residual block; The first convolution block is connected in series with the first residual block, the second convolution block is connected in series with the second residual block, and the third convolution block is connected in series with the third residual block; one output end of the first residual block is connected to the input end of the second residual block, and the other output end is connected to the shallow semantic information fusion attention module; One output end of the second residual block is connected to the input end of the third residual block, and the other output end is connected to the shallow semantic information fusion attention module; The output end of the third residual block is connected to the shallow semantic information fusion attention module; The first convolution block, the second convolution block and the third convolution block are all composed of a cascade combination of 3×3 convolutions and 1×1 convolutions; the first residual block, the second residual block and the third residual block are all composed of two 3×3 convolution blocks with residual connections and a Relu activation function.

3. The CBCT artifact correction system based on the multi-stage reconstruction network according to claim 1, characterized in that: The shallow semantic information fusion attention module includes a first resize module, a second resize module, a third resize module, a first hybrid attention module, a second hybrid attention module and a third hybrid attention module; The first hybrid attention module includes a first information fusion attention module, a first channel attention module and a first regulation module, the first information fusion attention module has three input terminals, and the input terminals of the first channel attention module and the first regulation module are connected to the third input terminal of the first information fusion attention module; The second hybrid attention module includes a second information fusion attention module, a second channel attention module and a second regulation module, the second information fusion attention module has three input terminals, and the input terminals of the second channel attention module and the second regulation module are connected to the third input terminal of the second information fusion attention module; The third hybrid attention module includes a third information fusion attention module, a third channel attention module and a third regulation module, wherein the third information fusion attention module has three input terminals, and the input terminals of the third channel attention module and the third regulation module are connected to the third input terminal of the third information fusion attention module; The output end of the first resize module is respectively connected to the third input end of the first information fusion attention module, the first input end of the second information fusion attention module and the first input end of the third information fusion attention module; the output end of the second resize module is respectively connected to the first input end of the first information fusion attention module, the third input end of the second information fusion attention module and the second input end of the third information fusion attention module; the output end of the third resize module is respectively connected to the second input end of the first information fusion attention module, the second input end of the second information fusion attention module and the third input end of the third information fusion attention module.

4. The CBCT artifact correction system based on the multi-stage reconstruction network according to claim 3, characterized in that: The output features of the shallow semantic information fusion attention module are: , , , in, represents the output features of the first hybrid attention module, represents the output features of the second hybrid attention module, Represents the output features of the third hybrid attention module; represents the output features of the first information fusion attention module; represents the output features of the second information fusion attention module; represents the output features of the third information fusion attention module; represents the output features of the first channel attention module; Represents the output features of the second channel attention module; Represents the output features of the third channel attention module; represents the output features of the first regulatory module; represents the output features of the second regulatory module; represents the output features of the third regulatory module; , , , , , , , , , express Activation function; Represents the output of the first resize module; Represents the output of the second resize module; Represents the output of the third resize module; and Respectively represent upsampling and downsampling; represents convolution; Represents the maximum pooling function; represents the average pooling function; Represents the ReLU activation function.

5. The CBCT artifact correction system based on the multi-stage reconstruction network according to claim 2, characterized in that: The single output decoder comprises a first convolutional layer, a second convolutional layer, a third convolutional layer, a first residual group, a second residual group, a third residual group and a multi-scale supervision mechanism; The first residual group is serially connected to the first convolutional layer, the second residual group is serially connected to the second convolutional layer, the third residual group is serially connected to the third convolutional layer, and the output end of the first convolutional layer is connected to the input end of the second residual group, and the output end of the second convolutional layer is connected to the input end of the third residual group.

6. A CBCT artifact correction method based on the CBCT artifact correction system based on a multi-stage reconstruction network according to any one of claims 1 to 5, characterized in that: The method comprises: Preprocessing the projection data acquired by CBCT scanning to obtain reference CBCT and sparse angle CBCT; Based on the sparse angle CBCT and the reference CBCT, a multi-stage hybrid attention reconstruction network is trained to obtain a sparse angle CBCT artifact correction model; The actual sparse angle CBCT is input into the sparse angle CBCT artifact correction model to generate a corrected CBCT.

7. The CBCT artifact correction method according to claim 6, characterized in that: The multi-stage hybrid attention reconstruction network is trained based on the sparse angle CBCT and the reference CBCT to obtain a sparse angle CBCT artifact correction model, specifically: Inputting the sparse angle CBCT into an image domain restoration module, training a first multi-input mixed attention convolutional neural network, and obtaining a trained first multi-input mixed attention convolutional neural network; Based on the trained first multi-input hybrid attention convolutional neural network, an image domain restoration module preliminarily removes artifacts of the sparse angle CBCT and restores a clean sparse angle CBCT; Inputting the clean sparse angle CBCT into a projection domain enhancement module, and forward-projecting the clean sparse angle CBCT using a forward projection module to obtain a projection image; Using the projection image, training a second multi-input mixed attention convolutional neural network to obtain a trained second multi-input mixed attention convolutional neural network; Performing projection domain enhancement on the projection image based on the trained second multi-input mixed attention convolutional neural network to obtain a projection domain enhanced projection image; Inputting the projection domain enhanced projection image into an image domain denoising module, and reconstructing the projection domain enhanced projection image through an FDK reconstruction module to obtain a reconstructed CBCT; The reconstructed CBCT is used to train a third multi-input mixed attention convolutional neural network to obtain a trained third multi-input mixed attention convolutional neural network, and a sparse angle CBCT artifact correction model is obtained.

8. The CBCT artifact correction method according to claim 7, characterized in that: The training of the first multi-input mixed attention convolutional neural network, the training of the second multi-input mixed attention convolutional neural network, and the training of the third multi-input mixed attention convolutional neural network are specifically: Inputting the input image into a multi-input encoder, the multi-input encoder extracts features from the input image to generate a first low-size feature map, a second low-size feature map, and a third low-size feature map; Inputting the first low-size feature map into a first convolution block, and then passing through a first residual block to obtain a first intermediate feature map; Inputting the second low-size feature map into a second convolution block, and then inputting the output features of the second convolution block and the first intermediate feature map into a second residual block to obtain a second intermediate feature map; Inputting the third low-size feature map into a third convolution block, and then inputting the output features of the third convolution block and the second intermediate feature map into a third residual block to obtain a third intermediate feature map; Inputting the first intermediate feature map, the second intermediate feature map and the third intermediate feature map into the shallow semantic information fusion attention module respectively to obtain a first enhanced intermediate feature map, a second enhanced intermediate feature map and a third enhanced intermediate feature map; Inputting the third enhanced intermediate feature map into the first residual group for refinement, and then passing through the first convolution layer to obtain a first large-size feature map; Inputting the first large-size feature map and the second enhanced intermediate feature map into a second residual group for refinement, and then passing through a second convolutional layer to obtain a second large-size feature map; Inputting the second largest size feature map and the first enhanced intermediate feature map into the third residual group for refinement, and then passing through the third convolution layer to obtain an output image; The multi-scale supervision mechanism supervises the output features of each stage step by step based on the reference CBCT. When the multi-scale supervision mechanism stabilizes at a lower value after multiple iterations and no longer decreases significantly, it means that the network training is completed.

Citation Information

Patent Citations

  • Double-domain CT image ring artifact removal method based on deep learning

    CN113554570A

  • Imaging correction method and system based on double-domain data, terminal and medium

    CN117934345A