CBCT artifact correction system and method based on lightweight network

Through the CBCT artifact correction system based on the lightweight parallel Transformer network, the problems of artifact and computing efficiency in sparse angle CBCT imaging are solved, and efficient artifact correction and image detail recovery are achieved.

CN120182407AInactive Publication Date: 2025-06-20SHANDONG NORMAL UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510245152.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-20
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

While reducing radiation dose, the sparse angle CBCT imaging method introduces serious image artifacts, affecting the diagnostic quality, and it is difficult for the prior art to find a balance between computing efficiency and image quality.

Method used

The CBCT artifact correction system based on the lightweight parallel Transformer network is adopted. Through the vertical and horizontal attention extraction links set in parallel, combined with dynamic jump layer connection and feature selection-fusion decoder module, efficient extraction and deep fusion of multidimensional features are achieved.

Benefits of technology

It significantly suppresses artifacts, restores the detailed structure of the image, reduces the computational complexity, meets the application needs of real-time and low-resource environments, and improves image quality and processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182407A_ABST
    Figure CN120182407A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of medical image processing, and provides a CBCT artifact correction system and method based on a lightweight network. The system comprises an image preprocessing unit, a model training unit and an image generation unit. The model training unit comprises an encoder, a decoder and a dynamic skip layer connection; the encoders comprise a plurality of longitudinal attention encoders connected in series and a plurality of transverse attention encoders connected in series, and the longitudinal attention encoders and the transverse attention encoders share weight parameters; the decoder comprises a plurality of decoder modules; output features of the longitudinal attention encoder and the transverse attention encoder at the same level are fused through learnable parameters and then are input to the decoder module at the same level; the learnable parameters can dynamically adjust the contribution ratio of the two output features in the weighted fusion process. According to the method, lightweight design is combined, the requirement for hardware resources is lowered while the artifact correction effect is remarkably improved, and the method is suitable for image reconstruction tasks in clinical medical imaging and other fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and particularly to a CBCT artifact correction system and method based on a lightweight network. Background Art

[0002] Cone Beam Computed Tomography (CBCT) is a three-dimensional imaging technology widely used in the fields of medicine, dentistry, and radiotherapy. Compared with traditional CT, CBCT has many advantages, such as low radiation dose, faster imaging process, and smaller equipment volume, so it is favored in clinical environments such as in-office imaging and bedside imaging. Especially in dentistry and tumor treatment, CBCT provides doctors with clear three-dimensional anatomical structures, which helps with accurate diagnosis and treatment planning. In order to further reduce the radiation dose received by patients, sparse-angle CBCT imaging methods are gradually becoming an ideal choice.

[0003] The main feature of sparse-angle CBCT imaging methods is to obtain less scan data by reducing the projection angles, thereby reducing the radiation dose. However, while this method reduces the patient's radiation exposure, it also introduces serious image artifacts. In conventional CT scans, sufficient projection data ensures image clarity and detail accuracy; but at sparse angles, due to insufficient projection data, the artifacts generated during image reconstruction affect the diagnostic quality. These artifacts usually appear as blurring, streaks, and distortion, masking the true edges and structures of tissues and organs, thus interfering with the clinical effectiveness of the images. Therefore, in the application of sparse-angle CBCT imaging methods, how to effectively correct artifacts has become a key technical challenge.

[0004] Traditionally, the artifact correction methods for sparse-angle CBCT imaging mainly include the filtered back-projection algorithm and the iterative reconstruction algorithm. The filtered back-projection algorithm achieves reconstruction through simple weighted averaging and filtering operations, with high computational efficiency, but performs poorly in the case of insufficient data and is prone to generating streak-like artifacts. The iterative reconstruction algorithm, on the other hand, achieves a better artifact suppression effect by gradually adjusting the error between the image and the projection data. However, the iterative reconstruction algorithm has a high computational complexity and requires a large amount of computing resources, making it difficult to meet the real-time requirements in clinical settings. In addition, existing deep learning-based artifact removal methods (such as convolutional neural networks and generative adversarial networks) have also been applied to the artifact correction of sparse-angle CBCT imaging, but these methods often rely on a large amount of training data and have a high model complexity, making it difficult to achieve efficient processing in low-resource hardware environments.

[0005] In recent years, due to its excellent performance in natural language processing and computer vision tasks, the Transformer network has been gradually introduced into the field of medical image processing. Its global self-attention mechanism can capture long-range feature dependencies, which is beneficial to the restoration of complex structures. However, traditional Transformer networks face problems of excessive computational complexity and high resource consumption in image reconstruction applications. Especially under the real-time processing requirements of sparse-angle CBCT, the standard Transformer architecture is not ideal. Therefore, to address this issue, it is of great significance to develop a lightweight Transformer network that is computationally efficient and can maintain the integrity of the image structure. Summary of the Invention

[0006] The present invention aims to solve the problems of artifacts in sparse-angle CBCT imaging and excessive computational complexity and high resource consumption of traditional Transformer networks in image reconstruction applications. A CBCT artifact correction system and method based on a lightweight network are proposed. By introducing a lightweight parallel Transformer network, artifacts are effectively suppressed, and the detailed structure of the image is restored. While ensuring the image quality, the computational complexity is significantly reduced, meeting the application requirements of real-time and low-resource environments.

[0007] In a first aspect, the present invention provides a CBCT artifact correction system based on a lightweight network, the system comprising: an image preprocessing unit, a model training unit, and an image generation unit;

[0008] The model training unit includes an encoder, a decoder, and a dynamic skip connection;

[0009] The encoder includes a longitudinal attention extraction link and a transverse attention extraction link arranged in parallel. The longitudinal attention extraction link includes a plurality of longitudinal attention encoders connected in series, and the transverse attention extraction link includes a plurality of transverse attention encoders connected in series. The longitudinal attention encoders and the transverse attention encoders share the same weight parameters;

[0010] The decoder includes a plurality of decoder modules connected in sequence;

[0011] The output features of the longitudinal attention encoder and the transverse attention encoder at the same level in the longitudinal attention extraction link and the transverse attention extraction link are weighted and fused through learnable parameters, and then input into the decoder module at the same level through a dynamic skip connection; the learnable parameters can dynamically adjust the contribution ratio of the two output features in the weighted fusion process.

[0012] In a possible implementation, the encoder further includes a downsampling module and a 1×1 convolution;

[0013] A downsampling module is connected to the output end of each longitudinal attention encoder and each transverse attention encoder;

[0014] The output features of the longitudinal attention extraction link and the transverse attention extraction link are fused and then input into a 1×1 convolution.

[0015] In a possible implementation manner, the decoder further includes an upsampling module;

[0016] An upsampling module is connected to the output end of each decoder module.

[0017] In a possible implementation manner, the longitudinal attention encoder includes a first layer normalization layer, a multi-head longitudinal attention module, a second layer normalization layer, a multi-layer perceptron, and a residual block connected in series;

[0018] The input end of the first layer normalization layer is connected to the output end of the multi-head longitudinal attention module through a skip connection; the input end of the second layer normalization layer is connected to the output end of the multi-layer perceptron through a skip connection;

[0019] The multi-head longitudinal attention module includes a first 1×1 convolution, multi-head longitudinal attention, and a second 1×1 convolution connected in series, and the input end of the first 1×1 convolution is connected to the output end of the second 1×1 convolution through a skip connection.

[0020] In a possible implementation manner, the decoder module includes a feature selection module, a channel attention module, and a feature fusion module.

[0021] In a possible implementation manner, the feature selection module includes spatial attention mechanisms connected in parallel, and the output features of the two spatial attention mechanisms are weighted and fused through learnable parameters;

[0022] The spatial attention mechanism includes a third 1×1 convolution and a sigmoid function connected in series, and the input end of the third 1×1 convolution is connected to the output end of the sigmoid function through a skip connection.

[0023] In a possible implementation manner, the channel attention module is composed of the addition of a global average pooling layer and a maximum pooling layer connected in parallel, and then a Softmax function connected in series;

[0024] A skip connection is provided between the input end of the channel attention module and the output end of the Softmax function.

[0025] In a possible implementation manner, the feature fusion module weights and fuses the input features of the feature selection module from the previous decoder module / 1×1 convolution and the output features of the channel attention module through learnable parameters.

[0026] Second aspect, the present invention provides a CBCT artifact correction method for a CBCT artifact correction system based on a lightweight network according to the first aspect, the method comprising:

[0027] Reconstruct the projection data to obtain sparse-angle images and reference images;

[0028] Based on the sparse-angle images and reference images, train to obtain a sparse-angle CBCT artifact correction model;

[0029] Input the actual sparse-angle images into the sparse-angle CBCT artifact correction model to obtain corrected images.

[0030] In a possible implementation, training to obtain a sparse-angle CBCT artifact correction model based on the sparse-angle images and reference images specifically includes:

[0031] Step S21: Input the sparse-angle images into the encoder. The longitudinal attention extraction link extracts longitudinal features, and the transverse attention extraction link extracts transverse features. The output features of the two links are fused and then passed through a 1×1 convolution to obtain the encoder output features;

[0032] The output features of the longitudinal attention encoder and the transverse attention encoder at the same level in the longitudinal attention extraction link and the transverse attention extraction link are weighted and fused through learnable parameters, and then input into the decoder module at the same level through a dynamic skip connection;

[0033] Step S22: Input the encoder output features into the decoder. Each decoder module of the decoder sequentially fuses and samples the input features with the fused features at the same level to obtain an artifact correction image;

[0034] Step S23: Compare the artifact correction image with the reference image, calculate the loss function of the two. If the loss function converges, the sparse-angle CBCT artifact correction model is obtained; if the loss function does not converge, update the parameters through the backpropagation algorithm, and repeat steps S21 - S22 until the loss function converges to obtain the sparse-angle CBCT artifact correction model.

[0035] The sparse-angle CBCT artifact correction system based on the lightweight parallel Transformer network proposed by the present invention solves the balance problem between computational efficiency and image quality in the prior art and has important clinical application value. Compared with the prior art, the beneficial effects of the present invention are:

[0036] (1) Efficient extraction of multi-dimensional features: By designing parallel vertical attention extraction links and horizontal attention extraction links, the feature information of the data is fully mined from different angles to achieve the efficient extraction of multi-dimensional features, significantly improving the effect of artifact correction. At the same time, except for the different multi-head vertical attention and multi-head horizontal attention, the structures of the vertical attention encoder and the horizontal attention encoder are the same, and they share the same weight parameters to achieve parameter sharing, reducing the redundant calculations and the number of parameters of the network. This not only realizes the lightweight of the network structure but also improves the efficiency and consistency of multi-dimensional feature extraction.

[0037] (2) Deep fusion of feature information: The decoder uses the spatial attention mechanism of the feature selection module to selectively fuse the features output by the encoder and the decoder, and combines the channel attention module to optimize the feature expression, effectively retaining and strengthening the key information, and further improving the image quality.

[0038] (3) Lightweight network structure: Through the multi-head horizontal attention module and the multi-head vertical attention module in the present invention, compared with the traditional multi-head self-attention module, the model parameters and computational complexity are reduced. While maintaining excellent performance, the demand for hardware resources is significantly reduced, making it more suitable for clinical practical applications.

[0039] (4) Dynamically adjust the fused information: Learnable parameters are introduced in modules such as dynamic skip connections and decoders for weighted fusion. The feature information of horizontal attention and vertical attention is dynamically adjusted and fused through the learnable parameters, enhancing the decoder's ability to perceive global information and significantly improving the artifact correction performance and computational efficiency of the network. Brief Description of the Drawings

[0040] To more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0041] Figure 1 It is the structural diagram of the CBCT artifact correction system based on a lightweight network provided by the embodiment of the present invention;

[0042] Figure 2 It is the structural diagram of the model training unit of the CBCT artifact correction system based on a lightweight network provided by the embodiment of the present invention;

[0043] Figure 3 It is the structural diagram of the vertical attention encoder of the CBCT artifact correction system based on a lightweight network provided by the embodiment of the present invention;

[0044] Figure 4 The structural diagram of the decoder module of the CBCT artifact correction system based on a lightweight network provided by an embodiment of the present invention;

[0045] Figure 5 The schematic flow chart of the CBCT artifact correction method based on a lightweight network provided by an embodiment of the present invention;

[0046] Figure 6 The comparison chart of sparse-angle CBCT images, CBCT images after artifact correction, and reference CBCT images provided by an embodiment of the present invention;

[0047] Figure 7 The comparison chart of the image processing time between the CBCT artifact correction system based on a lightweight network and the existing CBCT artifact correction system provided by an embodiment of the present invention;

[0048] Figure 8 The structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0049] To better understand the technical solution of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0050] It should be clear that the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0051] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms of "a", "the", and "said" used in the embodiments of the present invention and the claims are also intended to include the plural forms, unless the context clearly indicates otherwise.

[0052] It should be understood that the term " / and / " used herein is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, a and / or b may represent: a exists alone, a and b exist simultaneously, and b exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0053] See Figure 1 , which is the structural diagram of the CBCT artifact correction system based on a lightweight network provided by an embodiment of the present disclosure. As Figure 1 shown, the CBCT artifact correction system based on a lightweight network, the system includes an image preprocessing unit 101, a model training unit 102, and an image generation unit 103.

[0054] The image preprocessing unit 101 is used to reconstruct the projection data collected by CBCT scanning to obtain a sparse-angle CBCT image and a reference CBCT image.

[0055] Specifically: reconstruct the projection data collected by CBCT scanning, use FDK reconstruction for the data of 150 projection angles to obtain a sparse-angle CBCT image, and use FDK reconstruction for the data of 600 projection angles to obtain a reference CBCT image; pair the sparse-angle CBCT image and the reference CBCT image one by one to obtain a CBCT sample pair, and each CBCT sample pair includes a sparse-angle CBCT image and a corresponding reference CBCT image.

[0056] The model training unit 102 is used to train a sparse-angle CBCT artifact correction model based on the sparse-angle CBCT image and the reference CBCT image.

[0057] The image generation unit 103 is used to input the actual sparse-angle CBCT image into the sparse-angle CBCT artifact correction model to obtain a corrected CBCT image.

[0058] The model training unit 102 includes an encoder (compound parallel encoder), a decoder (feature selection-fusion decoder), and a dynamic skip connection.

[0059] The encoder is used to extract multi-dimensional feature information, including longitudinal global features and transverse local features, efficiently capture global context relationships and detail information through a parallel structure, and use a parameter sharing mechanism to reduce computational complexity and improve processing efficiency.

[0060] The encoder includes a longitudinal attention encoder (longitudinal attention Transformer encoder), a transverse attention encoder (transverse attention Transformer encoder), and a downsampling module.

[0061] The longitudinal attention encoder is used to extract the longitudinal features of the input data, capture the long-range dependence relationship of the input data in the longitudinal dimension through the longitudinal self-attention mechanism, and improve the context awareness ability.

[0062] The transverse attention encoder is used to extract the transverse features of the input data, model the dependence relationship in the transverse dimension through the transverse self-attention mechanism, effectively capture the information in the transverse direction, and enhance the breadth and accuracy of the overall feature expression.

[0063] The downsampling module is used to gradually reduce the resolution of the input data, reduce the data calculation amount while retaining key features, and provide a more efficient processing basis for encoding and decoding operations. The downsampling operation of the downsampling module works in conjunction with the attention mechanism, which can effectively aggregate multi-dimensional features and further optimize the expression of global and local information.

[0064] The decoder includes a decoder module (feature selection-fusion decoder module) and an upsampling module. Among them:

[0065] The decoder module is used to accurately correct artifacts by using the joint attention mechanism in the spatial and channel dimensions, improve the detail restoration ability and overall quality of the image, and generate a clear and accurate artifact-corrected CBCT image.

[0066] The upsampling module is used to upsample the input data and gradually restore the complete artifact-corrected CBCT image.

[0067] Based on the above embodiments, this embodiment further specifically discloses the specific implementation manner of the CBCT artifact correction system structure based on the lightweight network.

[0068] See Figure 2 , which is the structural diagram of the model training unit of the CBCT artifact correction system based on the lightweight network provided by the embodiments of the present disclosure. As Figure 2 shown, the model training unit includes an encoder, a decoder, and a dynamic skip connection.

[0069] The encoder includes a first longitudinal attention encoder, a first longitudinal downsampling module, a second longitudinal attention encoder, a second longitudinal downsampling module, a third longitudinal attention encoder, a third longitudinal downsampling module, a first transverse attention encoder, a first transverse downsampling module, a second transverse attention encoder, a second transverse downsampling module, a third transverse attention encoder, a third transverse downsampling module, and a 1×1 convolution.

[0070] The first longitudinal attention encoder, the first longitudinal downsampling module, the second longitudinal attention encoder, the second longitudinal downsampling module, the third longitudinal attention encoder, and the third longitudinal downsampling module are sequentially connected in series to form a longitudinal attention extraction link.

[0071] The first transverse attention encoder, the first transverse downsampling module, the second transverse attention encoder, the second transverse downsampling module, the third transverse attention encoder, and the third transverse downsampling module are sequentially connected in series to form a transverse attention extraction link.

[0072] It should be specifically noted that the horizontal downsampling module and the vertical downsampling module have the same structure. In this embodiment, the downsampling module is divided into a horizontal downsampling module and a vertical downsampling module according to the name of the specific feature of the downsampling of the downsampling module.

[0073] The vertical attention extraction link and the horizontal attention extraction link are arranged in parallel, and the output features of the vertical attention extraction link and the horizontal attention extraction link are fused and then passed through a 1×1 convolution to obtain the output features of the composite parallel encoder.

[0074] The decoder includes a first decoder module, a first upsampling module, a second decoder module, a second upsampling module, a third decoder module, and a third upsampling module connected in sequence.

[0075] The dynamic skip connection is used to efficiently transfer multi-level features between the encoder and the decoder and dynamically adjust the features transmitted from the vertical attention encoder module and the horizontal attention encoder module by means of weighted fusion.

[0076] The dynamic skip connection module weights and fuses the output features of the first vertical attention encoder and the output features of the first horizontal attention encoder through learnable parameters to obtain a first fusion feature, and then skip-inputs the first fusion feature into the third decoder module; weights and fuses the output features of the second vertical attention encoder and the output features of the second horizontal attention encoder through learnable parameters to obtain a second fusion feature, and then skip-inputs the second fusion feature into the second decoder module; weights and fuses the output features of the third vertical attention encoder and the output features of the third horizontal attention encoder through learnable parameters to obtain a third fusion feature, and then skip-inputs the third fusion feature into the first decoder module.

[0077] It should be specifically noted that in this embodiment, the learnable parameters in the dynamic skip connection module are preferably 1×1 convolutions to achieve the lightweight of the network structure.

[0078] The first vertical attention encoder, the second vertical attention encoder, and the third vertical attention encoder have the same structure. Refer to Figure 3 , which is the structural diagram of the vertical attention encoder of the CBCT artifact correction system based on the lightweight network provided by the embodiment of the present invention. As Figure 3As shown in the figure, the vertical attention encoder includes a first layer normalization layer (Layer Normalization, LN), a multi-head vertical attention module, a second layer normalization layer, a multi-layer perceptron, and a residual block connected in series in sequence. The input end of the first layer normalization layer is connected to the output end of the multi-head vertical attention module through a skip connection, so that the input features of the first layer normalization layer and the output features of the multi-head vertical attention module are added residually as the input features of the second layer normalization layer; the input end of the second layer normalization layer is connected to the output end of the multi-layer perceptron through a skip connection, so that the input features of the second layer normalization layer and the output features of the multi-layer perceptron are added residually as the input features of the residual block.

[0079] Further, the multi-head vertical attention module is used to process information in different feature subspaces and enhance the expression ability of global features. The multi-head vertical attention module includes a first 1×1 convolution, a multi-head vertical attention, and a second 1×1 convolution connected in series in sequence. The input end of the first 1×1 convolution is connected to the output end of the second 1×1 convolution through a skip connection, so that the input features of the first 1×1 convolution and the output features of the second 1×1 convolution are added residually as the output features of the multi-head vertical attention module.

[0080] The first 1×1 convolution and the second 1×1 convolution are used to compress and restore channel information.

[0081] The multi-head vertical attention is used to divide the input features into multiple heads, and each head independently calculates the vertical attention to capture the global context information along the vertical direction, and finally splices each head together.

[0082] The residual block includes two 3×3 convolutions and a 1×1 convolution connected in series in sequence. The input end of the first 3×3 convolution is connected to the output end of the 1×1 convolution through a skip connection.

[0083] Corresponding to the vertical attention encoder, the structures of the first horizontal attention encoder, the second horizontal attention encoder, and the third horizontal attention encoder are the same. The horizontal attention encoder includes a first layer normalization layer (Layer Normalization, LN), a multi-head horizontal attention module, a second layer normalization layer, a multi-layer perceptron, and a residual block connected in series in sequence. The input end of the first layer normalization layer is connected to the output end of the multi-head horizontal attention module through a skip connection, so that the input features of the first layer normalization layer and the output features of the multi-head horizontal attention module are added residually as the input features of the second layer normalization layer; the input end of the second layer normalization layer is connected to the output end of the multi-layer perceptron through a skip connection, so that the input features of the second layer normalization layer and the output features of the multi-layer perceptron are added residually as the input features of the residual block.

[0084] The multi-head horizontal attention module includes a first 1×1 convolution, a multi-head horizontal attention, and a second 1×1 convolution connected in series in sequence. The input end of the first 1×1 convolution is connected to the output end of the second 1×1 convolution through a skip connection, so that the input features of the first 1×1 convolution and the output features of the second 1×1 convolution are added residually and used as the output features of the multi-head horizontal attention module.

[0085] The multi-head horizontal attention is used to divide the input features into multiple heads, and each head independently calculates the horizontal attention to capture the horizontal information, and finally splices each head together.

[0086] It should be particularly noted that except for the differences in the multi-head vertical attention and the multi-head horizontal attention, the structures of the parallel vertical attention encoder and the horizontal attention encoder are the same, and they share the same weight parameters to achieve parameter sharing, reducing the redundant calculations and the number of parameters of the network. It not only realizes the lightweight of the network structure, but also improves the efficiency and consistency of multi-dimensional feature extraction.

[0087] The structures of the first decoder module, the second decoder module, and the third decoder module are the same. Refer to Figure 4 , which is the structural diagram of the decoder module of the CBCT artifact correction system based on the lightweight network provided by the embodiment of the present invention. As Figure 4 shown, the decoder module includes a feature selection module, a channel attention module, and a feature fusion module. Among them:

[0088] The feature selection module includes two parallel spatial attention mechanisms. One of the spatial attention mechanisms is the fused feature input end, which is used to connect the fused features coming from the dynamic skip connection module, and the other spatial attention mechanism is the decoding end, which is used to connect the input features. Specifically, the input features of the decoding end of the first decoder module are the output features of the composite parallel encoder, and the input features of the fused feature input end are the third fused features. The input features of the decoding end of the second decoder module are the output features of the first upsampling module, and the input features of the fused feature input end are the second fused features. The input features of the decoding end of the third decoder module are the output features of the second upsampling module, and the input features of the fused feature input end are the first fused features. The output features of the two spatial attention mechanisms are weighted and fused through learnable parameters to be used as the output features of the feature selection module, and the learnable parameters can dynamically adjust the contribution ratio of the output features of the two spatial attention mechanisms in the weighted fusion process.

[0089] The spatial attention mechanism filters out key global and local features by focusing on important regions and suppressing irrelevant information, enhancing the effectiveness of feature representation. The spatial attention mechanism includes a cascaded third 1×1 convolution and a sigmoid function. The input end of the third 1×1 convolution is connected to the output end of the sigmoid function through a skip connection, and the input feature of the third 1×1 convolution is multiplied by the residual of the output feature of the sigmoid function as the output feature of the spatial attention mechanism.

[0090] The third 1×1 convolution is used to compress the channels to 1.

[0091] The channel attention module is used to further optimize the expression of different feature channels, strengthen key channel information, and suppress redundant features. The introduction of the channel attention module enables the feature selection-fusion decoder to better capture the high-dimensional correlation information in the features and improve the quality of the decoding results.

[0092] The channel attention module consists of the addition of a global average pooling layer and a maximum pooling layer in parallel, followed by a cascaded Softmax function. A skip connection is provided between the input end of the channel attention module and the output end of the Softmax function, and the input feature of the channel attention module is multiplied by the residual of the output feature of the Softmax function as the output feature of the channel attention module.

[0093] The global average pooling layer and the maximum pooling layer are respectively used to compress the spatial information of the input feature.

[0094] The feature fusion module is used to perform weighted fusion on the input feature at the decoding end of the feature selection module and the output feature after passing through the channel attention module using learnable parameters, and dynamically adjust the contribution ratio of the two features in the weighted fusion process using learnable parameters.

[0095] It should be noted that in this embodiment, both learnable parameters in the feature selection-fusion decoder module are preferably 1×1 convolutions to achieve the lightweight of the network structure.

[0096] Based on the above embodiments, this embodiment provides a CBCT artifact correction method based on a lightweight network. Refer to Figure 5 , which is a schematic flowchart of the CBCT artifact correction method based on a lightweight network provided by the embodiment of the present invention. As Figure 5 shown, the method includes the following steps:

[0097] Step S1, reconstruct the projection data collected by CBCT scanning to obtain a sparse-angle CBCT image and a reference CBCT image. Specifically:

[0098] Using data of 150 projection angles, the CBCT projection data is reconstructed by the FDK algorithm to generate a sparse-angle CBCT image. Due to the small number of sampling angles, the reconstructed sparse-angle CBCT image has significant artifacts and noise. This sparse-angle image provides samples with artifacts for the artifact correction model, which helps the network learn how to eliminate the image artifact features caused by insufficient sampling during the training process.

[0099] Using data of 600 projection angles, a complete reconstruction is performed by the same FDK algorithm to obtain a high-quality reference CBCT image. Due to sufficient sampling angles, the reference image has fewer artifacts and noise, clear details, and can represent the ideal reconstruction result. This image is used as the control target of the model during training to guide the network to learn how to restore the sparse-angle image to a high-quality output similar to the reference image.

[0100] The image obtained after reconstructing the projection data is a 3D NII-format CBCT image, and the 3D NII image needs to be sliced into 2D images. The actual patient image data obtained is a 3D image, and the network model in this embodiment processes 2D images. Therefore, slicing is required for the network model to perform training and learning.

[0101] Pair the sparse-angle CBCT images with the corresponding reference CBCT images one by one to form sparse-reference sample pairs. Such image pairs are used to train the model, enabling the model to gradually optimize the artifact removal effect through contrastive learning. In each sample pair, the sparse image is input into the model, while the reference image serves as the standard for the artifact correction effect, providing accurate guidance for the model during the learning process.

[0102] Step S2: Based on the sparse-angle CBCT image and the reference CBCT image, a sparse-angle CBCT artifact correction model is trained. Specifically:

[0103] Step S21: Input the sparse-angle CBCT image into the encoder, and enter the longitudinally parallel attention extraction link and the laterally parallel attention extraction link in the encoder respectively. The longitudinally parallel attention extraction link extracts the longitudinal features of the sparse-angle CBCT image, and the laterally parallel attention extraction link extracts the lateral features of the sparse-angle CBCT image. The output features of the longitudinally parallel attention extraction link and the laterally parallel attention extraction link are fused and then passed through a 1×1 convolution to obtain the composite parallel encoder output features.

[0104] During this process, the output features of the first longitudinal attention encoder and the first lateral attention encoder are fused to obtain a first fused feature; the output features of the second longitudinal attention encoder and the second lateral attention encoder are fused to obtain a second fused feature; the output features of the third longitudinal attention encoder and the third lateral attention encoder are fused to obtain a third fused feature; the first fused feature, the second fused feature, and the third fused feature are respectively transmitted to the third decoder module, the second decoder module, and the first decoder module through dynamic skip connections. During the transmission process, the dynamic skip connections perform weighted adjustment on the feature information according to the complexity of the CBCT image artifacts, and perform selective fusion between the global features and the local features through learnable weight parameters to ensure that high-quality feature information can be accurately transmitted to the decoder and reduce the interference of redundant information.

[0105] Step S22: The encoder output feature and the third fused feature are respectively input into the first decoder module, and then upsampled by the first upsampling module to obtain a first high-resolution feature; the first high-resolution feature and the second fused feature are respectively input into the second decoder module, and then upsampled by the second upsampling module to obtain a second high-resolution feature; the second high-resolution feature and the third fused feature are respectively input into the third decoder module, and then upsampled by the third upsampling module to obtain the artifact-corrected CBCT image.

[0106] Step S23: Compare the artifact-corrected CBCT image with the reference CBCT image, calculate the loss function of the two. If the loss function converges, the sparse-angle CBCT artifact correction model is obtained; if the loss function does not converge, update the parameters through the backpropagation algorithm, and repeat steps S21 - S22 until the loss function converges to obtain the sparse-angle CBCT artifact correction model.

[0107] The loss function adopts a combination of mean squared error (MSE) and structural similarity (SSIM), which can not only minimize the error of image reconstruction, but also improve the contrast and detail consistency of the image. The specific formula is as follows:

[0108] L = α·MSE+(1 - α)·(1 - SSIM)

[0109] Among them, L represents the loss function, α represents the weight parameter, and in this embodiment, α = 0.5; MSE represents the mean squared error, and SSIM represents the structural similarity.

[0110] Step S3: Input the actual sparse-angle CBCT image into the sparse-angle CBCT artifact correction model to obtain the corrected CBCT image.

[0111] See Figure 6, which is a comparison diagram of sparse-angle CBCT images and reference CBCT images of the CBCT artifact correction method based on a lightweight network provided by an embodiment of the present invention. As Figure 6 shown, the leftmost figure is a sparse-angle CBCT image, the middle figure is the CBCT image after artifact correction generated by the system and method provided by an embodiment of the present invention, and the rightmost figure is a reference CBCT image. By comparison, it can be clearly seen that the CBCT image after artifact correction generated by the system and method provided by the embodiments of the present disclosure is very similar to the reference CBCT image, which proves the effectiveness of the system and method provided by the embodiments of the present disclosure.

[0112] Figure 7 This is a comparison diagram of the image processing time between the CBCT artifact correction system based on a lightweight network provided by an embodiment of the present invention and the existing CBCT artifact correction system. Among them, figure (a) is the image processing time of the CBCT artifact correction system based on a lightweight network provided by an embodiment of the present invention, and figure (b) is the image processing time of the CBCT artifact correction system in the prior art (the CBCT artifact correction system based on a multi-stage reconstruction network disclosed in Chinese Patent CN119273799A). As Figure 7 shown, a single image with a size of [1, 1, 384, 384] is input into the CBCT artifact correction system of the present disclosure and the existing CBCT artifact correction system respectively. The processing time of the system of this application is 0.0075 seconds, while the single-network processing time of the existing system is 2.229 seconds, and the overall processing time of the three networks of the existing system is 2.229 * 3 = 6.687 seconds. It can be clearly seen that, compared with the prior art, the processing time and efficiency of the system of the embodiments of the present disclosure have been exponentially improved, which proves the lightweight of the system provided by the embodiments of the present disclosure and the high efficiency of image processing brought by the lightweight.

[0113] Corresponding to the above embodiments, an embodiment of the present invention further provides an electronic device.

[0114] See Figure 8 , which is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 8 shown, the electronic device 800 may include: a processor 801, a memory 802, and a communication unit 803. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the electronic device shown in the figure does not constitute a limitation to the embodiments of the present invention. It can be a bus structure, a star structure, and may also include more or fewer components than shown, or combine some components, or different component arrangements.

[0115] Among them, the communication unit 803 is used to establish a communication channel so that the electronic device can communicate with other devices.

[0116] The processor 801 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 802, and by invoking the data stored in the memory, it executes various functions of the electronic device and / or processes data. The processor may be composed of an integrated circuit (IC), for example, it may be composed of a single packaged IC, or it may be composed of multiple packaged ICs with the same or different functions connected together. For example, the processor 801 may only include a central processing unit (CPU). In the embodiment of the present invention, the CPU may be a single arithmetic core or may include multiple arithmetic cores.

[0117] The memory 802 is used to store the execution instructions of the processor 801. The memory 802 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.

[0118] When the execution instructions in the memory 802 are executed by the processor 801, the electronic device 800 is enabled to execute some or all of the steps in the above method embodiments.

[0119] Corresponding to the above embodiments, an embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium may store a program. When the program runs, it can control the device where the computer-readable storage medium is located to execute some or all of the steps in the above method embodiments. Specifically, the computer-readable storage medium may be a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.

[0120] Corresponding to the above embodiments, an embodiment of the present invention further provides a computer program product. The computer program product contains executable instructions. When the executable instructions are executed on a computer, the computer is enabled to execute some or all of the steps in the above method embodiments.

[0121] In the embodiments of the present invention, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent the cases of A existing alone, A and B existing simultaneously, and B existing alone. Here, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one of the following" and its similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0122] Those of ordinary skill in the art can realize that the units and algorithm steps described in the embodiments disclosed herein can be implemented by a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.

[0123] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0124] In several embodiments provided by the present invention, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs, etc., which can store program codes.

[0125] The above is only the specific implementation manner of the present invention. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention and should be covered by the protection scope of the present invention.

Claims

1. A CBCT artifact correction system based on a lightweight network, characterized in that: The system comprises: an image preprocessing unit, a model training unit, and an image generating unit; The model training unit includes an encoder, a decoder, and dynamic skip-layer connections; The encoder includes a longitudinal attention extraction link and a lateral attention extraction link arranged in parallel, the longitudinal attention extraction link includes a plurality of longitudinal attention encoders connected in series, the lateral attention extraction link includes a plurality of lateral attention encoders connected in series, and the longitudinal attention encoder and the lateral attention encoder share the same weight parameter; The decoder includes a number of decoder modules connected in series; The output features of the vertical attention encoder and the horizontal attention encoder at the same level in the vertical attention extraction link and the horizontal attention extraction link are weighted fused through learnable parameters, and then input into the decoder module of the same level through dynamic skip-layer connection. The learnable parameters can dynamically adjust the contribution ratio of the two output features in the weighted fusion process.

2. The lightweight network-based CBCT artifact correction system according to claim 1, characterized in that: The encoder also includes a downsampling module and a 1×1 convolution; The output of each vertical attention encoder and horizontal attention encoder is connected to a downsampling module; The output features of the vertical attention extraction link and the horizontal attention extraction link are fused and input into the 1×1 convolution.

3. The lightweight network-based CBCT artifact correction system according to claim 1, characterized in that: The decoder also includes an upsampling module; The output of each decoder module is connected to an up-sampling module.

4. The lightweight network-based CBCT artifact correction system according to claim 2, characterized in that: The longitudinal attention encoder comprises a first normalization layer, a multi-head longitudinal attention module, a second normalization layer, a multi-layer perceptron and a residual block in series; The input end of the first normalization layer is connected to the output end of the multi-head vertical attention module through a skip layer; the input end of the second normalization layer is connected to the output end of the multi-layer perceptron through a skip layer; The multi-head vertical attention module includes a first 1×1 convolution, a multi-head vertical attention, and a second 1×1 convolution connected in series, and the input end of the first 1×1 convolution is connected to the output end of the second 1×1 convolution through a skip layer.

5. The lightweight network-based CBCT artifact correction system according to claim 2, characterized in that: The decoder module includes a feature selection module, a channel attention module and a feature fusion module.

6. The lightweight network-based CBCT artifact correction system according to claim 5, characterized in that: The feature selection module includes parallel spatial attention mechanisms, and the output features of the two spatial attention mechanisms are weighted and fused through learnable parameters; The spatial attention mechanism includes a third 1×1 convolution and a sigmoid function connected in series, and the input end of the third 1×1 convolution is connected to the output end of the sigmoid function through a skip layer.

7. The lightweight network-based CBCT artifact correction system according to claim 5, characterized in that: The channel attention module is composed of a global average pooling layer and a maximum pooling layer added in parallel, and then a Softmax function in series; The input end of the channel attention module and the output end of the Softmax function are connected via a skip layer.

8. The lightweight network-based CBCT artifact correction system according to claim 5, characterized in that: The feature fusion module performs weighted fusion of the input features of the feature selection module from the previous decoder module / 1×1 convolution and the output features of the channel attention module through learnable parameters.

9. The CBCT artifact correction method based on the lightweight network-based CBCT artifact correction system of claim 1, characterized in that: The method comprises: Reconstructing the projection data to obtain sparse angle images and reference images; Based on the sparse angle images and reference images, a sparse angle CBCT artifact correction model is trained; The actual sparse angle image is input into the sparse angle CBCT artifact correction model to obtain the corrected image.

10. The CBCT artifact correction method according to claim 9, characterized in that: Based on the sparse angle image and the reference image, the sparse angle CBCT artifact correction model is trained, specifically: Step S21, input the sparse angle image into the encoder, extract the longitudinal features by the longitudinal attention extraction link, extract the lateral features by the lateral attention extraction link, and fuse the output features of the two links and obtain the encoder output features through 1×1 convolution; The output features of the vertical attention encoder and the horizontal attention encoder at the same level in the vertical attention extraction link and the horizontal attention extraction link are weighted and fused through learnable parameters, and then input into the decoder module at the same level through dynamic skip-layer connection; Step S22, the encoder outputs features and inputs them to the decoder, and each decoder module of the decoder sequentially fuses and samples the input features with the fusion features of the same level to obtain an artifact-corrected image; Step S23, compare the artifact-corrected image with the reference image, and calculate the loss function of the two. If the loss function converges, the sparse angle CBCT artifact correction model is obtained; if the loss function does not converge, the parameters are updated through the back propagation algorithm, and steps S21-S22 are repeated until the loss function converges to obtain the sparse angle CBCT artifact correction model.

Citation Information

Patent Citations

  • CBCT artifact correction system and method based on multi-stage reconstruction network

    CN119273799A