Feature enhancement and bone structure constraint fused three-dimensional CBCT synthetic CT image model and method

Through the FEB-CycleGAN model combined with ResNet and ViT feature extraction modules, bone structural constraints were introduced, which solved the problem of inaccurate Hounsfield units in CBCT images, and achieved high-quality three-dimensional CBCT to synthetic CT conversion, improving image quality and dose calculation accuracy.

CN120374780APending Publication Date: 2025-07-25HEFEI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510666773.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing CBCT images have problems with low signal-to-noise ratio and inaccurate Hounsfield units caused by x-ray scattering in adaptive radiation therapy, which affects the accuracy of dose calculation. Traditional methods rely on paired data and lack information interaction capabilities in three-dimensional space, resulting in loss or distortion of structural information.

Method used

The FEB-CycleGAN model is adopted, combined with the ResNet and ViT feature extraction modules, and the structural similarity of the bone region is enhanced through the mutual information loss function constraint of the bone structure contour, and the transformation from three-dimensional CBCT to synthetic CT is realized, thereby improving the structural retention performance.

Benefits of technology

The conversion effect of three-dimensional CBCT to synthetic CT is significantly improved, the difference in HU value at the connection between bone and soft tissue is reduced, the image quality and dose calculation accuracy is improved, and the generalization ability of the model in different devices and scenarios is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374780A_ABST
    Figure CN120374780A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional CBCT synthetic CT image model fusing feature enhancement and bone structure constraint and a method, and relates to the technical field of medical information. The model fuses a medical image generator of Resnet and a medical image generator of VIT, the Resnet is used for extracting local features reflecting image structure details, and the ViT is used for capturing global features reflecting image long-range dependency; global and local features are fused through a compression channel, so that the network can pay attention to local details and capture wide context information at the same time; in addition, according to the model FEB-CycleGAN, a bone profile mutual information loss function is introduced to constrain the structural similarity between the original CBCT and the synthetic CBCT, so that the synthetic CT has a better structure retention characteristic. According to the method, the effectiveness and design rationality of the model are comprehensively verified through experiments such as image quality evaluation, ablation, dose comparison and sensitivity analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical information technology, and specifically relates to a three-dimensional CBCT synthetic CT image model and method integrating feature enhancement and bone structure constraint. Background Art

[0002] Image-guided adaptive radiotherapy uses medical imaging to improve the accuracy and precision of dose delivery during radiotherapy. Currently, cone beam computed tomography (CBCT) is commonly used in clinics to provide three-dimensional (3D) imaging for patient setup and adaptive treatment. However, due to phenomena such as low signal-to-noise ratio and high x-ray scattering in CBCT images, the Hounsfield unit (HU) in CBCT images is inaccurate, which seriously affects dose calculation and adaptive planning before each fraction of dose delivery. Currently, the commonly used method in clinics is to align CBCT with the planned CT based on the deformable registration algorithm. However, the accuracy of this registration is limited by scatter artifacts and anatomical differences between the planned CT and daily CBCT, posing a great challenge for accurate alignment. In recent years, generating synthetic CT (sCT) by using the density of CT and the detailed anatomical information obtained from CBCT to provide the latest anatomical structure and calibrated HU values has become a research hotspot in the field of ART.

[0003] Traditional CT image synthesis methods, such as Monte Carlo simulation, CT prior knowledge, histogram matching, and random forest, although showing robustness and theoretical interpretability in specific tasks, still have limitations such as strong dependence on data quality, weak adaptability to complex scenarios, and over-reliance on manual feature extraction. With the large-scale application of deep learning in fields such as image synthesis and image conversion, researchers have also carried out a series of studies on CT image synthesis methods based on deep learning, mainly including CT synthesis methods based on paired data and CT synthesis methods based on unpaired data.

[0004] CT synthesis methods based on paired data usually require paired CBCT and CT to estimate artifacts in CBCT images or learn the CBCT-to-CT mapping relationship. For example, in 2019, Li et al. proposed a 2D U-Net neural network with an encoder-decoder structure to achieve CBCT image synthesis of CT images (Li, Yinghui, et al. "A preliminary study of using a deep convolution neural network to generate synthesized CT images based on CBCT for adaptive radiotherapy of nasopharyngeal carcinoma." Physics in Medicine & Biology 64.14 (2019): 145010.). To improve the accuracy of the HU values of the synthesized sCT, in 2020, Chen et al. proposed a U-Net-based neural network, sCTU-Net, which generates higher-quality CT images by designing a loss function that combines mean absolute error (MAE) and structural similarity (DSSIM) (Chen, Liyuan, et al. "Synthetic CT generation from CBCT images via deep learning." Medical physics 47.3 (2020): 1115-1125.). Although supervised learning-based methods can generate synthetic CTs with relatively high image quality and dose calculation accuracy, the accuracy of such methods depends on a large amount of high-quality paired medical image data. However, in the case of specific diseases or rare situations, high-quality paired data are often difficult to obtain. In addition, medical data are subject to strict privacy protection restrictions, further increasing the difficulty of high-quality CT synthesis.

[0005] To address the above problems, CT synthesis based on unpaired data of CBCT and CT images has become the focus of research, and a series of good synthesis results have been achieved. For example, in 2021, Zhang et al. implemented the conversion from CBCT to CT based on the generative adversarial network (GAN) (Zhang, Yang, et al. "Improving CBCT quality to CT level using deep learning with generative adversarial network." Medical physics 48.6 (2021): 2816-2826.). However, GAN-based models still require correctly aligned paired datasets for training, and some studies have also found that they cannot retain details during the conversion process. In 2022, Rusanov et al. proposed a systematic review to explore the application of deep learning (DL) in improving the quality of CBCT images, focusing on the effectiveness and limitations of using cycleGAN for synthetic CT generation, which pointed out the direction for the development of synthetic CT research (Rusanov, Branimir, et al. "Deep learning methods for enhancing cone-beam CT image quality toward adaptive radiation therapy: A systematic review." Medical Physics 49.9 (2022): 6019-6054.). To address the problems of insufficient soft tissue contrast and detail performance, in 2022, Deng et al. proposed respath-cycleGAN, which solved the information loss caused by network downsampling by designing a residual path in the generator (Deng, Liwei, et al. "Synthetic CT generation based on CBCT using respath-cycleGAN." Medical physics 49.8 (2022): 5317-5329.).To address the problem of potentially introducing incorrect structural information and noise, Deng et al. introduced the Diversity Branch Block (DBB) module into the generator of CycleGAN in 2023 to obtain low-resolution supplementary semantic information and overcome the challenges brought by noise and defects in CBCT images (Deng, Liwei, et al. "Synthetic CT generation from CBCT using double-chain-CycleGAN." Computers in Biology and Medicine 161 (2023): 106889.).

[0006] The above research has wide applications in the field of 2D image processing and also provides new ideas for the processing and generation of 3D images and CT research. In terms of three-dimensional CT synthesis, in 2020, Liu et al. enhanced the network's ability to suppress redundant information in CBCT images in three-dimensional space by adding attention gates to the generator, improving the image quality (Liu, Yingzi, et al. "CBCT-based synthetic CT generation using deep-attention cycleGAN for pancreatic adaptive radiotherapy." Medical physics 47.6 (2020): 2472-2483.). In 2022, Sun et al. replaced 3D convolution with 2.5D image input and proposed a U-Net discriminator network to improve the accuracy of synthesis (Sun, Bin, et al. "Double U-Net CycleGAN for 3D MR to CT image synthesis." International Journal of Computer Assisted Radiology and Surgery 18.1 (2023): 149-156). At the same time, comparative studies on 2D and 3D models in image conversion tasks have shown that by appropriately selecting the size of 3D patches and constructing a cycle consistency loss function by combining SSIM and L1 norm, it can serve as an indirect constraint to effectively improve the performance of the model (Hadzic, Ibrahim, et al. "Optimizing CycleGAN design for CBCT-to-CT translation: insights into 2D vs 3D modeling, patch size, and the need for tailored evaluation metrics." Medical Imaging 2024: Image Processing. Vol. 12926. SPIE, 2024.).In addition, in 2022, Hu also proposed an unsupervised learning method based on a dual attention module, which integrates context information through scale-aware position and channel attention blocks to effectively handle large intensity variations between CT and CBCT images (Hu, Rui, et al. "Unsupervised computed tomography and cone-beam computed tomography image registration using a dual attention network." Quantitative Imaging in Medicine and Surgery 12.7 (2022): 3705.). In 2024, Hu et al. proposed an improved U-Net architecture based on Vision Transformer (ViT), which significantly improved the quality of the generated synthetic CT images (Hu, Yuxin, et al. "Synthetic CT generation based on CBCT using improved vision transformer CycleGAN." Scientific Reports 14.1 (2024): 11455.). It is worth noting that this method only uses the cycle consistency loss function as an indirect constraint.

[0007] Although the current 2.5D method can achieve 3D CT synthesis, this method usually combines multiple 2D slices to generate images with a certain depth perception, lacking the ability to interact information between distant voxels. Therefore, it cannot fully process all-round information in three-dimensional space, resulting in possible deficiencies in the integrity of the three-dimensional structure of the generated images and prone to loss or distortion of structural information. In addition, there is a lack of direct constraint between the input CBCT image and the synthetic CT image in the current method, which cannot guarantee the structural consistency between the two, resulting in a large difference in HU values around the junction of bone structure and soft tissue. Summary of the Invention

[0008] Cone-beam computed tomography (CBCT) is one of the most commonly used 3D imaging modalities in image-guided radiotherapy. However, there are severe artifacts and inaccurate Hounsfield units (HU) in CBCT images, which pose challenges to CBCT-based adaptive dose planning. To address the above challenges, the present invention proposes a novel three-dimensional CBCT synthetic CT framework model (FEB-CycleGAN), which introduces bone structure contour constraints, enhances the bone structure similarity between the synthetic image and the input image, and thus improves the structure retention performance. At the same time, it combines ResNet and ViT feature extraction modules, and through feature fusion and enhancement, effectively improves the overall feature extraction ability, resulting in a significant improvement in the conversion effect from three-dimensional CBCT to synthetic CT (sCT).

[0009] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0010] A three-dimensional CBCT synthetic CT image model (FEB-CycleGAN) that fuses feature enhancement and bone structure constraints. The model FEB-CycleGAN fuses a medical image generator of Resnet and VIT. Among them, Resnet is used to extract local features reflecting the structural details of the image, and ViT aims to capture global features reflecting the long-range dependencies of the image; the global and local features are fused by compressing channels, so that the network can simultaneously focus on local details and capture extensive context information; in addition, the model FEB-CycleGAN introduces a bone contour mutual information loss function to constrain the structural similarity between the original CBCT and the synthetic CBCT, so that the synthetic CT has better structure retention characteristics.

[0011] As a preferred technical solution of the present invention, the model FEB-CycleGAN is trained by the following steps:

[0012] ① Construct an improved CycleGAN generator: This method adopts a dual-generator structure, including a CBCT-to-CT generator (Gct) and a CT-to-CBCT generator (Gcbct); among them, Gct is responsible for converting CBCT images into sCT, and Gcbct is responsible for restoring sCT to CBCT, so as to achieve cyclic consistency learning;

[0013] ②ViResBlock module enhances the generator: The ViResBlock (Vision Transformer Residual Block) is introduced to combine ResNet and ViT (Vision Transformer) for feature extraction. By combining local convolutional feature extraction (ResNet) with global self-attention (ViT), the ViResBlock improves the model's feature learning ability in the bone region, enabling the generated sCT to have a higher ability to retain details.

[0014] ③Optimization of bone region structure preservation: A bone contour extraction module is specifically introduced in the generator. By comparing the bones in CBCT and CT images, the information expression in the bone region is enhanced, background interference is reduced, and the consistency of HU values is improved.

[0015] ④Adversarial learning mechanism: CycleGAN is trained with two discriminators (Dct and Dcbct). Dct is responsible for distinguishing real CT from the sCT generated by Gct, and Dcbct is responsible for distinguishing real CBCT from the CBCT generated by Gcbct. Through adversarial training, it is ensured that the generated sCT not only retains the CBCT structure information but also conforms to the distribution of real CT images.

[0016] ⑤Multi-loss optimization: The cycle consistency loss (Lcycle) is used to ensure that the converted image still retains the original structure, the adversarial loss (LGAN) is used to improve the realism of the image, and the bone loss (Lbone) is used to enhance the ability to retain details in the bone region and optimize the consistency of HU value distribution.

[0017] The present invention also proposes a method for realizing 3D CBCT synthetic CT based on this model, and the steps are as follows:

[0018] ①Train the model FEB-CycleGAN;

[0019] ②Input CBCT image: Input the CBCT image to be converted into the trained FEB-CycleGAN model to generate the corresponding sCT image;

[0020] ③Output the generated sCT image.

[0021] The present invention proposes a three-dimensional CBCT synthetic CT image model (Feature-Enhanced and Bone-Structure-Constrained, FEB-CycleGAN) that fuses feature enhancement and bone structure constraints. Based on 128*128*64 3D CBCT images, it captures more spatial relationships and synthesizes sCT images. A large number of experiments show that the method of the present invention has obvious advantages in terms of bone structure details, mainly manifested in:

[0022] ① Innovative model design: A three-dimensional CBCT synthetic CT image model (FEB-CycleGAN) that combines ResNet and ViT is proposed, integrating the local feature extraction ability of ResNet and the global dependency modeling ability of ViT.

[0023] ② Structural consistency constraint: By introducing the mutual information loss function of bone contours, the similarity between the original CBCT and the synthetic sCT in the bone structure area is enhanced, effectively reducing the HU value difference at the junction of bone and soft tissue, and improving the structural retention and consistency of the images.

[0024] ③ Comprehensive experimental verification: Through image quality assessment, ablation experiments, dose comparison experiments, and sensitivity analysis, etc., the effectiveness and design rationality of the model are comprehensively verified. At the same time, tests based on real hospital datasets prove the generalization ability and stability of this method under different devices and scenarios, laying a foundation for practical clinical applications. Experimental results show that the three-dimensional CBCT synthetic CT method based on FEB-CycleGAN with feature enhancement and bone structure constraints performs excellently in both image quality and dose accuracy evaluations, significantly outperforming the ordinary 3D-CycleGAN method. In terms of image quality metrics such as mean absolute error (MAE), structural similarity (SSIM), and peak signal-to-noise ratio (PSNR), the MAE of the FEB-CycleGAN method is 79.65 ± 19.49 HU, a 6.1% reduction compared to the ordinary 3D-CycleGAN (84.79 ± 20.50 HU); the SSIM reaches 0.87 ± 0.05; the PSNR is 28.86 ± 1.95 dB, a 1.4% increase compared to the ordinary 3D-CycleGAN (28.47 ± 2.06 dB). At the same time, in key dose accuracy metrics such as the gamma index passing rate of the clinical dataset, FEB-CycleGAN also performs outstandingly, fully verifying its clinical application potential in adaptive radiotherapy and providing important support for improving the quality and precision of radiotherapy. Description of the Drawings

[0025] Figure 1 It is a schematic diagram of the training of the FEB-CycleGAN model. This figure shows the training process of the FEB-CycleGAN model, including the feature extraction module of the input data, the constraint mechanism of the loss function, and the overall architecture of the generative adversarial network.

[0026] Figure 2 It is a schematic diagram of the test of the FEB-CycleGAN model.

[0027] Figure 3 It is the network structure of the FEB-CycleGAN generator.

[0028] Figure 4 Schematic diagram of ViResBlock showing (a) ResNet Block and (b) Vit Block.

[0029] Figure 5 Network structure of the FEB-CycleGAN discriminator.

[0030] Figure 6 Result comparison of converting CBCT to sCT by different methods. The figure shows the conversion effects of the original pCT, CBCT, standard CycleGAN, and improved FEB-CycleGAN. Among them, BA, BB, and BC respectively represent the three-dimensional slice comparison results of different datasets, and the red box area highlights the target bone structure. FEB-CycleGAN has better performance in retaining bone contours and restoring detail information.

[0031] Figure 7 Three-dimensional slice display of the BA dataset. The figure successively shows the sagittal, axial, and coronal slice views of the three-dimensional data, used to visually present the comparison effect of the original CBCT and the generated sCT in three-dimensional space.

[0032] Figure 8 Calculation of the per-pixel HU value difference, and the obtained result is clamped to [-900, 900].

[0033] Figure 9 Image and HU value comparative analysis of the 168th layer slice. (a) Schematic diagram of the image slice and the analysis position; (b) Comparative curve of the vertical HU contour lines of CT, CycleGAN, and FEB-CycleGAN; (c) Comparative curve of the horizontal HU contour lines of CT, CycleGAN, and FEB-CycleGAN, and the red dashed box is the enlarged detail area.

[0034] Figure 10 Comparison of bone contour images generated by different methods.

[0035] Figure 11 Model performance indicators with the change in the number of ViT modules.

[0036] Figure 12 Comparison display of (a) CBCT and (b) sCT images on ZJU-RTDS. Specific implementation manner

[0037] The present invention proposes a three-dimensional CBCT synthetic CT image model (FEB-CycleGAN) that integrates feature enhancement and bone structure constraints. The model FEB-CycleGAN integrates a medical image generator of Resnet and ViT. Among them, Resnet is used to extract local features reflecting the structural details of the image, and ViT aims to capture global features reflecting the long-range dependencies of the image; the global and local features are fused by compressing channels, so that the network can simultaneously focus on local details and capture extensive context information; in addition, the model FEB-CycleGAN introduces a bone contour mutual information loss function to constrain the structural similarity between the original CBCT and the synthetic CBCT, so that the synthetic CT has better structure retention characteristics.

[0038] The following further details the model and method of the present invention in conjunction with embodiments and drawings.

[0039] 1 Materials and methods

[0040] 1.1 Dataset

[0041] The present invention uses the dataset of the SynthRAD2023 challenge (Thummerer, Adrian, et al. "SynthRAD2023 Grand Challenge dataset: Generating synthetic CT for radiotherapy." Medical physics 50.7 (2023): 4664-4674.). The dataset of this challenge collected the imaging data of patients receiving radiotherapy in the brain or pelvic region from three Dutch institutions between 2018 and 2022: Radboud University Medical Center (A), Utrecht University Medical Center (B), and University of Groningen Medical Center (C). All preprocessing and postprocessing were applied to 3D images. There were a total of 180 CBCT-CT pairs of the brain available, and a random 80 / 10 / 10 split was used to create the training and validation sets and the test set during training. The training, validation, and test set splits of each institution classified by the center name (A, B, C) are shown in Table 1.

[0042] Table 1 Distribution of training, validation, and test sets

[0043]

[0044] In addition, to further verify the feasibility of the method in practical applications, the present invention also introduced a real dataset from a hospital for testing. This dataset was sourced from the radiotherapy department of the First Affiliated Hospital of Zhejiang University School of Medicine, and consisted of CBCT and CT image pairs of the head and neck region, with a total of 30 pairs of samples. This dataset covered different equipment types and imaging scenarios, providing a solid foundation for evaluating the generalization ability and stability of the model.

[0045] 1.2 Preprocessing

[0046] In the present invention, all data underwent a standardized preprocessing process to ensure the consistency and reliability of model training and evaluation. First, binary masks were generated using thresholding techniques and hole filling algorithms in the ITK image processing toolkit, and the masks were dilated to include the air margin around the patient. In addition, to unify the image resolution, brain images of all anatomical regions were resampled to a voxel spacing of 1×1×1 mm 3 .

[0047] After the initial preprocessing was completed, the data was further processed for use by the model. Specifically, it included:

[0048] (1) Conversion from DICOM to compressed NIfTI (nii.gz);

[0049] (2) Rigorous registration between CT and MR / CBCT;

[0050] (3) Anonymization of brain patient data, including removal of facial features, to protect patient privacy;

[0051] (4) Generation of binary masks through patient contour segmentation;

[0052] (5) Cropping of MR / CBCT, CT images and their corresponding masks to remove irrelevant background regions, thereby further reducing the file size and focusing on the target region.

[0053] 1.3 Method

[0054] 1.3.1 FEB-CycleGAN Neural Network Framework

[0055] As Figure 1 and 2As shown in the figure, the present invention proposes an overall training framework for CBCT synthetic CT based on FEB-CycleGAN. This method aims to generate high-quality sCT images to meet the requirements of clinical applications. The framework mainly consists of two generative adversarial networks (GANs), and each GAN contains a generator and a discriminator. Specifically, the CBCT to CT generator (Gct) is used to convert CBCT-style images into sCT-style images, while the CT to CBCT generator (Gcbct) is used to restore sCT-style images to CBCT-style images. Through this bidirectional mapping method, the model can achieve cycle consistency of styles between the input and output images, while ensuring the high quality and structural consistency of the generated images. In addition, the discriminator Dct is used to distinguish real sCT images from the sCT images generated by Gct, while the discriminator Dcbct is used to distinguish real CBCT images from the CBCT images generated by Gcbct. Through the adversarial learning mechanism, these two discriminators respectively guide Gct and Gcbct to generate more realistic target-style images, thereby further improving the visual quality and detail retention ability of the generated results. To further improve the quality of the generated images and the performance of specific regions (such as the bone region), the model introduces three key loss functions:

[0056] (1) Cycle consistency loss (Lcycle): Ensure that the structure of the image is consistent with the original input after cross-modal cycle conversion, and avoid information distortion.

[0057] (2) Adversarial loss (LGAN): Improve the quality of the generated images through the dynamic game between the generator and the discriminator.

[0058] (3) Bone loss (Lbone): Design weighted constraints for the bone region, optimize the consistency of detail retention and HU value distribution, and reduce information loss in key regions of medical images.

[0059] In the entire data stream, CBCT images are first converted into sCT images through the generator Gct and then further input into the generator Gcbct to be restored to CBCT images. Similarly, after sCT images are converted into CBCT images through the generator Gcbct, they are input into the generator Gct to be restored to sCT images. Through this cyclic mapping method, the framework not only realizes the efficient conversion between the two image styles, but also improves the authenticity of the generated images and the ability to retain structural details through the collaborative optimization of the above-mentioned multiple loss functions. In addition, the framework introduces the ViResBlock module into the generator structure to enhance the feature extraction ability of the network, especially in the processing of complex anatomical structures (such as bone contours). To solve the problem of uneven information expression between the bone region and the background region, the framework designs a specific bone contour extraction module and integrates it into the calculation of the loss function, thus significantly improving the model's ability to refine and generate the bone region.

[0060] 1.3.2 The Generator of FEB-CycleGAN

[0061] Each generator of FEB-CycleGAN consists of three parts: downsampling, feature extraction, and upsampling, as Figure 3 shown.

[0062] (1) Downsampling part: The design of the downsampling module plays a crucial role in 3D image processing, aiming to balance the efficiency and performance of the model by reducing the computational complexity and expanding the receptive field. This module consists of consecutive convolutional operations, and each layer has carefully designed parameter settings to gradually compress the spatial resolution and enhance the feature expression ability. This layer-by-layer progressive structure achieves two core goals in the feature extraction process: on the one hand, by reducing the spatial resolution, it significantly reduces the computational complexity of the model and the GPU memory overhead, providing a basis for subsequent high-dimensional feature processing; on the other hand, combined with large convolutional kernels and stride operations, it significantly expands the receptive field, enabling the model to capture context dependencies within a larger range, thereby enhancing the sensitivity of features to global information. This design not only optimizes the running efficiency of the model but also supports the accurate capture of multi-scale features of 3D images.

[0063] (2) Feature extraction part: It consists of parallel ResNet blocks and ViT branches (ViResBlock). As Figure 4As shown in (a), in the ResNet branch, the present invention designs a residual block (ResNet Block) for constructing a deep neural network. Its core is a three-dimensional convolutional block (conv block), aiming to enhance the feature learning and expression ability of the network. Specifically, the ResNet Block is composed of two identical sub-modules. Each sub-module includes three-dimensional convolution, normalization, and non-linear activation operations. Each convolution operation realizes feature extraction through a specified padding strategy and the number of channels. The normalization layer is used to stabilize the training, and the ReLU activation function enhances the non-linear expression ability of the model. This repetitive structure aims to fully extract local features and improve the richness of feature representation. As Figure 4 As shown in (b), in the ViT branch, the present invention adopts a 3D vision transformer structure to enhance the network's ability to extract global features of three-dimensional images. This module processes the downsampled three-dimensional feature map. First, it divides the feature map into small blocks of a fixed size and maps the number of input channels to a high-dimensional feature space through an embedding layer. Subsequently, a multi-layer multi-head self-attention mechanism is used to model global dependencies, especially long-range spatial information, to enhance the globality and robustness of feature expression. Through this structure, the ViT branch can effectively capture the complex relationships between image patches and extract global features. Finally, the output feature dimension is mapped back to the initial number of channels for fusion. When combining the features of ResNet and ViT, first, the features of the two are concatenated through channels to form a high-dimensional feature map, and then the number of channels is compressed through downsampling to match the target dimension, thereby achieving effective multi-branch feature fusion. The global three-dimensional spatial dependence helps the model better capture the connections between anatomical structures, thereby accurately restoring the complex details and key features in the image. However, since the convolution operation is essentially local, it has limitations in modeling long-range spatial associations, resulting in incomplete structural information and thus limiting its performance. By effectively combining the features of ResNet and ViT, this method fully utilizes their respective advantages, thereby improving the overall performance of feature extraction.

[0064] (3) Upsampling part: It consists of three transposed convolution layers and combines the skip connection operation in downsampling so as to connect with the feature map during downsampling in the upsampling process, thereby better retaining the detail information. This module gradually realizes the restoration of spatial resolution through multi-layer transposed convolution operations and optimizes feature fusion by combining the skip connection mechanism. At each upsampling stage, the transposed convolution layer simultaneously expands the spatial size and adjusts the number of channels to extract higher-level features with stronger semantic expression ability. In addition, the output feature of each layer is connected with the feature of the corresponding downsampling stage through a skip connection, fusing the low-level detail information and high-level semantic information, thereby achieving a balance between detail retention and global consistency in the feature reconstruction process and improving the quality of the final generated result.

[0065] 1.3.3 Discriminator of FEB-CycleGAN

[0066] Each discriminator of FEB-CycleGAN adopts a PatchGAN structure with a receptive field of 58×58×58, which consists of five convolutional layers, as Figure 5 shown. The first few layers achieve feature extraction and dimensionality reduction through larger strides and gradually increasing numbers of channels, and the last few layers further refine the features through smaller strides. Finally, the network outputs a low-resolution discrimination result for evaluating the authenticity of local regions. This structure effectively captures detailed features through local receptive fields while preserving global consistency, adapting to the discrimination task requirements of small image patches.

[0067] 1.3.4 Loss Function

[0068] In the original CycleGAN, adversarial loss, cycle consistency loss, and reconstruction loss are used in the loss function for training. The expression of the total loss function is:

[0069]

[0070] where λ1 and λ2 are both weight coefficients, and their values are 10 and 5 respectively.

[0071] Adversarial loss L GAN , expressed as and is used to constrain the "game" between the generator and the discriminator. As the training progresses, the generated images become more and more similar to the real images, so that the discriminator cannot correctly distinguish them. At the same time, the discriminator adapts to the gradually improving ability of the generator and sets higher standards for the authenticity of the generated images.

[0072]

[0073] where is the complete objective of the generator (Gcbct) and the generator (Gct), is the complete objective of the discriminator (Dcbct) and the discriminator (Dct).

[0074] The images generated by the generator should not only be similar to the target images in style, but also retain the content of the original images as much as possible. To achieve this, the present invention uses cycle consistency loss to constrain the training process of the generator, which is defined as follows:

[0075]

[0076] In addition, if the input image already meets the target style, the output of the generator should be consistent with the input image. Therefore, identity loss is used to constrain the training process of the generator, which is defined as follows:

[0077]

[0078] To better preserve the anatomical structures in CBCT and generate sCT images, the present invention proposes a structural similarity loss based on the mutual information constraint of bone contours. The specific method is as follows:

[0079] First, extract the bone contours of the original CBCT and the synthetic CT. Then, under the constraint of the mutual information loss function, the bone contour of the synthetic CT is more similar to that of the original CBCT. To effectively extract the bone contours of the original CBCT and the synthetic CT, first, for the medical image data, find its cumulative distribution function (CDF):

[0080]

[0081] and set probability thresholds P lower and P upper to determine the upper and lower limits of the bone HU value:

[0082] HU lower = arg min(CDF(h i ) ≥ P lower ) (8)

[0083] HU upper = arg min(CDF(h i ) ≤ P upper ) (9)

[0084] where h i is the only HU value in the medical image data. Here, n j represents the number of times h j appears in the image data, and N is the number of unique HU values. Calculate the upper and lower limits of the HU value. Then, limit the image intensity within the upper and lower bounds according to the window width and window level of the bone HU value, and clip the values outside the range. Set the pixel values greater than a certain threshold in the image after adjusting the window width and window level to 1, and the pixel values less than the threshold to 0, so as to obtain the bone region segmentation results of the two different modality images and effectively retain the bone contour. Assume the original CBCT image is I CBCT , I CT . Through the range of the bone HU value, calculate the window width W and window level L. Limit the intensities of CBCT and sCT within the range [HU lower , HU upper :

[0085]

[0086] Then binarize the image through the set threshold T to extract the bone contour:

[0087]

[0088] Here, B CBCT and B sCT represent the bone contours extracted from the original CBCT and synthetic CT. Subsequently, in order to effectively measure the similarity of the bone contours in CBCT and CT images, the present invention proposes a contour loss calculation method based on mutual information. Assume P(B CBCT , B sCT ) is the joint probability distribution of the bone contours of CBCT and sCT, and P(B CBCT ) and P(B sCT ) are their respective marginal probability distributions. The mutual information (MI) is defined as:

[0089]

[0090] To minimize the difference between the bone contours of CBCT and sCT, the expression of mutual information MI is calculated in the following way, which is actually a numerical change to the mutual information value. The change formula is:

[0091]

[0092] In this way, the mutual information value is non-linearly processed, so that a larger mutual information (high similarity) corresponds to a lower loss, while a lower mutual information (low similarity) corresponds to a higher loss value. This processing method can effectively increase the penalty of the network for regions with lower similarity, while encouraging higher mutual information. Finally, the loss is calculated using mutual information, no longer focusing on the pixel value differences of each pixel at the bottom layer, but on the data distribution characteristics of the bone contours in different patterns, thus paying more attention to the similarity of the overall contours at a higher level. By adding this supervision to the loss function, the complete objective LG can be written as an equation.

[0093]

[0094] 2 Implementation details

[0095] In the present invention, the CycleGAN program used is sourced from the project 3D-CycleGAN-Pytorch-MedImaging-Runninggator publicly released on GitHub. The input and output interfaces of this program have been modified to adapt to the processing requirements of CBCT images and CT images. On this basis, the implementation of FEB-CycleGAN has been developed. In terms of the model architecture, experiments show that under the current memory conditions, the optimal configuration of the number of Resnet and ViT in the model is 9 to 3. The development environment of the model is Python 3.11.0 and PyTorch 2.1.1, and the hardware configuration uses a single A100 GPU (80GB VRAM). During the training process, the Adam optimizer is adopted, the learning rate is set to 0.0002, and the total number of training epochs is 3500. Among them, the learning rate remains unchanged in the first 1750 epochs, and linearly decays to zero in the last 1750 epochs. In the training data processing, the method of randomly extracting overlapping patches is used, with the patch size of 128×128×64 and the overlap between patches of 32×32×32. All data are normalized to [0,255] and further normalized to [-1,1]. In the test data processing, the sliding window technique is adopted to extract overlapping patches, and the prediction results are synthesized into the original image by the method of weighted average, and the generated image range is [0,255]. To ensure fair comparison, the training settings of all comparison models are consistent with those of FEB-CycleGAN, including the number of training epochs, learning rate, etc., and they are run on the same workstation (as shown in Table 2).

[0096] Table 2 Implementation Detail Parameters

[0097]

[0098] 3 Evaluation

[0099] To accurately compare the similarity between the sCT images and CT images generated by different models, quantitative evaluation metrics such as the mean absolute error (MAE), peak signal-to-noise ratio (PSNR), and structural similarity index (SSIM) are introduced. These metrics are defined as follows:

[0100]

[0101] where CT(x,y,z) and sCT(x,y,z) are the pixel values (x,y,z) in the planned CT and sCT respectively, and n i n j n k is the total number of pixels. MAX is the maximum intensity in CT and SCT. μ CT and μ sCT are the average values of the CT and sCT images. σ CTand σ sCT are the standard deviations of CT and sCT images. MAE is the magnitude of the voxel-based Hounsfield unit (HU) difference between the original CT and sCT. PSNR measures whether the predicted sCT intensity is uniformly distributed or sparse. SSIM is used to evaluate the structural similarity between sCT and CT images. Lower MAE values and higher PSNR and SSIM values indicate better sCT image quality.

[0102] 4 Experiments

[0103] The present invention aims to verify the effectiveness and practicality of the improved method by comprehensively evaluating and optimizing the FEB-CycleGAN model. The specific objectives include: (1) evaluating the performance of the generated sCT images in retaining skeletal structures and overall image features through image quality comparison; (2) conducting ablation experiments to analyze the contributions of each module to the generation effect and verify the rationality of the design; (3) performing dose comparison experiments to evaluate the clinical application value of the generated sCT images in dose calculation; (4) conducting sensitivity analysis to explore the robustness of the model to different parameter settings; (5) verifying through hospital datasets to evaluate the generalization ability of the method under different devices and scenarios and ensure its reliability and stability in practical applications.

[0104] 4.1 SCT Image Quality Comparison

[0105] Based on the average results of the test patient cohort (the average number of patients is 12), the present invention comprehensively evaluated the overall quality of the synthetic CT (sCT) images and conducted a detailed comparative analysis between the proposed FEB-CycleGAN model with skeletal constraints and the traditional unsupervised method (CycleGAN). Table 3 details the comparison results of key metrics such as mean absolute error (MAE), structural similarity index (SSIM), and peak signal-to-noise ratio (PSNR) in the test patient cohort. Through analysis, it can be seen that FEB-CycleGAN shows significant advantages in all metrics. Specifically, the mean MAE of FEB-CycleGAN is 79.65 ± 19.49, which is approximately 41.3% lower than that of CBCT (135.74 ± 37.32) and further reduced by approximately 6% compared to CycleGAN (84.79 ± 20.50), indicating that the model has excellent performance in error control for CT generation, which also verifies the effectiveness of the skeletal constraint mechanism. At the same time, in terms of SSIM, the mean value of FEB-CycleGAN reaches 0.87 ± 0.05, which is approximately 17.6% higher than that of CBCT (0.74 ± 0.05), showing its significant advantage in retaining image structure information. Although numerically close to that of CycleGAN (0.87 ± 0.05), FEB-CycleGAN shows its superiority by introducing skeletal structure constraints (such asFigure 6 , Figure 10 As shown in Figure 10 , the details of the bone region are further optimized, making the generated sCT images perform more excellently in the fidelity of key structures. In addition, in terms of the PSNR index, FEB-CycleGAN reached 28.86 ± 1.95, significantly higher than 26.54 ± 1.96 of CBCT and better than 28.47 ± 2.06 of CycleGAN, demonstrating its advantages in reducing noise, enhancing image contrast, and retaining details.

[0106] Through the above analysis, it can be clearly seen that with the support of the bone constraint mechanism, FEB-CycleGAN effectively reduces errors, improves structural similarity, and improves the signal-to-noise ratio of images, thus being superior to the traditional CycleGAN method in generating high-quality sCT images. These results further verify the key role of the bone constraint mechanism and the improved network design in enhancing the quality during the generation of sCT.

[0107] Table 3 Comparison results of SCT image quality indicators

[0108]

[0109] FEB-CycleGAN comprehensively improves the quality of synthetic CT (sCT) images from a three-dimensional perspective, showing significant advantages compared with traditional two-dimensional slice methods. Three-dimensional data processing can comprehensively utilize the spatial information of volume data to ensure the comprehensive presentation of anatomical structure features from multiple perspectives. Figure 6 The coronal slice results of patients in three datasets A, B, and C are shown. By comparing the CBCT images with the sCT images generated by different methods, it can be clearly seen that FEB-CycleGAN has a significant improvement in the sharpness of bone contours, the uniformity of bone density, and soft tissue contrast. In addition, the region of interest (ROI) marked by the red dashed box is further enlarged to show the details, verifying the significant advantages of FEB-CycleGAN in retaining key anatomical structure information.

[0110] To further demonstrate the role of the three-dimensional method in improving image quality, Figure 7Axial, sagittal, and coronal slices of the BA center were selected for detailed analysis, and the differences in the presentation of anatomical structures between 2D and 3D perspectives were compared. In axial slices, the distribution of bones and soft tissues in the cross section can be observed; sagittal slices intuitively reflect the continuity of the bone structure and the midline anatomical characteristics; coronal slices show the distribution of bone density and anatomical hierarchy. Compared with the two-dimensional slice method with a single perspective, three-dimensional slices can present anatomical features more comprehensively and effectively avoid the limitations of two-dimensional perspective. FEB-CycleGAN further enhances the detail retention and overall consistency of key bone areas through three-dimensional data processing, making the generated sCT images superior to traditional methods in terms of bone edge sharpness, uniformity of density distribution, and soft tissue contrast.

[0111] pass Figure 6 and Figure 7 The comparative analysis of the proposed method can verify that the quality of the generated images can be significantly improved in three dimensions. It not only improves the clarity of the bone contour, but also effectively retains the key anatomical information in the CBCT image from both the global and local levels, providing a more reliable basis for the diagnosis and analysis of medical images.

[0112] Figure 8 The difference between basecyclegan and FEB-CycleGAN and pCT images of the above test patients is shown. The darker the color, the greater the deviation between sCT and dpCT images. This shows that in terms of HU, sCT images are closer to pCT images than CBCT images. Sct FEB-CycleGAN The images have fewer or lighter red and blue areas. This shows that FEB-CycleGAN is better than basecycleGAN in generating sCT images, and the generated sCT images are closer to pCT images in structure and hu value.

[0113] To further evaluate the performance of the network, two slices were randomly selected from the test image and the vertical and horizontal contours corresponding to the red lines were drawn to plot the HU values, as shown in Figure 9 As shown. The vertical and horizontal contour lines pass through the brain tissue and skull regions. In both profile lines, the HU values range from -1000 to 3000. The figure shows that in both contour line plots, the blue line corresponding to FEB-CycleGAN is closest to the black line corresponding to pCT. This indicates that our image exhibits the best correction with the least deviation in HU values compared to the pCT image.

[0114] 4.2 CycleGAN for bone structure improvement

[0115] During the testing phase, in addition to the overall image quality, special attention was paid to the retention of the bone region in the sCT image, and its anatomical consistency and the integrity of the bone structure were analyzed in detail. To evaluate the integrity of the bone structure on sCT, Figure 10 A detailed comparison of the bone regions in the sCT image and the pCT image is shown. Especially under the same CBCT input, the binary bone image generated by FEB-CycleGAN is compared with the bone region in the pCT, intuitively showing the differences in bone structure retention among different methods.

[0116] From the comparison results, it can be seen that the binary bone image generated by FEB-CycleGAN is superior to the traditional method in terms of the sharpness, continuity of the bone boundary, and detail retention. Especially in complex bone regions, such as small bone fragments, sutures, and edge transitions, this method shows higher fidelity and integrity, avoiding problems such as blurred, broken bone edges or information loss that may occur in traditional methods. In addition, for important bone structure regions, the anatomical consistency between the image generated by FEB-CycleGAN and the pCT is significantly higher, and the binary bone contour is more consistent with the bone distribution of the real CT.

[0117] 4.3 Ablation experiments

[0118] To comprehensively evaluate the effectiveness of each component in the proposed network, systematic ablation experiments were conducted, mainly evaluating the network performance through quantitative analysis. The average image quality evaluation metrics (MAE, SSIM, and PSNR) of six test patients were selected as the basic criteria. Key components were gradually removed or adjusted, and the changes in image quality under different settings were recorded. The results show that the design of each module has an important impact on the quality of image generation, and the contribution degree of each component to the overall performance was verified in the data analysis (as shown in Table 4). On the BA dataset, the MAE obtained by the basic Cyclegan model was 103.11±22.23, and the SSIM and PSNR were 0.82±0.05 and 27.02±2.35 respectively. After adding the Bone Structure Constraint, the MAE decreased to 99.74±26.19, and the PSNR and SSIM also increased to 27.21±2.41 and 0.82±0.06 respectively, indicating that this module has a preliminary effect on reducing errors and enhancing image details. After further introducing the ViResBlock generator module, the MAE decreased to 97.34±23.98, and the PSNR increased to 27.40±2.39, showing effective suppression of background redundant information and improvement in feature extraction ability. When the two work together (ViResBlock + Bone Structure Constraint), the MAE is the lowest, at 95.95±21.26, and the SSIM and PSNR reach 0.83±0.05 and 27.58±2.16 respectively, verifying the collaborative optimization ability between modules. On the BB dataset, the MAE of the Base model was 74.64±6.35, and it gradually decreased to 71.92±4.74 after adding the ViResBlock and Bone Structure Constraint. The PSNR and SSIM metrics also increased accordingly, reaching 29.76±0.63 and 0.92±0.01 respectively under the combination of ViResBlock-G + Bone Structure Constraint, indicating that these modules can effectively enhance the stability of detail generation and improve the overall image quality. On the BC dataset, the contribution of each component was also significant. The MAE of the basic Cyclegan model was 76.60±14.86, and it decreased to 71.09±16.31 after introducing ViResBlock + Bone Structure Constraint; the PSNR and SSIM increased from 28.74±1.77 and 0.86±0.04 to 29.23±1.95 and 0.88±0.04 respectively.These results indicate that on complex datasets, ViResBlock + Bone Structure Constraint can not only improve the accuracy of detail restoration but also enhance the structural consistency between the generated results and real images.

[0119] From these results, it can be clearly seen that the design of each module is crucial for improving network performance. The ViResBlock module performs outstandingly in key feature extraction and background suppression, while the Bone Structure Constraint module makes significant contributions to detail optimization and texture enhancement. The combination of the two can not only effectively reduce errors but also improve the overall quality and stability of the images. These data supports provide important references for future model optimization.

[0120] Table 4 Quantitative evaluation results (MAE, PSNR, SSIM) of different methods on three datasets. Base: Basic model. Base + BSC: Combination of the basic model and bone structure constraint (BSC). VRB - G: Model using the ViResBlock generator (VRB - G). VRB - G + BSC: Model using the VRB - G generator and combined with BSC.

[0121] BA dataset

[0122]

[0123] BB dataset

[0124]

[0125] BC dataset

[0126]

[0127] 4.4 Dose comparison experiment

[0128] The main purpose of sCT is to serve as the basis for subsequent clinical tasks, especially dose calculation. Therefore, dose calculation provides the most accurate method to verify the effectiveness of sCT generation and its clinical applicability. For this reason, the sCT generated by the method at different dose levels was compared with the real CT.

[0129] Dose accuracy was evaluated by dose recalculation on synthetic CT based on gamma analysis and dose - volume histogram (DVH) parameters. The in - vivo gamma passing rate was calculated according to the criteria of 2% / 2mm and 3% / 3mm, and the dose below 10% of the maximum dose was not calculated for gamma values.

[0130] 4.5 Sensitivity analysis

[0131] In the present invention, a sensitivity analysis of the number of ViT modules in the generator is carried out to evaluate its impact on the performance of the FEB-CycleGAN framework. As the core component for enhancing the global feature extraction ability, the change in the number of ViT modules is directly related to the generation quality and training efficiency of the model. However, the experimental results show that the increase in the number of ViT modules does not always bring a linear improvement in performance. Excessive modules may lead to an increase in model complexity, thus causing problems such as overfitting or optimization difficulties. From Figure 11 It can be seen from the experimental data and the line chart that as the number of ViT modules gradually increases from 2 to 3, the model improves in various performance metrics (MAE, SSIM, and PSNR). Among them, MAE drops to 79.65±19.49, SSIM and PSNR increase to 0.87±0.05 and 28.86±1.95 respectively, which indicates that increasing the number of modules helps to improve the feature extraction ability of the generator, thus enhancing the structural and texture consistency of the generated images. However, when the number of modules further increases to 4 and 6, the performance shows a certain degree of decline. MAE gradually rises back to 86.65±34.89 (4 modules) and 84.76±27.83 (6 modules), PSNR fluctuates less but is still slightly lower than the best result of 3 modules. At the same time, SSIM remains at the level of 0.86±0.07 without showing obvious improvement. This performance decline may be due to the fact that excessive ViT modules introduce additional computational burden and model complexity, making the training process more likely to fall into local optimal solutions. In addition, a larger number of modules may lead to the generator having too strong learning ability and overfitting to the training data, thus affecting the generalization performance of the generated images to unseen data. Combining the experimental results, when the number of modules is 3, an optimal balance is achieved in various indicators, which indicates that in the FEB-CycleGAN framework, an appropriate setting of ViT modules can achieve an ideal trade-off between feature extraction ability and model complexity.

[0132] 4.6 Effect verification of Radiology DataSet (ZJU-RTDS) of the First Affiliated Hospital of Zhejiang University School of Medicine

[0133] In the present invention, in order to verify the effectiveness of the proposed model in the actual clinical scenario, the Radiology DataSet of the First Affiliated Hospital of Zhejiang University School of Medicine (ZJU-RTDS) was used for evaluation (as shown in Table 5). Compared with the publicly available datasets, the dataset provided by the hospital is more in line with the actual clinical scenario, and targeted data collection and preprocessing were carried out for the target task to ensure the high quality and clinical relevance of the data. This enables the model to more comprehensively evaluate the generation ability in different anatomical structures and lesion regions, and verify its actual application performance. It should be noted that the reason for the relatively high evaluation index values is mainly due to the differences in the configurations of the original CBCT and CT, especially in terms of scanning parameters and imaging quality. This difference has a certain impact on the evaluation index, but at the same time reflects the strong robustness of the model when processing actual clinical data. Through this verification, the effectiveness and adaptability of the proposed model in the actual clinical scenario were further confirmed.

[0134] In addition, in order to visually demonstrate the model performance, a comparison of CBCT and synthetic CT (sCT) images was carried out (as Figure 12 ). The comparison results show that the sCT images generated by the proposed model well preserve the anatomical structure information in the CBCT images, especially showing excellent performance in the detail restoration of the bone region. This result visually verifies the effectiveness and performance of the proposed model in the conversion process from CBCT to sCT, especially in terms of anatomical structure preservation and image quality improvement, providing reliable support for clinical practical applications.

[0135] Table 5 Quantitative evaluation results of CBCT images and generated sCT images on ZJU-RTDS

[0136]

[0137] In summary, the present invention developed an improved CycleGAN model, called FEB-CycleGAN, for generating synthetic CT (sCT) images from cone-beam CT (CBCT) images, thereby indirectly realizing the correction of CBCT images. By introducing the bone structure contour constraint, the bone structure similarity between the synthetic image and the input image was enhanced, thus improving the structure preservation performance. At the same time, the ResNet and ViT feature extraction modules were combined, and through feature fusion and enhancement, the overall feature extraction ability was effectively improved, resulting in a significant improvement in the conversion effect from three-dimensional CBCT to synthetic CT (sCT).

[0138] The method of the present invention was verified on the SynthRAD2023 Grand Challenge dataset. The experimental results show that FEB-CycleGAN has significant advantages in enhancing the ability to capture all-round information and retaining details in the bone region. Compared with the traditional CycleGAN method, the sCT images generated by FEB-CycleGAN are closer to the original CBCT images in terms of bone region correction and detail restoration. At the same time, in terms of clinical effects: in key dose accuracy indicators such as the gamma index passing rate, FEB-CycleGAN also performs excellently, fully verifying its clinical application potential in adaptive radiotherapy and providing important support for improving the quality and accuracy of radiotherapy.

[0139] The above content is only an example and illustration of the concept of the present invention. Those skilled in the art of the present technology can make various modifications or supplements to the described specific embodiments or use similar methods for substitution, as long as they do not deviate from the concept of the invention or exceed the scope defined by the claims of the present invention, they should all fall within the protection scope of the present invention.

Claims

1. A three-dimensional CBCT synthetic CT image model integrating feature enhancement and bone structure constraint, characterized in that The model FEB-CycleGAN is a medical image generator that integrates Resnet and ViT. Among them, Resnet is used to extract local features that reflect the structural details of the image, and ViT aims to capture global features that reflect the long-range dependencies of the image; the global and local features are fused by compressing channels, enabling the network to simultaneously focus on local details and capture extensive context information; in addition, the model FEB-CycleGAN introduces a bone contour mutual information loss function to constrain the structural similarity between the original CBCT and the synthesized CBCT, so that the synthesized CT has better structure retention characteristics.

2. The three-dimensional CBCT synthetic CT image model integrating feature enhancement and bone structure constraint according to claim 1, characterized in that The model FEB-CycleGAN is trained using the following steps: ①Construct an improved CycleGAN generator: This method adopts a dual-generator structure, including a CBCT-to-CT generator (Gct) and a CT-to-CBCT generator (Gcbct); among them, Gct is responsible for converting CBCT images into sCT, and Gcbct is responsible for restoring sCT to CBCT, thus realizing cyclic consistency learning; ②Enhance the generator with the ViResBlock module: The ViResBlock (Vision Transformer ResidualBlock) is introduced, combining ResNet and ViT (Vision Transformer) for feature extraction; the ViResBlock improves the model's feature learning ability in the bone region through the combination of local convolutional feature extraction (ResNet) and global self-attention (ViT), enabling the generated sCT to have higher detail retention ability; ③Optimize the structural preservation in the bone region: A bone contour extraction module is specifically introduced in the generator. By comparing the bones in the CBCT and CT images, the information expression in the bone region is enhanced, background interference is reduced, and the HU value consistency is improved; ④Adversarial learning mechanism: CycleGAN is trained using dual discriminators (Dct and Dcbct). Among them, Dct is responsible for distinguishing real CT from the sCT generated by Gct, and Dcbct is responsible for distinguishing real CBCT from the CBCT generated by Gcbct; through adversarial training, it is ensured that the generated sCT not only retains the CBCT structure information but also conforms to the distribution of real CT images; ⑤Optimize multiple losses: The cyclic consistency loss (Lcycle) is used to ensure that the converted image still retains the original structure, the adversarial loss (LGAN) is used to improve the realism of the image, and the bone loss (Lbone) is used to strengthen the detail retention ability in the bone region and optimize the consistency of the HU value distribution.

3. A method for implementing three-dimensional CBCT synthetic CT based on the model described in claim 1, characterized in that, The steps are as follows: ①Train the model FEB-CycleGAN; ②Input the CBCT image: Input the CBCT image to be converted into the trained FEB-CycleGAN model to generate the corresponding sCT image; ③Output the generated sCT image.

Citation Information

Cited By

  • CBCT high-quality CT image synthesis method based on structure prior guidance

    CN121120833A