CBCT image enhancement method and system based on perception enhancement cyclic generative adversarial network

Through the improved U-Net structure and multi-perceptual loss function, combined with dual contrast loss, a perceptual enhancement cycle generation adversarial network is built to enhance CBCT images, which solves the problem of insufficient CBCT image quality and realizes high-quality image conversion and dose calculation.

CN120047330APending Publication Date: 2025-05-27TCM INTEGRATED HOSPITAL OF SOUTHERN MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510057605.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing CBCT image quality is insufficient, especially in soft tissue contrast and CT value accuracy, which is difficult to directly use for high-precision dose calculations. The traditional CycleGAN network architecture is simple, with a single loss function, making it difficult to ensure multi-dimensional constraints on image quality.

Method used

The method of generating adversarial network based on perceptual enhancement loop is adopted. By improving the U-Net structure, integrating residual connections, and introducing multi-perceptual loss functions and dual contrast loss, a perceptual enhancement loop generation adversarial network is built to enhance CBCT images.

Benefits of technology

It significantly improves the model performance, enhances feature extraction ability and contrast learning ability, and achieves all-round quality constraints. The generated sCT images are better than existing methods in indicators such as MAE, PSNR, and SSIM, meeting clinical needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047330A_ABST
    Figure CN120047330A_ABST
Patent Text Reader

Abstract

The invention discloses a CBCT image enhancement method and system based on a perception enhancement cyclic generative adversarial network, and relates to the technical field of medical image processing, and the method comprises the steps: employing a U-Net-based improved structure in a generator based on a CycleGAN network, integrating residual connection, and introducing a multi-perception loss function, negative samples are introduced into a discriminator to form dual contrast loss, so that contrast learning is realized, and a perception-enhanced cyclic generative adversarial network is constructed; enhancing the CBCT image based on the perception enhancement cyclic generative adversarial network; by quantifying the average absolute error, the peak signal-to-noise ratio and the structural similarity index of the sCT and the reference CT, the dose gamma passing rate of the sCT and the reference CT and the relative dose deviation of a target region and an organ at risk, the sCT image generated by the method is superior to the existing method in the aspects of image quality and dose calculation precision; and a more accurate dose calculation basis can be provided for self-adaptive radiotherapy of nasopharyngeal carcinoma.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and in particular, to a CBCT image enhancement method and system based on a perception-enhanced cycle generative adversarial network. Background Art

[0002] Precision radiotherapy is the main treatment method for nasopharyngeal carcinoma, but anatomical structure changes during the treatment process may cause the dose distribution to deviate from the original plan. Adaptive Radiotherapy (ART) solves this problem by dynamically adjusting the plan, but traditional ART relies on repeated Computed Tomography (CT) scans, which increases the radiation dose and treatment time of patients.

[0003] To overcome this difficulty, researchers have tried to use the cone-beam CT (CBCT) built into radiotherapy equipment for ART. However, the image quality of CBCT has obvious deficiencies, especially in terms of soft tissue contrast and CT value accuracy, and it is difficult to be directly used for high-precision dose calculation. To solve this problem, researchers have proposed various methods, such as CBCT image post-processing correction and registration-based synthetic CT (sCT). However, these methods have limited accuracy when dealing with complex anatomical structures and large deformations, and the calculation efficiency is low, making it difficult to meet the real-time needs of the clinic.

[0004] In recent years, deep learning technology has developed rapidly in the field of medical image processing. Among them, the method based on the Cycle Generative Adversarial Network (CycleGAN) can achieve the conversion from CBCT to CT without paired data, providing a new idea for solving the problem of insufficient paired training data. Some researchers have tried to improve the quality of the generated images by adding gradient consistency loss, shape consistency loss, etc. to CycleGAN. However, these improved methods still have the following deficiencies: (1) The network architecture is relatively simple, and there are limitations in maintaining fine texture details, suppressing noise and artifacts in medical images; (2) The loss function design is single, lacking multi-dimensional constraints on the quality of the generated images, and it is difficult to ensure the authenticity and accuracy of the converted images; (3) The ability to capture and reconstruct complex anatomical structures is insufficient, especially in cases where the parts are complex and the deformations are obvious, such as nasopharyngeal carcinoma.

[0005] Therefore, there is an urgent need to develop a method that can effectively improve the quality of CBCT images, maintain the accuracy of anatomical structures, and achieve stable and reliable conversion to meet the clinical needs of nasopharyngeal carcinoma adaptive radiotherapy. In particular, new technological breakthroughs are needed in maintaining image details, improving soft tissue contrast, and ensuring CT value accuracy. Summary of the Invention

[0006] To solve the above problems, the objective of the present invention is to provide an image generation method that can effectively improve the quality of CBCT images and maintain the accuracy of anatomical structures, aiming to overcome the problems such as insufficient quality of existing CBCT images, simple architecture of traditional CycleGAN networks, and single loss function.

[0007] To achieve the above technical objectives, the present application provides a CBCT image enhancement method based on a perception-enhanced cycle generative adversarial network, including the following steps: Based on the CycleGAN network, in the generator, an improved U-Net structure is adopted, residual connections are incorporated, and a multi-perception loss function is introduced. In the discriminator, negative samples are introduced to form a dual contrast loss, and a perception-enhanced cycle generative adversarial network is constructed. Based on the perception-enhanced cycle generative adversarial network, the CBCT image is enhanced.

[0008] Preferably, in the process of improving the generator, an improved U-Net structure is adopted. The generator includes an encoder, a decoder, and cross-layer connections. The encoder gradually extracts multi-scale semantic features through downsampling. The decoder restores image details through upsampling. The cross-layer connections enable the feature maps of the encoder to be directly transmitted to the corresponding layers of the decoder. At the same time, ResNet-style residual blocks are introduced. And by fusing the structural similarity index SSIM, mean absolute error MAE, and cross-entropy CE loss as the multi-perception loss function, the quality of the generated image is constrained from multiple dimensions.

[0009] Preferably, in the process of improving the generator, the generator is composed of an input layer, an encoder, a converter, and a decoder, where the input layer adopts a convolutional layer with a 7×7 convolutional kernel; the encoder includes two downsampling layers with a stride of 2. Each layer uses a 3×3 convolution with stride = 2 to halve the size of the feature map, followed by an instance normalization layer and a ReLU activation function; the converter includes 9 residual blocks; the decoder includes two transposed convolutional layers for 2-fold upsampling, and finally generates the target image through a 7×7 convolutional layer followed by a Tanh activation function.

[0010] Preferably, in the process of improving the discriminator, the discriminator adopts an improved PatchGAN structure; the first layer: a 4×4 convolutional layer with a stride of 2; the middle layer: 4 convolutional blocks, each block including a convolutional layer, an instance normalization layer, and a LeakyReLU activation function; the output layer: a 1×1 convolutional layer → Sigmoid, outputting a discriminant probability map.

[0011] Preferably, in the process of introducing the dual contrast loss, according to the adversarial loss and the source-target domain distribution difference loss, as the dual contrast loss, wherein the adversarial loss is used to distinguish the difference between the generated image and the real CT image; The source-target domain distribution difference loss is used to enhance the ability to distinguish CT and CBCT features.

[0012] Preferably, in the process of introducing the dual contrast loss, the least squares loss function is used as the adversarial loss, and the L2 distance metric is used as the source-target domain distribution difference loss.

[0013] Preferably, in the process of enhancing the CBCT image, the CBCT image and the CT image are obtained, and the CBCT is registered to the CT to construct an initial data set. By preprocessing the initial data set, a data set for training the perceptual enhancement cycle generative adversarial network is generated, wherein the preprocessing is performed through image unification, normalization, image enhancement, and data augmentation.

[0014] Preferably, in the process of training the perceptual enhancement cycle generative adversarial network, the training strategy is optimized, including: Using the Adam optimizer with an initial learning rate of 0.0002; Training for 200 epochs with a batch size of 8; Evaluating the model performance on the validation set every 20 training epochs; When there is no obvious improvement in performance for 5 consecutive epochs, stop training; During the training process, the update ratio of the generator and the discriminator is 1:1; Optimizing with the grid search method.

[0015] Preferably, in the process of training the perceptual enhancement cycle generative adversarial network, an early stopping mechanism is set to avoid overfitting.

[0016] The present invention also discloses a CBCT image enhancement system based on the perceptual enhancement cycle generative adversarial network, which is used to implement the CBCT image enhancement method based on the perceptual enhancement cycle generative adversarial network mentioned above, including: A network construction module, which is used to construct a perceptual enhancement cycle generative adversarial network based on the CycleGAN network, where the generator improves the U-Net structure and incorporates residual connections and introduces a multi-perceptual loss function, and the discriminator introduces negative samples to form a dual contrast loss; An image enhancement module, which is used to enhance the CBCT image based on the perceptual enhancement cycle generative adversarial network.

[0017] The present invention discloses the following technical effects: 1. The improved network structure design and dual contrast loss mechanism significantly enhance the model performance: (1) The U-Net residual structure and multi-scale feature fusion of the generator enhance the feature extraction ability; (2) The dual contrast loss of the discriminator improves the model's ability to identify features in different image domains and its contrast learning ability; (3) The training process is more stable, avoiding the mode collapse problem.

[0018] 2. The introduction of the multi-perceptual loss function realizes all-round quality constraints: (1) The SSIM loss ensures the accurate preservation of structural information; (2) The MAE loss provides pixel-level precise supervision; (3) The CE loss enhances semantic consistency; (4) The synergistic effect of each loss component significantly improves the quality of the generated images.

[0019] 3. The optimized training strategy improves the practicality of the model: (1) The convergence speed is increased by more than 30%; (2) The generated sCT images are superior to existing methods in terms of indicators such as MAE, PSNR, and SSIM; (3) In the dose calculation verification, more than 97% of the voxels pass the 3mm / 3% gamma analysis, meeting the clinical requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0021] Figure 1 It is a schematic flowchart of the method described in the present invention; Figure 2 It is a schematic diagram of the structure of the perception-enhanced cyclic generative adversarial network described in the present invention; Figure 3 It is a schematic diagram of the generator architecture described in the present invention; Figure 4 It is a schematic diagram of the discriminator architecture described in the present invention; Figure 5 It is a comparison diagram of CBCT, PE-CycleGAN synthesized sCT, and reference CT images of nasopharyngeal carcinoma patients described in the present invention; Figure 6The original reference CT image of the nasopharyngeal carcinoma patient described in the present invention and the LineProfile and HU Difference of the PE-CycleGAN synthesized sCT; Figure 7 It is the dose distribution of the patient described in the present invention on the CT and PE-CycleGAN synthesized sCT images. Among them, the leftmost one in the upper row is the planned CT dose distribution, the second one from the left in the upper row is the dose-volume histogram of the planned CT and sCT, the leftmost one in the lower row is the sCT dose distribution, and the second one from the left in the lower row is the difference between the planned CT and sCT dose distributions; Figure 8 It is the percentage deviation heat map of the sCT and CT dose parameters of the test patient described in the present invention. Detailed implementation manners

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Usually, the components of the embodiments of the present application described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but merely represents the selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.

[0023] As Figure 1-8 shown, the CBCT image enhancement method based on the perception-enhanced cycle generative adversarial network provided by the present invention includes the following steps: (1) Preprocess the input CBCT and CT images, including size unification, normalization, and data augmentation; (2) Construct a generator network based on the improved U-Net, including an encoder, a converter, and a decoder; (3) Construct a discriminator network based on PatchGAN and introduce a double contrast loss; (4) Fuse a multi-perception loss function in the generator; (5) Adopt an optimized training strategy for model training; (6) Evaluate the performance of the trained model, including evaluating the quality of the sCT image through indicators such as MAE, PSNR, and SSIM.

[0024] Specifically, image preprocessing (1) Image size standardization: - Uniformly adjust the CBCT and CT images to a size of 256×256; - After normalization, the image data is distributed within the range of [0, 1]; - Extract the MASK for the image to focus on the key areas; - Use data augmentation methods to expand the training samples.

[0025] Generator network structure design: (1) Overall network structure: - Design a generator network based on the improved U-Net; - The input layer uses a convolutional layer with a 7×7 convolutional kernel, 1→64 channels; - The encoder contains two downsampling layers with a stride of 2. Each layer uses a 3×3 convolution with stride = 2 to halve the size of the feature map, followed by an instance normalization layer and a ReLU activation function; - The transformer contains 9 residual blocks, and the structure of each residual block is: - Conv(3×3) → InstanceNorm → ReLU → Conv(3×3) → InstanceNorm + Identity mapping → ReLU; - The decoder contains two transposed convolutional layers for 2x upsampling, and finally generates the target image through a 7×7 convolutional layer followed by a Tanh activation function.

[0026] - Add skip connections between the corresponding layers of the encoder and decoder. Connect the 128-channel and 256-channel feature maps of the encoder to the corresponding 128-channel and 256-channel layers of the decoder respectively, and achieve multi-scale feature fusion through the concatenation operation.

[0027] - Input layer: 1→64 channels - Two layers of encoder: 64→128→256 channels - 9 residual blocks of the transformer: Keep 256 channels unchanged - Two layers of decoder: 256→128→64 channels - Output layer: 64→1 channel.

[0028] Discriminator network structure design: (1) Discriminator network structure: - Adopt an improved PatchGAN structure; - First layer: 4×4 convolutional layer, stride of 2; - Middle layer: 4 convolutional blocks, each block contains a convolutional layer, an instance normalization layer, and a LeakyReLU activation function; - Output layer: 1×1 convolutional layer → Sigmoid, output the discriminant probability map.

[0029] Introduction to Loss Functions: (1) Dual Contrast Loss Function: L_D = L_adv(D_CT, CT, sCT) + λ_DC·L_DC(D_CT, CT, CBCT) Where: - L_D represents the total dual contrast loss, which is used to optimize the overall performance of the discriminator; - L_adv represents the adversarial loss, which is used to distinguish the difference between the generated image and the real CT image; - L_DC is the source - target domain distribution difference loss, which is used to enhance the ability to distinguish CT and CBCT features; - D_CT represents the CT discriminator; it evaluates the consistency between the generated image and the real CT image; - CT represents the real CT image, which serves as a reference standard; - sCT represents the generated pseudo - CT image, that is, the target image generated by the model; - CBCT represents the cone - beam CT image, which serves as the source - domain image; - λ_DC represents the weight coefficient of the source - target domain distribution difference loss (L_DC), which is used to balance the relative importance of the adversarial loss (L_adv) and the distribution difference loss (L_DC), and its value range is [0.5, 2].

[0030] (2) Design of the Generator's Multi - perception Loss Function: L_G = L_adv(D_CT, sCT) + λ_rec·L_rec(G_CT, G_CBCT) + λ_SSIM·L_SSIM(sCT, CT) + λ_MAE·L_MAE(sCT, CT) + λ_CE·L_CE(sCT, CT).

[0031] Where: - L_G represents the overall loss of the generator, comprehensively measuring the quality of the generated image; - L_adv represents the adversarial loss, which is used to distinguish the difference between the generated image and the real CT image; - L_rec represents the cycle - consistency loss, ensuring the reversibility of image conversion; - L_SSIM represents the structural similarity loss, which measures image similarity from three dimensions: brightness, contrast, and structure; - L_MAE represents the mean absolute error loss, which is used for pixel - level supervision; - L_CE represents the cross - entropy loss, enhancing semantic consistency; - G_CT represents the generator in the CT domain, which converts the CBCT image into a CT image; - G_CBCT represents the generator in the CBCT domain, which converts CT images into CBCT images; - λ_rec represents the weight of the cycle consistency loss, ensuring the reversibility of the conversion, with a value of 10.0; - λ_SSIM represents the weight of the structural similarity loss, ensuring the preservation of the image structure features, with a value of 0.5; - λ_MAE represents the weight of the mean absolute error loss, controlling the accuracy at the pixel level, with a value of 1.5; - λ_CE represents the weight of the cross-entropy loss, enhancing the semantic consistency, with a value of 1.0; (3) Definitions of each loss term: ① L_rec = ||G_CT(G_CBCT(x)) - x|| 1 + ||G_CBCT(G_CT(y)) - y|| 1 Among them: - L_rec represents the cycle consistency loss, ensuring the reversibility of the image conversion; - x represents the input CBCT image, serving as the source domain image; - y represents the input CT image, serving as the target domain image; - ||·|| 1 represents the L1 norm (Manhattan distance), calculating the sum of the absolute values of the image differences; - G_CT(G_CBCT(x)) represents converting the CBCT image to the CT domain and then back to the CBCT domain; - G_CBCT(G_CT(y)) represents converting the CT image to the CBCT domain and then back to the CT domain; ② L_adv(D_CT, CT, sCT) = E[log(D_CT(CT))] + E[log(1 - D_CT(sCT))] Among them: - E[·] represents the expected value; - D_CT(CT) represents the discrimination result of the discriminator for the real CT image; - D_CT(sCT) represents the discrimination result of the discriminator for the generated sCT image; ③ L_DC(D_CT, CT, CBCT) = E[||f_CT(CT) - f_CT(CBCT)|| 2 ] Among them: - f_CT represents the feature extraction layer of the CT discriminator; - ||·|| 2 represents the L2 norm (Euclidean distance); ④L_MAE(sCT, CT) = 1 / N·∑|sCT_i - CT_i| Among them: - N represents the total number of image pixels; - i represents the pixel position index; - sCT_i represents the value of the i-th pixel of the generated image; - CT_i represents the value of the i-th pixel of the real CT image; ⑤L_CE(sCT, CT) = -∑(CT_i·log(sCT_i)) Among them: - i represents the pixel position index; - CT_i represents the probability distribution of the real CT image; - sCT_i represents the probability distribution of the generated image; ⑥L_SSIM = 1 - SSIM(sCT, CT) - SSIM(x, y) = [l(x, y)]^α·[c(x, y)]^β·[s(x, y)]^γ Among them: - l(x, y) represents the luminance similarity comparison function, which evaluates the matching degree of image luminance; - c(x, y) represents the contrast similarity comparison function, which evaluates the matching degree of image contrast; - s(x, y) represents the structure similarity comparison function, which evaluates the matching degree of image structure; - α represents the weight coefficient of luminance comparison, which adjusts the importance of luminance similarity; - β represents the weight coefficient of contrast comparison, which adjusts the importance of contrast similarity; - γ represents the weight coefficient of structure comparison, which adjusts the importance of structure similarity; - ^ represents the power operation symbol, which is used to adjust the contribution degree of each similarity index.

[0032] - l(x, y) = (2μx·μy + C 1 ) / (μx² + μy² + C 1 ) - c(x, y) = (2σx·σy + C 2 ) / (σx² + σy² + C 2 ) - s(x, y) = (σxy + C 3 ) / (σx·σy + C 3 ) Among them: - μx and μy represent the means of images x and y; - σx and σy represent the standard deviations of images x and y; - σxy represents the covariance of images x and y; - C 1 、C 2 、C 3 are small constants to prevent the denominator from being zero.

[0033] Training strategy optimization: (1) Optimizer configuration: - Use the Adam optimizer with an initial learning rate set to 0.0002; - Train for 200 epochs with a batch size of 8; - Evaluate the model performance on the validation set every 20 training epochs; - Stop training when there is no obvious improvement in performance for 5 consecutive epochs; - The update ratio of the generator and discriminator during training is 1:1; - Tuning by grid search method; - Set an early stopping mechanism to avoid overfitting.

[0034] Example: The present invention provides a method for generating sCT from CBCT based on a perception-enhanced recurrent generative adversarial network for adaptive radiotherapy of nasopharyngeal carcinoma. The specific implementation steps are as follows: 1. Data collection and preprocessing: 1.1 Data collection: Retrospectively include 87 nasopharyngeal carcinoma patients who received intensity-modulated radiotherapy in the hospital from 2020 to 2023, and the data is desensitized. The inclusion criteria are: pathologically confirmed, aged 18 - 70 years old, Karnofsky performance status score ≥ 70 points, enhanced CT localization and first CBCT scan before treatment, and signed informed consent. The exclusion criteria are: recurrent or metastatic tumors, concurrent radiotherapy and chemotherapy, severe side effects, missing data, etc.

[0035] 1.2 Image acquisition: All patients underwent SIMENS CT (120 kV, 200 mA, 3 mm slice thickness, 0.97 mm pixel pitch, 512×512 matrix) and CBCT scans (120 kV, 20 mA, 25 cm field of view, 3 mm slice thickness, 0.97 mm pixel pitch, 512×512 matrix). The CBCT was registered to the CT, and 80 patient cases with a total of 6510 pairs of CT-CBCT slices were selected to form the training set, and 7 cases with a total of 680 pairs of slices were used as the test set.

[0036] 1.3 Data preprocessing: - Image normalization: Adjust the image size to 256×256; - Normalization processing: Normalize the pixel values to the range [0, 1]; - Extract MASK from the image to focus on the key areas; - Data augmentation: Use methods such as random flipping and rotation to expand the dataset.

[0037] 2. Network structure implementation: 2.1 Generator module implementation: (1)Input layer configuration: - Convolution kernel size: 7×7; - Stride: 1; - Padding method: 0-value padding; - Number of output channels: 64; - Activation function: ReLU.

[0038] (2)Encoder configuration: - The first downsampling layer: Conv(3×3,stride=2)→InstanceNorm→ReLU; - The second downsampling layer: Conv(3×3,stride=2)→InstanceNorm→ReLU.

[0039] (3)Transformer configuration: - Number of residual blocks: 9; - Residual block structure: Conv(3×3)→InstanceNorm→ReLU→Conv(3×3)→InstanceNorm + Identity mapping → ReLU.

[0040] (4)Decoder configuration: - The first upsampling layer: TransConv(3×3,stride=2)→InstanceNorm→ReLU; - The second upsampling layer: TransConv(3×3,stride=2)→InstanceNorm→ReLU; - Output layer: Conv(7×7,stride=1)→Tanh.

[0041] 2.2 Discriminator module implementation: (1)Network structure configuration: - Input layer: Conv(4×4,stride=2,channels=64)→LeakyReLU(0.2); - Intermediate layer: 4 convolutional blocks, each block contains a convolutional layer, an instance normalization layer and LeakyReLU; - Output layer: Conv(1×1) → Sigmoid; 3. Loss function implementation: 3.1 Discriminator dual contrast loss: L_D = L_adv(D_CT, CT, sCT) + λ_DC·L_DC(D_CT, CT, CBCT); Implementation details: - The adversarial loss uses the least squares loss function; - The source-target domain distribution difference loss uses the L2 distance metric; - λ_DC is set to 1.0.

[0042] 3.2 Generator multi-perception loss: L_G = L_adv(D_CT, sCT) + λ_rec·L_rec + λ_SSIM·L_SSIM + λ_MAE·L_MAE + λ_CE·L_CE Implementation details: - λ_rec = 10.0: Cycle consistency loss weight; - λ_SSIM = 0.5: Structural similarity loss weight; - λ_MAE = 1.5: Mean absolute error loss weight; - λ_CE = 1.0: Cross-entropy loss weight.

[0043] 4. Training process implementation: 4.1 Training environment configuration: - Hardware platform: RTX 4090 GPU; - Deep learning framework: Tensorflow 2.5.0; - CUDA version: 11.2.

[0044] 4.2 Training parameter settings: - Batch size: 8; - Initial learning rate: 0.0002; - Total number of training epochs: 200; - Adam optimizer parameters: β1 = 0.9, β2 = 0.999.

[0045] 4.3 Training process monitoring: - Evaluate performance on the validation set every 20 epochs; - Stop training if there is no improvement for 5 consecutive epochs; - Save the model with the best validation performance.

[0046] 5. Experimental result verification: 5.1 Image quality assessment: Evaluate the quality of sCT images using metrics such as MAE, PSNR, and SSIM: - MAE: For PE - CycleGAN, it is 56.89 ± 13.84 HU; - PSNR: 26.69 ± 2.41 dB; - SSIM: 0.92 ± 0.02.

[0047] 5.2 Dose calculation verification: Use 3D gamma analysis to evaluate the accuracy of dose calculation: - Pass rate under the 2mm / 2% standard: 90.13 ± 3.75%; - Pass rate under the 3mm / 3% standard: 97.20 ± 2.52%.

[0048] 5.3 Target volume and OAR dose evaluation: Evaluate the dose deviation of the target volume and organs at risk: - Mean dose deviation of the target volume: < 3%; - Mean dose deviation of OAR: < 3% (except for the lens at 3.38%); - Mean dose deviation of most structures: < 1%.

[0049] The experimental results show that the method of the present invention is not only significantly superior to the existing methods in terms of image quality evaluation metrics, but also fully meets the clinical requirements in terms of dose calculation accuracy, and can provide reliable technical support for the adaptive radiotherapy of nasopharyngeal carcinoma.

[0050] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general - purpose computers, special - purpose computers, embedded processors, or other programmable data - processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data - processing devices generate means for realizing the functions specified in one or more flows or multiple flows and / or blocks Figure 1 one or more flows and / or blocks Figure 1 one or more blocks or multiple blocks.

[0051] In the description of the present invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality of" means two or more unless otherwise specifically defined.

[0052] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these modifications and variations.

Claims

1. A CBCT image enhancement method based on perceptual enhancement recurrent generative adversarial network, characterized in that: The following steps are involved: Based on the CycleGAN network, the generator improves the U-Net structure and integrates residual connections and introduces multi-perceptual loss functions, and the discriminator introduces negative samples to form a dual contrast loss to build a perceptual enhanced cyclic generative adversarial network. Based on the perception-enhanced cyclic generative adversarial network, the CBCT image is enhanced.

2. The CBCT image enhancement method based on perceptual enhancement recurrent generative adversarial network according to claim 1, characterized in that: In the process of improving the generator, an improved U-Net structure is adopted. The generator includes an encoder, a decoder and cross-layer connections. The encoder gradually extracts multi-scale semantic features through downsampling, and the decoder restores image details through upsampling. The cross-layer connection enables the feature map of the encoder to be directly passed to the corresponding layer of the decoder. At the same time, the ResNet-style residual block is introduced. And by fusing the structural similarity index SSIM, the mean absolute error MAE and the cross entropy CE loss as the multi-perceptual loss function, the generated image quality is constrained from multiple dimensions.

3. The CBCT image enhancement method based on perceptual enhancement recurrent generative adversarial network according to claim 2, characterized in that: In the process of improving the generator, the generator is composed of an input layer, an encoder, a converter and a decoder, wherein, The input layer uses a convolution layer with a 7×7 convolution kernel; The encoder contains two downsampling layers with a stride of 2. Each layer uses a 3×3 convolution with stride=2 to halve the feature map size, followed by an instance normalization layer and a ReLU activation function. The transformer includes 9 residual blocks; The decoder includes two transposed convolutional layers for 2x upsampling, and finally generates the target image through a 7×7 convolutional layer followed by a Tanh activation function.

4. The CBCT image enhancement method based on perceptual enhancement recurrent generative adversarial network according to claim 3 is characterized by: In the process of improving the discriminator, the discriminator adopts an improved PatchGAN structure; First layer: 4×4 convolutional layer with a stride of 2; Middle layer: 4 convolution blocks, each block contains a convolution layer, an instance normalization layer and a LeakyReLU activation function; Output layer: 1×1 convolution layer → Sigmoid, outputs the discriminant probability map.

5. The CBCT image enhancement method based on perceptual enhancement recurrent generative adversarial network according to claim 4, characterized in that: In the process of introducing the dual contrast loss, the adversarial loss and the source-target domain distribution difference loss are used as the dual contrast loss, wherein the adversarial loss is used to discriminate the difference between the generated image and the real CT image; The source-target domain distribution difference loss is used to enhance the ability to distinguish CT and CBCT features.

6. The CBCT image enhancement method based on perceptual enhancement recurrent generative adversarial network according to claim 5, characterized in that: In the process of introducing the dual contrast loss, the least square loss function is used as the adversarial loss, and the L2 distance metric is used as the source-target domain distribution difference loss.

7. The CBCT image enhancement method based on perceptual enhancement recurrent generative adversarial network according to claim 6, characterized in that: In the process of enhancing the CBCT image, the CBCT image and the CT image are acquired, and the CBCT is registered to the CT, an initial data set is constructed, and a data set for training a perception-enhanced cyclic generative adversarial network is generated by preprocessing the initial data set, wherein the preprocessing is performed through image unification, normalization processing, image enhancement and data enhancement.

8. The CBCT image enhancement method based on perceptual enhancement recurrent generative adversarial network according to claim 7, characterized in that: During the training of the perception-enhanced recurrent generative adversarial network, the training strategy is optimized, including: The Adam optimizer is used, and the initial learning rate is set to 0.0002; Train for 200 epochs with a batch size of 8; The model performance is evaluated on the validation set every 20 training epochs; When there is no significant performance improvement for 5 consecutive epochs, stop training; During the training process, the update ratio of the generator and the discriminator is 1:1; Grid search method tuning.

9. The CBCT image enhancement method based on perceptual enhancement recurrent generative adversarial network according to claim 8, characterized in that: During the training of the perception-enhanced recurrent generative adversarial network, an early stopping mechanism is set to avoid overfitting.

10. A CBCT image enhancement system based on a perceptual enhancement recurrent generative adversarial network, used to implement the CBCT image enhancement method based on a perceptual enhancement recurrent generative adversarial network as claimed in any one of claims 1 to 9, characterized in that: include: The network construction module is used to build a perceptually enhanced cyclic generative adversarial network based on the CycleGAN network. The generator improves the U-Net structure and integrates residual connections and introduces multi-perceptual loss functions, and the discriminator introduces negative samples to form a dual contrast loss. The image enhancement module is used to enhance the CBCT image based on the perceptual enhancement cycle generative adversarial network.