Low-count PET image quality enhancement method and system based on fusion multi-input cyclic consistent generative adversarial network
By optimizing the model structure and loss function through the MI-CycleWGAN network, the problems of insufficient model generalization ability and low detail fidelity in low-count PET image quality enhancement are solved, achieving efficient low-count PET image quality enhancement and meeting clinical diagnostic needs.
Patent Information
- Application Number
- CN202510804212.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-05-30
- Filing Date
- 2025-06-16
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-06-16
AI Technical Summary
Existing methods for enhancing the quality of low-count PET images suffer from problems such as insufficient model generalization ability, low detail fidelity, insufficient quantitative accuracy, and insufficient attention to actual clinical needs, especially in terms of lesion detectability and quantitative accuracy of standard uptake values in low-count PET images.
We employ a low-count PET image quality enhancement method based on MI-CycleWGAN. This method involves constructing a multi-input cyclic consistent generative adversarial network (CAN), training it with a large sample size of whole-body PET image data, and designing a total loss function that includes perceptual loss, adversarial loss, and cycle consistency loss. We then use a U-Net structure, residual blocks, and attention mechanisms to enhance the image quality.
It significantly improves the image quality of low-count PET images, bringing them closer to standard-count PET images, enhancing lesion detectability and SUV value accuracy of typical anatomical regions, reducing radiation dose and economic costs, and improving the efficiency of PET examinations.
Smart Images

Figure CN120807320A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of PET image processing, and particularly relates to a low-count PET image quality enhancement method and system based on a fusion multi-input cycle-consistent generative adversarial network. BACKGROUND
[0002] Positron Emission Tomography (PET) is an advanced molecular imaging technique with unique clinical application value in oncology, cardiology and neurology. PET imaging is based on coincidence detection technology, and its image quality is affected by intrinsic physical factors, detector structure performance and coincidence photon count values. In clinical work, the PET image quality mainly depends on the coincidence photon count value, which is positively correlated with the radio tracer dose and the acquisition time. Higher coincidence photon count values in the clinic mean higher ionizing radiation dose, tracer cost and longer scan time, thus limiting the scope and efficiency of PET clinical use. Low count (LC) PET corresponds to lower radio tracer dose and fast PET scanning in clinical work. Lower radio tracer dose can reduce the radiation dose of the subject and the staff, and reduce the cost of the tracer. Fast PET scanning can reduce motion artifacts caused by physiological motion of the subject, while improving the efficiency of PET work. However, they will all cause problems such as reduced signal-to-noise ratio, blurred image and lost details due to the reduction of coincidence photon count value, and it is difficult to meet the needs of clinical diagnosis. Therefore, how to improve the image quality of low count PET (LCPET) imaging to obtain the same diagnostic information as the standard dose / standard acquisition time, i.e. standard count (SC) PET imaging, commonly used in current clinical work, so as to reduce the radiation dose of PET examination and improve the efficiency of PET work, has important significance and practical value in PET clinical application.
[0003] Current research can be divided into two directions: the first direction is to integrate the improvement method into the PET image reconstruction process, such as the direct reconstruction method using end-to-end network to map the PET raw data to the PET image, and the deep learning reconstruction method (DLR) combining deep learning with Bayesian reconstruction method. This direction belongs to computationally intensive, often requiring a large data set. The second direction is to directly use the improvement method for the reconstructed image, which belongs to the image post-processing method, and the related research can be mainly divided into two categories: traditional method and deep learning method.
[0004] Traditional methods such as nonlinear Gaussian filtering, non-local mean (NLM), wavelet transform, etc. are difficult to obtain ideal output due to the super parameter setting, and the generated image loses a lot of geometric shape and texture information, which limits the effect of traditional methods.
[0005] In recent years, deep learning methods have shown obvious advantages and great application potential in the generation of medical images. Deep learning techniques have been shown to better remove noise while predicting texture details in PET images, preserving lesion information as much as possible. Kaplan et al. divided PET images by body part and trained a residual convolutional neural network to estimate full-dose PET images from 1 / 10 dose PET images. They conducted a study on the prediction and generation of high-quality PET images for different body regions (brain, chest, abdomen, and pelvis). Wang, Nie et al.'s research showed that compared with a single convolutional neural network (CNN), a generative adversarial network (GAN) has better performance in medical image denoising and generation modeling. However, since both the input and output data of the LCPET to SCPET training model have noise, it is difficult to ensure that the generator of the GAN learns a meaningful mapping, resulting in a unique output for a given input. The generator may create non-existent features or collapse into a narrow distribution. Therefore, the invention uses a cycle-consistent generative adversarial network (CycleGAN) architecture based on GAN, which introduces an inverse transformation in a cyclic manner, adding more constraints to the generator, effectively avoiding model collapse and ensuring that the generator finds a unique meaningful mapping. Experimental results also show that for the LCPET image quality enhancement task, the performance of CycleGAN is better than that of GAN. Lei et al. used a CycleGAN model to predict full-dose PET images from 1 / 8 dose PET images. The model was trained and tested using 25 and 10 PET whole-body images, respectively. Amirhossein et al. used data from 100 patients to compare the performance of CycleGAN and Residual Neural Network (ResNET) in generating full-dose PET images from low-dose PET images. The results showed that CycleGAN had better qualitative and quantitative indicators. Their research also demonstrated the advantages of CycleGAN networks in PET image quality enhancement. Lei et al. further developed a CT-assisted CycleGAN network in 2020 for low-count PET image denoising, which improved the accuracy and quality of the reconstructed images. However, the training and use of this model require structural information provided by CT images, and it cannot be used for PET single modality images. In addition, Xue et al. proposed a domain transformation combined CycleGAN network (LCPR-Net) in 2021, which used PET image data from 30 patients to directly reconstruct high-quality full-count PET images from LC sinusoidal PET data.However, the study simulated different noise levels of sinogram data by system matrix forward projection, introduction of Poisson noise and normalization method, which is different from the real clinical data. In addition, the LCPR-Net network uses the least square loss, which can make the training more stable compared with the cross-entropy loss, but still faces the problem of gradient vanishing or saturation. For PET image generation, a better network structure and loss function are still needed to optimize the training process and image generation quality.
[0006] Although the low-count PET image quality enhancement method based on CycleGAN has made progress, there are still the following problems and limitations: first, the model generalization ability is insufficient: small sample training and part-specific limit the universality and clinical application range of the model. The sample size of previous studies is usually small, which reduces the robustness of the model and affects the generalization of the model, especially for abnormal cases; many models are trained for different parts of the body, while in clinical practice, the body is usually scanned as a whole, so if the model is only effective for one part, the complexity of the model is increased. Some models need CT or MR images to provide anatomical structures as an aid, which cannot be used for PET single modality images, reducing the universality. Second, the detail fidelity is low, and the traditional architecture is easy to lose the edge of small lesions or metabolic heterogeneity information. Third, previous studies have focused more on quantitative indicators of image quality in model evaluation, and the attention to clinical actual needs such as lesion detectability is insufficient, lacking multi-dimensional verification for standard uptake value (SUV) quantitative error and lesion detectability. Most studies only evaluate the overall universal quantitative indicators of the image, such as normalized mean square error (NRMSE), structural similarity index (SSIM), peak signal-to-noise ratio (PSNR), etc., and do not comprehensively evaluate the indicators that are crucial to clinical actual work, such as lesion detectability and SUV values of lesion regions and typical structure regions such as liver pool and mediastinal pool. Fourth, the quantitative accuracy is not enough, especially the SUV value quantitative accuracy which is more concerned by the clinic, and there is still a significant difference between the generated SCPET image and the real SCPET image. SUMMARY
[0007] In view of the problems existing in the prior art, the present application proposes a low-count PET image quality enhancement method and system based on MI-CycleWGAN, which realizes the prediction and generation of SCPET image from LCPET image. Through the optimization of network structure and loss function, the training of large sample whole body PET image data and the combination of comprehensive qualitative and quantitative evaluation methods in clinical practice, the problems of poor PET image quality and poor clinical quantitative indicators after processing by the enhancement model in the prior art are improved.
[0008] In a first aspect, the present application provides a low-count PET image quality enhancement method based on a fusion multi-input cycle-consistent generative adversarial network, comprising:
[0009] obtaining a PET image set, the image set comprising a plurality of training sample pairs, each training sample pair comprising an SCPET image sample and corresponding three LCPET image samples with different count ratios; SCPET is standard count PET, and LCPET is low count PET;
[0010] constructing a fusion multi-input cycle-consistent generative adversarial network, comprising: expanding the input of the generator and the discriminator in the LCPET image domain into three input channels, expanding the output of the generator in the SCPET image domain into three output channels, one input channel or output channel corresponding to one PET image, and the count ratios of the three PET images corresponding to the three input channels or output channels being different;
[0011] designing a total loss function, the total loss function comprising a perception loss, an adversarial loss and a cycle consistency loss; the perception loss comprises the difference between a first generated image and the SCPET image sample in a feature space and the difference between a second generated image and the LCPET image sample in the feature space; the first generated image is an image reconstructed by the LCPET image sample through the generator in the LCPET image domain; the second generated image is an image reconstructed by the SCPET image sample through the generator in the SCPET image domain;
[0012] based on the total loss function, training the fusion multi-input cycle-consistent generative adversarial network using the image set to obtain a trained image quality enhancement model;
[0013] inputting a low-count PET image to be enhanced into the trained generator in the LCPET image domain to obtain an enhanced PET image.
[0014] Further, the generator adopts a U-Net structure, and residual blocks are introduced in the encoder and the decoder, and an attention mechanism and adaptive instance normalization are used at the end of the decoder.
[0015] Further, the discriminator comprises three parallel sub-discriminators, each sub-discriminator adopts a patchGAN structure and adopts spectral normalization; correspondingly, for the same generated image, the generated image is down-sampled to three different resolutions, and the generated images with three different resolutions are input into the three sub-discriminators respectively.
[0016] Further, the total loss function L is:
[0017] L=λ cycle L cycle +Ladv +λ pept L pept
[0018]
[0019] Among them, L cycle is the cycle consistency loss; L adv To combat the loss, the Wasserstein distance formula with gradient penalty term is adopted; L pept is the perceptual loss, L pept1 is the difference between the first generated image and the SCPET image sample in the feature space, L pept2 is the difference between the second generated image and the LCPET image sample in the feature space, L pept =L pept1 +L pept2 ;λ cycle and λ pept is the corresponding weight, N is the batch size, φ is the given feature extractor, G(x i ) and y i are generated images and real images respectively, A and B are defined as LCPET image domain and SCPET image domain respectively, X A Represents the original LCPET images with multiple different count ratios, X B Represents the corresponding original SCPET image, generator G A Represents the mapping from A to B, generator G B Represents the mapping from B to A.
[0020] Furthermore, the training process also includes: clinical evaluation, quantitative index analysis and key anatomical area analysis of the enhanced PET images to optimize model parameters; wherein the key anatomical areas include lesion areas and normal tissue areas.
[0021] Furthermore, the clinical evaluation specifically includes: visual assessment of image quality using a five-point Likert scale; and binary evaluation of lesion detectability.
[0022] Furthermore, the quantitative index analysis specifically includes: quantitative evaluation of the entire image from the normalized root mean square error, structural similarity index, and peak signal-to-noise ratio; analysis of the joint histogram of the generated image and the real SCPET image; and comparative evaluation combined with linear regression parameters.
[0023] Further, the key anatomical region analysis specifically comprises: drawing ROIs in the key anatomical region, calculating the average standard uptake value and the maximum standard uptake value of each ROI according to the requirement of standard uptake value quantitative accuracy in clinical diagnosis, using the SCPET image as a reference standard, and using a Bland-Altman graph to evaluate the consistency of the standard uptake value.
[0024] In a second aspect, the present application provides a low-count PET image quality enhancement system based on a fusion multi-input cycle-consistent generative adversarial network, comprising:
[0025] A data acquisition and preprocessing module is configured to acquire a PET image set, wherein the image set comprises a plurality of training sample pairs, each training sample pair comprising an SCPET image sample and three LCPET image samples corresponding to the SCPET image sample and having different count ratios; SCPET refers to standard count PET, and LCPET refers to low-count PET;
[0026] A network model construction module is configured to construct a fusion multi-input cycle-consistent generative adversarial network, comprising: expanding the input of the generator and the discriminator in the LCPET image domain into three input channels, expanding the output of the generator in the SCPET image domain into three output channels, one input channel or output channel corresponding to one PET image, and the three PET images corresponding to the three input channels or output channels having different count ratios; and designing a total loss function, wherein the total loss function comprises a perception loss, an adversarial loss, and a cycle-consistency loss; the perception loss refers to the difference between the first generated image and the SCPET image sample in the feature space and the difference between the second generated image and the LCPET image sample in the feature space, the first generated image being an image reconstructed by the LCPET image domain generator from the LCPET image sample, and the second generated image being an image reconstructed by the SCPET image domain generator from the SCPET image sample;
[0027] A training optimization module is configured to train the fusion multi-input cycle-consistent generative adversarial network based on the total loss function using the image set to obtain a trained image quality enhancement model.
[0028] An image quality enhancement module is configured to input a low-count PET image to be enhanced into the trained LCPET image domain generator to obtain an enhanced PET image.
[0029] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method of the first aspect.
[0030] In a fourth aspect, the present application provides a non-transitory computer-readable storage medium having stored thereon a computer program which, when executed by a processor, implements the method of the first aspect.
[0031] The present application has the following advantages:
[0032] 1. The present application has better performance by redesigning and optimizing the CycleGAN network (new generator structure, introducing attention mechanism, residual block, perceptual loss and multiple sub-discriminators), which can make the processed low-count PET image closer to the standard count image while better preserving the detail information. Compared with previous studies, the low-count PET image enhanced by the method of the present application has better global indicators such as PSNR and SSIM. At the same time, it has comparable detection effect as the standard count image in lesion detectability, and the SUV value of the lesion area and typical anatomical area also has high precision. Using the method of the present application, the problem of PET image quality decline affecting clinical diagnosis due to reduction of tracer dose, shortening of acquisition time and equipment in current clinical PET examination can be effectively solved, which can greatly reduce the radiation dose of patients and workers in PET examination, reduce the environmental cost and economic cost brought by radioactive tracers, and improve the work efficiency of PET examination.
[0033] 2. Supervised learning method, higher efficiency
[0034] Although the CycleGAN network is originally designed for unsupervised learning tasks, according to the characteristics of PET images and the requirements of image precision for clinical diagnosis, the present application still adopts a paired supervised learning method for training to maintain quantitative pixel values, remove large geometric mismatches, enable the network to focus on mapping details and speed up training, and is more conducive to improving the training speed and precision of the model. The low-count image is obtained by the retrospective reconstruction method of shortening the acquisition time, without the need to increase additional scanning. The paired supervised learning mode only involves the model training stage. After the model training is completed, for later users, only the low-count PET image needs to be provided, and no paired data is needed, which does not increase the complexity of model use.
[0035] 3. Deep learning image enhancement technology, more feasible
[0036] Unlike the deep learning image reconstruction (DLR) method which directly uses deep learning in the PET image reconstruction process, the present application is a post-processing of the image reconstructed by the PET device, which belongs to the deep learning image enhancement technology (DLE). Compared with the DLR method, it requires less computing resources, and can be trained and deployed directly in the back end of the current clinical workflow, rather than the reconstruction algorithm reconstruction which must be participated by the device manufacturer, and is more clinically feasible.
[0037] 4. Through comparative experiments, a clinically acceptable single-bed shortest scan time and a lowest count ratio are obtained. The lowest level of radio tracer activity that can be used in clinical and the shortest acquisition time that can be achieved for a single bed are provided as a beneficial reference. BRIEF DESCRIPTION OF DRAWINGS
[0038] Figure 1 A flowchart of the low-count PET image quality enhancement method based on the fusion of the multi-input cycle consistent generative adversarial network of the application;
[0039] Figure 2 A schematic diagram of the MI-CycleWGAN network model used in the application;
[0040] Figure 3 A schematic diagram of the MI-CycleWGAN network generator structure used in the application;
[0041] Figure 4 A schematic diagram of the MI-CycleWGAN network discriminator structure used in the application;
[0042] Figure 5 A schematic diagram of the clinical-oriented comprehensive evaluation scheme flowchart of the application;
[0043] Figure 6 (a) is a real standard count PET image (corresponding to an acquisition time of 180S), also known as SCPET image and corresponding lesions; (b) is a PET image generated after processing using the MI-CycleWGAN network of the application, also known as MI-CycleWGAN generated image; (c) is an original low-count PET image (corresponding to an acquisition time of 30S), also known as LCPET image and corresponding lesions;
[0044] Figure 7 (a) is a joint histogram of the LCPET image and the SCPET image; (b) is a joint histogram of the MI-CycleWGAN generated image and the SCPET image;
[0045] Figure 8 Bland-Altman plot of typical anatomical region SUVmean (first row) and lesion region SUVmax (second row). Taking the SCPET image as the standard, the LCPET (left) and the MI-CycleWGAN generated image (right) are calculated respectively. The solid red line and the dashed blue line respectively represent the mean and the 95% confidence interval of the SUV difference;
[0046] Figure 9 A structure schematic diagram of a low-count PET image quality enhancement system based on the fusion of the multi-input cycle consistent generative adversarial network of the application;
[0047] Figure 10 A structural block diagram of an electronic device provided for implementation of the present application. DETAILED DESCRIPTION
[0048] In order to make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0049] The present application provides a low-count PET image quality enhancement method and system based on a fusion multi-input cycle consistent generative adversarial network (MI-CycleWGAN), which is characterized by establishing a non-linear end-to-end mapping model through deep learning technology to map low-count PET (LCPET) images to standard-count PET (SCPET) image domain, thereby achieving significant improvement in LCPET image quality.
[0050] The present application significantly enhances the feature expression and image reconstruction performance of the model through multiple improvements to the CycleGAN model and redesign of the generator network. The technical solution of the present application includes six parts: data acquisition and preprocessing, network architecture design, loss function optimization, model training strategy, image enhancement generation, and comprehensive evaluation in a clinical orientation.
[0051] In combination with Figure 1 and Figure 2 , the embodiments of the present application provide a low-count PET image quality enhancement method based on a fusion multi-input cycle consistent generative adversarial network, which specifically includes the following steps:
[0052] S101: Data acquisition and preprocessing. This includes obtaining a PET image set, which includes a plurality of training sample pairs, each training sample pair including an SCPET image sample and a corresponding LCPET image sample. SCPET is standard-count PET, and LCPET is low-count PET.
[0053] S102: Construct a fusion multi-input cycle consistent generative adversarial network, referred to as MI-CycleWGAN.
[0054] Specifically, in the MI-CycleWGAN constructed by the embodiment of the present application, the input of the generator and the discriminator in the LCPET image domain is expanded to three input channels, the output of the generator in the SCPET image domain is expanded to three output channels, and one input channel or output channel corresponds to one PET image, and the count ratios of the three PET images corresponding to the three input channels or output channels are different.
[0055] Specifically, the original CycleGAN architecture usually uses a single channel as input, but this input form has the problem of limited information dimension when processing PET medical images. In order to capture the change process of images with different count ratios, obtain more intermediate state information, and thus more fully utilize the detailed features in the images, the embodiment changes the input form of the network from the original single channel to three-channel input. The multiple channels respectively represent PET images with different count ratios.
[0056] The original data of the PET image is the "coincidence photon count value" stored in the list mode. In clinical practice, the coincidence photon count value is positively correlated with the tracer activity and the acquisition time. The present application adopts a retrospective reconstruction method to obtain low-count PET images by shortening the acquisition time. Since the acquisition time of the standard count image is 180 seconds, the count ratio of the PET image with the acquisition time shortened to 15 seconds is 1 / 12, the count ratio of the PET image with the acquisition time shortened to 20 seconds is 1 / 9, and the count ratio of the PET image with the acquisition time shortened to 30 seconds is 1 / 6, and so on. In short, PET images with different count ratios represent images obtained by reducing the coincidence photon count value of the PET original data to different ratios of the standard count image.
[0057] This multi-channel fusion strategy not only enriches the input feature dimension of the model, but also helps to capture more complex spatial structures and texture details, thereby improving the recognition and recovery ability of the model for the lesion area in the low-count PET image, and providing a more solid data foundation for generating images with more biological consistency.
[0058] S103: design a total loss function, the total loss function includes a perception loss, an adversarial loss and a cycle consistency loss; the perception loss includes a difference between a first generated image and an SCPET image sample in a feature space and a difference between a second generated image and an LCPET image sample in the feature space; the first generated image is an image reconstructed by an LCPET image domain generator from the LCPET image sample; the second generated image is an image reconstructed by an SCPET image domain generator from the SCPET image sample;
[0059] Specifically, due to the change of the network structure, an adaptive instance normalization AdaIN is further introduced into the generator to replace the ontology mapping loss.
[0060] S104: Based on the total loss function, the image set is used to train the fused multi-input cyclically consistent generative adversarial network to obtain a trained image quality enhancement model;
[0061] Specifically, LCPET image samples and SCPET image samples were input into the MI-CycleWGAN network for supervised training. Through a series of ablation experiments, the optimal number of residual blocks and downsampling times were obtained.
[0062] S105: Inputting the low-count PET image to be enhanced into the generator of the trained LCPET image domain to obtain an enhanced PET image.
[0063] Furthermore, in one embodiment, this embodiment also provides a clinically oriented comprehensive evaluation method. Correspondingly, during the training process, the method further includes optimizing the model parameters in combination with the comprehensive evaluation method.
[0064] Specifically, the comprehensive evaluation method includes clinical evaluation, quantitative index analysis and key anatomical area analysis; key anatomical areas include lesion areas and normal tissue areas.
[0065] In one embodiment, the data acquisition and preprocessing process is as follows: a large amount of whole-body data of the patient collected by the full digital PET / CT system is collected. 18 F-FDG PET images, excluding data without obvious lesions, were used to form a PET image set encompassing a wide range of disease types. Low-count PET images, designated as sample low-count PET images (LCPET), were obtained using a retrospective reconstruction method with shortened acquisition times: 15 seconds (1 / 12 counts), 20 seconds (1 / 9 counts), 30 seconds (1 / 6 counts), 40 seconds (2 / 9 counts), and 60 seconds (1 / 3 counts). Reconstruction parameters were identical to those for standard acquisition time images (SCPET). After retrospective reconstruction, over 200,000 low-count PET and standard-count PET image pairs were obtained.
[0066] Each standard-count PET sample image was paired with simulated low-count PET sample images of 15 seconds (1 / 12 counts), 20 seconds (1 / 9 counts), and 30 seconds (1 / 6 counts) to form a dataset. After desensitization of the raw data, the PET images were resized from 288 × 288 to 256 × 256, and the .dcm format PET cross-sectional images were converted to the .npy format.
[0067] The network training adopts a supervised training mode, and the required data is arranged in sequence according to the cross-sectional level, accumulated case by case, and ensured that the LCPET data and the SCPET data correspond to the same sequence and adopt the same naming format. The data set is divided into a training set, a verification set and a test set according to a ratio of 8:1:1.
[0068] It should be noted that the quality of the PET image in the clinic mainly depends on the coincidence photon count value, which is positively correlated with the radio tracer dose and the acquisition time. Low count PET (LCPET) corresponds to low tracer dose and short scan time in clinical practice, so the present application can be used for low count PET images and low quality PET images in all cases, including low dose PET images, fast scan PET images and quality enhancement and evaluation of low quality PET images due to equipment reasons.
[0069] In one embodiment, the network architecture of the MI-CycleWGAN is as shown in Figure 2 The MI-CycleWGAN provided in this embodiment is based on the CycleGAN architecture, and is obtained by improving the generator, the discriminator and the loss function in the CycleGAN, and can complete the generation task of multiple-to-one, that is, can generate one SCPET image according to the input of multiple LCPET images with different count ratios.
[0070] Specifically, Figure 2 In this embodiment, A and B are defined as the LCPET image domain and the SCPET image domain respectively; the generator G A represents the mapping from A to B; the generator G B represents the mapping from B to A. The two discriminators D A and D B are used to judge whether the input image is generated by the generator or a real image; specifically, the discriminator D A is used to discriminate between the multiple LCPET images reconstructed by the generator G B and the original multiple LCPET images; and the discriminator D B is used to discriminate between the SCPET image reconstructed by the generator G A and the original SCPET image.
[0071] In this embodiment, since the generator adopts a multiple input mode, there is no ontology mapping loss (L identity ) in the MI-CycleWGAN network. In addition, the perceptual loss is introduced on the basis of the CycleGAN network in this embodiment, so that the generator can retain more image detail information in the process of restoring the low count PET image to the standard count PET image. The perceptual loss (L pept), so the loss function of the MI-CycleWGAN network includes the adversarial loss (L adv ), cycle consistency loss (L cycle ) and perceptual loss (L pept ), the total loss function is calculated as follows:
[0072] Therefore, the total network loss is defined as formula (1):
[0073] L=λ cycle L cycle +L adv +λ pept L pept (1)
[0074] by Figure 2 For example, set X A Represents the original LCPET images with multiple different count ratios, X B Represents the corresponding original SCPET image, and the calculation process of each loss is as follows:
[0075] (1) Perceptual loss as a guiding signal for the generator. A After the generator G A Get G A (X A ), that is, the first generated image; calculate G A (X A ) and X B The perceptual loss L between pept1 ; similarly, X B After the generator G B Get G B (X B ), that is, the second generated image; calculate G B (X B ) and X A The perceptual loss L between pept2 ;
[0076] In this example, the VGG-19 network is used as the feature extractor. VGG-19 contains 16 convolutional layers and 3 fully connected layers. The feature map is extracted from the output of the 16th convolution layer, and the difference between the generated image and the target image in the feature space is calculated. The perceptual loss serves as a guiding signal for the generator, prompting the image output by the generator to be closer to the target image. The calculation of the perceptual loss is shown in formula (2):
[0077]
[0078] Where N is the batch size, φ is the pre-trained VGG-19 network, G(x i ) and y irespectively are generated image and real image.
[0079] It should be noted that the perceptual loss can also be realized by other pre-training models, such as using ResNet, Inception, DenseNet as a feature extractor to extract feature maps of generated images and real images respectively, and then calculating the perceptual loss according to the above formula (2).
[0080] (2)X B After the generator G B , calculate the adversarial loss L B adversarial loss L B adversarial loss L B adversarial loss L B adversarial loss L A adversarial loss L adv1 ; similarly, X A after the generator G A , calculate the adversarial loss L A adversarial loss L A adversarial loss L A adversarial loss L A adversarial loss L B adversarial loss L adv2 ;
[0081] In this embodiment, the Wasserstein distance is used instead of the JS divergence in the original GAN network, which alleviates the instability problem when the JS divergence compares the data distribution, thereby improving the convergence of the network and the generation quality. The adversarial loss function adopts the Wasserstein distance formula with gradient penalty term, as formula (3):
[0082]
[0083] where x represents a sample from the real data distribution P r , z represents a sample from the prior noise distribution, the first two terms are the distance estimation of Wasserstein, and the last term is the gradient penalty term for network regularization. λ is the weight hyperparameter of the gradient penalty term.
[0084] (3) The goal of the cycle consistency loss is to ensure that LDPET and SDPET can be converted to each other, which ensures that the input image processed by the two generators is as close to the original image as possible, so that the network maintains consistency and stability when performing image conversion tasks. The cycle consistency loss consists of two parts: forward cycle consistency loss L cycle1 and backward cycle consistency loss L cycle2 . X A after the generator G A , and then after the generator G B , G B (GA (X A )),compute G B (G A (X A )) with loss L A on X cycle1 ; X B goes through generator G B then through generator G A to get G A (G B (X B )),compute G A (G B (X B )) with loss L B on X cycle2 . The formula is:
[0085]
[0086] In an embodiment, the network architecture of the two generators G A and G B of the MI-CycleWGAN is the same, as shown in Figure 3 . In order to capture the change process of PET images of different count ratios, so as to obtain more intermediate state information for image generation, the generator structure is designed to be multi-channel input. The basic architecture of the generator adopts U-Net, which generally includes symmetric encoder and decoder structures, and transmits features between corresponding layers through skip connection to enhance the ability of detail preservation.
[0087] In this embodiment, the encoding process of the encoder includes: the input is a three-channel LCPET image with an image size of 256x256x1 in each channel, which is subjected to a convolution layer to extract preliminary features (also the first downsampling operation, as shown by the downward green arrow in Figure 3 ), the convolution kernel size of the convolution layer is 4, the stride is 2, and the padding is 1; then, the input image is subjected to 4 times of linearly connected downsampling operations (as shown by the downward red and purple arrows in Figure 3 ), and batch normalization is performed, features are extracted, and the size of the feature map is gradually reduced, the number of channels of the feature map is gradually increased in the downsampling process, and the output feature map is reserved for use in the skip connection of the subsequent decoding path. Figure 3 The numbers beside the feature maps in Figure 3 and Figure 3 are the sizes of the feature maps, and in combination with the number of channels, it can be seen that the size (widthxheightxnumber of channels) of the feature map is 128x128x64, 64x64x128, 32x32x256, 16x16x512, and 8x8x512 in turn in the downsampling process.
[0088] The decoding process of the decoder includes that the feature map output by the last layer of the encoder is subjected to 6 times of layer-by-layer upsampling operations (such as Figure 3 In the decoding process, the spatial size of the feature map is gradually restored by transposed convolution through the upward 5 arrows and the adjacent 1 left yellow arrow shown in FIG. 6. The output of each upsampling operation is subjected to normalization and ReLU activation, and is spliced with the output of the corresponding layer in the encoder through a jump connection, so that more local features are retained.
[0089] Further, the attention mechanism and AdaIN (adaptive instance normalization) are used to compress the channels after the end of the upsampling stage to complete the noise reduction.
[0090] Specifically, the AdaIN is used instead of the ontology mapping loss (L identity ). The ontology mapping loss is used to limit the autonomous generation of the generator, but due to the change of the network structure, the ontology mapping loss cannot be calculated, so the AdaIN is introduced into the network, which is used to adjust the mean and standard deviation of the feature mapping, so that the feature distribution of the low-count PET image is closer to the standard count image. This not only reduces the noise interference in the low-count PET image, but also adaptively adjusts the contrast and tissue details of the image, thereby improving the quality and biological consistency of the generated image. Overall, combined with the channel attention mechanism and the AdaIN, the model can effectively improve the attention to key areas, improve the accuracy of image conversion, reduce artifacts in the low-count PET image, and improve the quality and biological credibility of the generated image.
[0091] As an implementable manner, the attention mechanism introduced in the tail structure of the generator in the embodiment is the Squeeze-and-Excitation (SE) attention mechanism, so as to strengthen the modeling ability of the network to key features. The SE module performs global average pooling on the spatial features of each channel through the "squeezing" operation, extracts the global information at the channel level, and then generates the importance weight of each channel through the "excitation" operation. This mechanism can adaptively enhance the feature channels highly related to the anatomical structure, the lesion area and the like, while suppressing irrelevant or redundant information, thereby improving the recognition ability of the network to the key areas in the low-count PET image. The introduction of the SE module not only significantly enhances the selective attention ability of the model in the feature extraction stage, but also effectively improves the accuracy of detail restoration and structure preservation in the image generation process. Combined with the U-Net architecture and the subsequent residual module, the SE mechanism further enhances the expressiveness of the feature transmission path, and provides the generator with more robust and biologically consistent image construction capability. The experimental results show that after introducing the SE attention mechanism, the generated image has obvious improvement in structural similarity, detail preservation and artifact suppression.
[0092] Further, a plurality of residual blocks are introduced in the generator structure for enhancing the feature extraction capability, each residual block being composed of two convolution layers and a ReLU activation function layer. Specifically, the last two down-sampling layers of the encoder and the last two down-sampling layers of the decoder are all set to be residual blocks. The residual modules are integrated into the generator network in combination with the U-Net-based architecture to retain more image details and improve the feature expression capability of the network.
[0093] Specifically, high-level features can be effectively extracted by the residual blocks to enhance the learning ability of the model. With the support of the skip connection, the decoder can effectively recover the local features of the low-count image and complete the low-count PET image denoising in combination with the global features to obtain the final denoised PET image. Meanwhile, the AdaIN in combination with the residual connection enables the key anatomical information to be retained and enhances the stability and generalization ability of the generative adversarial network.
[0094] The generator provided by the embodiment of the application introduces a multi-channel input and effectively improves the modeling capability of the model for noise distribution and image structure through the synergistic effect of the U-Net architecture, the AdaIN normalization, the attention mechanism and the residual module, thereby significantly enhancing the feature expression and image reconstruction performance. Experiments show that the generator structure improves the detail quality and subjective perception effect of the generated image while maintaining the consistency of the image structure.
[0095] In one embodiment, a discriminator in the SCPET image domain takes the generated SCPET image or the real SCPET image as input to determine whether it is true. A discriminator in the LCPET image domain takes the generated LCPET image or the real LCPET image as input to determine whether it is true.
[0096] In the embodiment, the discriminator includes three parallel sub-discriminators, each of which adopts a patchGAN structure and spectral normalization. Correspondingly, for the same generated image, the generated image is down-sampled to three different resolutions, and the generated images of the three different resolutions are input into the three sub-discriminators, respectively. Figure 4 As shown in FIG. 5, patchGAN is used to gradually reduce the spatial dimension and extract high-dimensional features of the image through multiple convolution layers. The discriminator is designed as a multi-scale discriminator composed of three sub-discriminators of different scales, and the spectral normalization technology is used in each sub-discriminator to stabilize the training process of the discriminator and improve its discrimination ability at different scales. Through the average pooling down-sampling of the input image, the sub-discriminator can extract multi-scale features, so that the discriminator can comprehensively evaluate the generated image at different scales and capture more details.
[0097] As shown in FIG. 5, patchGAN is used to gradually reduce the spatial dimension and extract high-dimensional features of the image through multiple convolution layers. The discriminator is designed as a multi-scale discriminator composed of three sub-discriminators of different scales, and the spectral normalization technology is used in each sub-discriminator to stabilize the training process of the discriminator and improve its discrimination ability at different scales. Through the average pooling down-sampling of the input image, the sub-discriminator can extract multi-scale features, so that the discriminator can comprehensively evaluate the generated image at different scales and capture more details. Figure 4As shown, the set input is a single-channel image of 256x256x1. First, the input image is processed by a 4x4 convolution layer (stride=2, padding=1), the channel number is expanded to 64, and LeakyReLU (0.2) is used as the activation function. Subsequently, a second 4x4 convolution layer (stride=2, padding=1) is used, the channel number is increased to 128, and Batch Normalization is applied, followed by LeakyReLU (0.2). The third layer of 4x4 convolution (stride=2, padding=1) continues to extract deep features, and the feature map size is reduced to 32x32, and the channel number is increased to 256. The fourth layer of convolution has a stride of 1 (stride=1, padding=1), and the feature map size becomes 31x31x512, further enhancing the feature extraction capability, and still using Batch Normalization and LeakyReLU (0.2) for processing. Finally, a 4x4 convolution layer (stride=1, padding=1) is used to output a single-channel feature map (size 30x30x1) for predicting the authenticity of the local region (Patch). The entire network extracts multi-scale features, enabling the discriminator to effectively distinguish the authenticity of the generated image and enhancing the stability of the training.
[0098] To verify the effectiveness of the scheme, the present application also provides the following experimental data.
[0099] Experimental environment: Based on Python3.12 PyTorch 2.3.0cuda12.1 platform, the scheme of the present application is implemented on a computer using NVIDIA GeForce RTX4090 GPU. Different count ratios of LCPET images and their corresponding SCPET images in the training set are input into the MI-CycleWGAN model for supervised training.
[0100] First, the network structure is optimized through pre-experiment. Through small sample pre-experiment, the model is verified, and through a series of ablation experiments, the influence of different down-sampling times and the introduction of different residual blocks on the network performance is evaluated. The experimental results show that the generator architecture with 5 times of down-sampling performs best in image quantization indicators. Then, the model parameters are optimized in combination with the comprehensive evaluation method guided by clinical evaluation.
[0101] In the training stage, the Adam optimizer and the CosineAnnealingWarmRestarts learning rate scheduler are used. In the experiment, the parameters T0=10 and Tmult=2 are set; the learning rate experiences frequent decay in the initial period, and then the cycle length is gradually lengthened at each restart, so that the model can converge quickly in the early stage of training, and also can keep the learning rate in a suitable range through the gradually increasing cycle length in the later stage, so as to promote the model to better explore the loss space in different training stages and achieve better generalization performance. The MI-CycleWGAN denoising model is trained in small batches of 32 image patches. The weight hyperparameter λ of the gradient penalty term is set to 10, the weight λ of the perceptual loss is set to 1.5, and the weight λ of the cycle consistency loss function is set to 4. The determination of the hyperparameters is based on the experimental results under the condition of setting the period to 60 times. The learning rate is selected to be 1e-4, and the two exponential decay rates of the momentum estimation are β1=0.5 and β2=0.999, respectively. pept cycle
[0102] For the trained MI-CycleWGAN model, the generator network G A is used to convert the LCPET into SCPET images. The LCPET images that need to be quality enhanced or the low-quality PET images caused by devices and other reasons are input into the generator network G A , and the corresponding SCPET images, i.e., high-quality PET images, can be output.
[0103] A comprehensive evaluation scheme is designed. The purpose of LCPET image quality enhancement is to obtain images that can meet the needs of clinical diagnosis under the premise of minimizing radiation dose and shortening scanning time, so the present application combines more clinical practice in the model evaluation link and focuses on the most critical part of clinical diagnosis: lesion detectability and quantitative indicators. Unlike the evaluation method of the model in previous studies, which often uses only part of the general image quantitative indicators, the present application comprehensively evaluates the images generated by the model from the aspects of lesion contrast and image quality, as shown in FIG. 1, mainly including the following steps: Figure 5
[0104] First, clinical evaluation is performed. The image quality is visually evaluated by multiple experienced diagnostic physicians using the Likert five-point scale, which is divided into five categories: (1) cannot be explained, (2) poor, (3) sufficient, (4) good, and (5) excellent. Then, the lesion detectability is evaluated in a binary manner: accepted or not accepted.
[0105] Then, quantitative index analysis was performed. The analysis was performed for the whole image and the key anatomical regions respectively, to evaluate the differences and similarities between the generated image and the real SCPET image in terms of tracer uptake. First, the whole image was quantitatively evaluated, and the index was normalized mean square error (NRMSE), structural similarity index (SSIM), and peak signal-to-noise ratio (PSNR), as shown in the following formulas. Then, joint histogram analysis was performed on the generated image and the real SCPET image to describe the voxel-level correlation between them in terms of tracer uptake, and the linear regression parameters were compared and evaluated, as shown in Figure 7 The fitting curve of the joint histogram to the above data is shown.
[0106]
[0107] wherein μ x and σ x are the mean and variance of the generated image x, μ y and σ y are the mean and variance of the real image y, σ xy is the covariance of the images x and y, C1 and C2 are the stability constants introduced to prevent the denominator from being zero, and the values are related to the pixel value range, generally small. MAX is the maximum pixel value possible in the image, MSE is the mean square error, and PSNR is used to measure the similarity between images, and the higher the value, the smaller the distortion, and the unit is usually dB; x ij is the gray value of the generated image x at the i, j pixel; y ij is the gray value of the real image y at the i, j pixel; N, M are the size (rows and columns) of the image.
[0108] Then, the evaluation of the key anatomical regions was performed, and the main regions included the lesion region and the part of the normal tissue region that was more concerned in clinical practice. The normal tissue region included the brain, the left and right lungs, the liver (ROI with a diameter of 3 cm in the right lobe of the liver, avoiding the large blood vessel region), the thoracic aorta (ROI with a diameter of 1 cm in the thoracic aorta, avoiding the blood vessel wall, representing the mediastinal blood pool), and the abdominal aorta. ROIs were drawn in these key anatomical regions. In combination with the need for quantitative accuracy of the standard uptake value (i.e., SUV value) in clinical diagnosis, the average standard uptake value (i.e., SUVmean) and the maximum standard uptake value (i.e., SUVmax) of each ROI were calculated, and the Bland-Altman plot was used to evaluate the consistency of the SUV values, with the SCPET image as the reference standard. The experimental results are shown in Figure 6 , Figure 7 , Figure 8 and Table 1. In Table 1, the prefix S represents SCPET, the prefix G represents the generated PET, and the prefix L represents LCPET.
[0109] Table 1
[0110]
[0111]
[0112] It should be noted that in the current scheme, some typical normal tissue regions of clinical concern are selected: brain and left and right lung, liver, mediastinal blood pool, abdominal aorta, and representative anatomical regions such as bone marrow, basal ganglia, and spleen can be added according to actual conditions to increase the evaluation of quantitative indicators thereof.
[0113] To sum up, the present application optimizes the generator structure systematically, and significantly enhances the feature modeling and image generation ability of the model through the synergistic integration of multi-channel input, residual module, AdaIN normalization and SE attention mechanism. The multi-channel input expands the information dimension and improves the model's perception of complex tissue structures. The residual module enhances the expression and transmission ability of deep features. AdaIN guides the feature distribution to align with the standard count image, improving the consistency of image style and details. The SE attention mechanism further strengthens the response of key regions and suppresses redundant interference. Under the support of the U-Net backbone architecture, these improvements complement each other, and together build a low-count PET image generation model with better robustness and biological consistency. Experimental results show that the joint use of the above modules is superior to the original CycleGAN structure in terms of image quality, structure restoration and artifact suppression, verifying the effectiveness and necessity of the comprehensive structural optimization.
[0114] Based on the same inventive concept, as shown in Figure 9 The embodiment of the present application also provides a low-count PET image quality enhancement system based on a fusion multi-input cycle-consistent generative adversarial network, which comprises a data acquisition and preprocessing module, a network model construction module, a training optimization module and an image quality enhancement module.
[0115] Specifically, the data acquisition and preprocessing module is configured to obtain a PET image set, the image set comprising a plurality of training sample pairs, each training sample pair comprising an SCPET image sample and a corresponding LCPET image sample; SCPET is standard count PET, and LCPET is low count PET; the network model construction module is configured to construct a multi-input cycle-consistent generative adversarial network, comprising: extending the input of the generator and the discriminator in the LCPET image domain to a plurality of input channels, extending the output of the generator in the SCPET image domain to a plurality of output channels, one input channel or output channel corresponding to one PET image, the count ratios of the plurality of PET images corresponding to the plurality of input channels or output channels being different; and designing a total loss function, the total loss function comprising a perceptual loss, an adversarial loss and a cycle-consistency loss; the perceptual loss refers to the difference between a first generated image and the SCPET image sample in a feature space and the difference between a second generated image and the LCPET image sample in the feature space, the first generated image being an image reconstructed by the LCPET image domain generator from the LCPET image sample, and the second generated image being an image reconstructed by the SCPET image domain generator from the SCPET image sample; the training optimization module is configured to train the multi-input cycle-consistent generative adversarial network based on the total loss function using the image set to obtain a trained image quality enhancement model; and the image quality enhancement module is configured to input a low count PET image to be enhanced into the trained LCPET image domain generator to obtain an enhanced PET image.
[0116] It should be noted that the low count PET image quality enhancement system based on the multi-input cycle-consistent generative adversarial network provided by the embodiments of the present application is to realize the above-mentioned method, and the functions thereof can be referred to the above-mentioned method embodiments, which will not be described here.
[0117] Figure 10 An example of an entity structure diagram of an electronic device is shown in FIG. 1. Figure 10As shown, the electronic device can include a processor 1001, a communications interface 1002, a memory 1003, and a communications bus 1004, wherein the processor 1001, the communications interface 1002, and the memory 1003 complete mutual communication through the communications bus 1004. The processor 1001 can invoke a logical instruction in the memory 1003 to execute a low-count PET image quality enhancement method based on a fusion multi-input cycle-consistent generative adversarial network, the method comprising: obtaining a PET image set, the image set comprising a plurality of training sample pairs, each training sample pair comprising an SCPET image sample and three corresponding LCPET image samples with different count ratios; SCPET is standard count PET, and LCPET is low count PET; constructing a fusion multi-input cycle-consistent generative adversarial network, comprising: expanding the input of the generator and the discriminator in the LCPET image domain to three input channels, expanding the output of the generator in the SCPET image domain to three output channels, one input channel or output channel corresponding to one PET image, and the count ratios of the three PET images corresponding to the three input channels or output channels being different; designing a total loss function, the total loss function comprising a perception loss, an adversarial loss, and a cycle consistency loss; the perception loss comprising the difference between a first generated image and the SCPET image sample in a feature space and the difference between a second generated image and the LCPET image sample in the feature space; the first generated image being an image reconstructed by the LCPET image domain generator from the LCPET image sample; the second generated image being an image reconstructed by the SCPET image domain generator from the SCPET image sample; based on the total loss function, training the fusion multi-input cycle-consistent generative adversarial network using the image set to obtain a trained image quality enhancement model; inputting a low count PET image to be enhanced into the trained LCPET image domain generator to obtain an enhanced PET image.
[0118] In addition, the logic instructions in the memory 1003 described above can be implemented in the form of software function units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, etc.
[0119] The embodiments of the present application also provide a computer program product, which comprises a computer program stored on a non-transitory computer readable storage medium, and the computer program comprises program instructions, and when the program instructions are executed by a computer, the computer can execute the low-count PET image quality enhancement method based on the fusion multi-input cycle-consistent generative adversarial network provided by the above-mentioned method embodiments.
[0120] The embodiments of the present application also provide a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the low-count PET image quality enhancement method based on the fusion multi-input cycle-consistent generative adversarial network provided by the above-mentioned method embodiments.
[0121] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be implemented by means of software plus necessary universal hardware platforms, and of course can also be implemented by hardware. Based on such understanding, the above technical solutions essentially or the parts that contribute to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0122] It should be pointed out finally that the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit the same; and although the present application has been described in detail with reference to the foregoing embodiments, it should be appreciated by those skilled in the art that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features thereof can be replaced equivalently; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A low-count PET image quality enhancement method based on a fusion multi-input cyclic consistent generative adversarial network, characterized by: include: Acquire a PET image set, the image set comprising a plurality of training sample pairs, each training sample pair comprising an SCPET image sample and three corresponding LCPET image samples with different count ratios; the SCPET is a standard count PET, and the LCPET is a low count PET; Constructing a fused multi-input cycle-consistent generative adversarial network, including: expanding the inputs of the generator and discriminator in the LCPET image domain to three input channels, and expanding the output of the generator in the SCPET image domain to three output channels, where one input channel or output channel corresponds to one PET image, and the count ratios of the three PET images corresponding to the three input channels or output channels are different; Design a total loss function, the total loss function including perceptual loss, adversarial loss and cycle consistency loss; the perceptual loss includes the difference between the first generated image and the SCPET image sample in the feature space and the difference between the second generated image and the LCPET image sample in the feature space; the first generated image is an image of the LCPET image sample reconstructed by the generator of the LCPET image domain; the second generated image is an image of the SCPET image sample reconstructed by the generator of the SCPET image domain; Based on the total loss function, the image set is used to train the fused multi-input cyclic consistent generative adversarial network to obtain a trained image quality enhancement model; The low-count PET image to be enhanced is input into the generator of the trained LCPET image domain to obtain the enhanced PET image.
2. The low-count PET image quality enhancement method based on fusion of multi-input cycle-consistent generative adversarial networks according to claim 1 is characterized in that: The generator adopts a U-Net structure, introduces residual blocks in the encoder and decoder, and uses an attention mechanism and adaptive instance normalization at the end of the decoder.
3. The low-count PET image quality enhancement method based on fusion of multi-input cycle-consistent generative adversarial networks according to claim 1 is characterized in that: The discriminator includes three parallel sub-discriminators, each of which adopts a patchGAN structure and spectral normalization; correspondingly, for the same generated image, after it is downsampled to three different resolutions, the generated images of the three different resolutions are respectively input into the three sub-discriminators.
4. The low-count PET image quality enhancement method based on fusion of multi-input cycle-consistent generative adversarial networks according to claim 1, characterized in that: The total loss function L is: L=λ cycle L cycle +L adv +λ pept L pept Among them, L cycle is the cycle consistency loss; L adv To combat the loss, the Wasserstein distance formula with gradient penalty term is adopted; L pept is the perceptual loss, L pept1 is the difference between the first generated image and the SCPET image sample in the feature space, L pept2 is the difference between the second generated image and the LCPET image sample in the feature space, L pept =L pept1 +L pept2 ;λ cycle and λ pept is the corresponding weight, N is the batch size, φ is the given feature extractor, G(x i ) and y i are generated images and real images respectively, A and B are defined as LCPET image domain and SCPET image domain respectively, X A Represents the original LCPET images with multiple different count ratios, X B Represents the corresponding original SCPET image, generator G A Represents the mapping from A to B, generator G B Represents the mapping from B to A.
5. The low-count PET image quality enhancement method based on fusion of multi-input cycle-consistent generative adversarial networks according to claim 1, characterized in that: The training process also includes: clinical evaluation, quantitative index analysis and key anatomical area analysis of the enhanced PET images to optimize model parameters; wherein the key anatomical areas include lesion areas and normal tissue areas.
6. The low-count PET image quality enhancement method based on fusion of multi-input cycle-consistent generative adversarial networks according to claim 5, characterized in that: The clinical evaluation specifically includes: visual assessment of image quality using a five-point Likert scale; and binary evaluation of lesion detectability.
7. The low-count PET image quality enhancement method based on fusion of multi-input cycle-consistent generative adversarial networks according to claim 5, characterized in that: The quantitative index analysis specifically includes: quantitative evaluation of the entire image from the normalized root mean square error, structural similarity index, and peak signal-to-noise ratio; analysis of the joint histogram of the generated image and the real SCPET image; and comparative evaluation based on linear regression parameters.
8. The low-count PET image quality enhancement method based on fusion of multi-input cycle-consistent generative adversarial networks according to claim 5, characterized in that: The key anatomical region analysis specifically includes: drawing ROIs in key anatomical regions, calculating the average standard uptake value and maximum standard uptake value of each ROI based on the requirements for quantitative accuracy of standard uptake values in clinical diagnosis, and using SCPET images as the reference standard to evaluate the consistency of standard uptake values using Bland-Altman plots.
9. A low-count PET image quality enhancement system based on a fusion multi-input cyclically consistent generative adversarial network, characterized by: include: a data acquisition and preprocessing module for acquiring a PET image set, wherein the image set includes a plurality of training sample pairs, each training sample pair including an SCPET image sample and three corresponding LCPET image samples with different count ratios; SCPET is a standard count PET image and LCPET is a low count PET image; A network model construction module is used to construct a fused multi-input cycle-consistent generative adversarial network, including: expanding the inputs of the generator and discriminator in the LCPET image domain into three input channels, and expanding the output of the generator in the SCPET image domain into three output channels, wherein one input channel or output channel corresponds to one PET image, and the count ratios between the three PET images corresponding to the three input channels or output channels are different; and designing a total loss function, wherein the total loss function includes perceptual loss, adversarial loss, and cycle-consistency loss; the perceptual loss refers to the difference in feature space between a first generated image and an SCPET image sample and the difference in feature space between a second generated image and an LCPET image sample, wherein the first generated image is an image of the LCPET image sample reconstructed by the generator in the LCPET image domain; and the second generated image is an image of the SCPET image sample reconstructed by the generator in the SCPET image domain; A training optimization module, configured to train the fused multi-input cyclically consistent generative adversarial network using the image set based on the total loss function to obtain a trained image quality enhancement model; The image quality enhancement module is used to input the low-count PET image to be enhanced into the generator of the trained LCPET image domain to obtain the enhanced PET image.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Image denoising method based on multi-channel GAN
CN112270654A
Multi-task learning type generative adversarial network generation method and system for low-dose PET reconstruction
CN112508175A
Three-dimensional PET image reconstruction method and system based on diffusion multi-scale generative adversarial network, and storage medium
CN118941718A
Method and system for generating multi-task learning-type generative adversarial network for low-dose pet reconstruction
US20220188978A1
Low-dose image enhancement method and system based on multiple dose levels, and computer device, and storage medium
WO2021168920A1
Cited By
Self-adaptive strategy optimization method and system for robot with body based on interactive feedback
CN121223793A