A PET / CT head and neck tumor automatic segmentation method based on fusion diffusion model
By using a customized feature extraction and task-oriented supervised PET/CT fusion diffusion model, the problem of insufficient modal prior modeling in PET/CT tumor segmentation is solved, achieving high-precision automated tumor segmentation and improving segmentation accuracy and robustness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENYANG UNIVERSITY OF TECHNOLOGY
- Filing Date
- 2025-10-22
- Publication Date
- 2026-07-03
AI Technical Summary
Existing PET/CT tumor segmentation methods suffer from insufficient modal prior modeling and inappropriate fusion timing and path, leading to the dilution of fine-grained information and the easy omission of small-volume and blurred-boundary lesions.
An automated segmentation method for PET/CT head and neck tumors based on a fusion diffusion model is adopted. A customized feature extractor is used to extract features from the imaging characteristics of PET and CT images. Combined with a task-oriented auxiliary supervision mechanism and the DDPM backbone network, the feature enhancement and correction of PET and CT images are achieved. Conditional denoising segmentation is used to improve segmentation accuracy.
It significantly enhances the ability to express modal features, improves the accuracy of tumor region localization and the fineness of boundary characterization, reduces the dilution of small target information and missed detection, and improves the recall rate of small-volume lesions.
Smart Images

Figure CN121437878B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of medical image processing and computer-aided diagnosis technology, and in particular to an automated segmentation method for head and neck tumors based on a fusion diffusion model in PET / CT. Background Technology
[0002] Head and neck cancer (H&N) is one of the most serious malignant tumors threatening public health worldwide. PET and CT imaging, due to their complementary functional and structural properties, have become key tools for the diagnosis, staging, and subsequent monitoring of H&N tumors. PET uses radiopharmaceuticals to reflect tumor metabolic activity and can sensitively detect lesions, but it has low spatial resolution and is prone to false positives. CT provides high-resolution anatomical structures, which is beneficial for delineating boundaries, but lacks functional information. Combining the two can improve the sensitivity and clarity of tumor detection, providing crucial information for clinical treatment planning.
[0003] Accurate identification and volume assessment of tumor lesions are prerequisites for subsequent radiotherapy and surgical treatment. Inaccurate tumor localization can lead to insufficient target dose or organ damage. In clinical practice, relying on radiologists to manually delineate tumor volumes is not only time-consuming and cumbersome but also susceptible to subjective factors, resulting in inconsistent outcomes. Therefore, developing high-precision automated segmentation technology for head and neck tumors has become a research hotspot.
[0004] In recent years, models such as Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) have been widely applied to tumor segmentation tasks. With the development of deep learning, Denoising Diffusion Models (DDPMs) have emerged as a powerful generative model and have been introduced into the field of tumor segmentation. DDPMs recover segmentation results close to the true labels by simulating a progressive denoising process. Compared with traditional methods, diffusion models achieve more stable convergence, possess stronger nonlinear modeling capabilities and better generalization performance, and can more effectively capture complex structural and detailed information in images. In recent years, research based on diffusion models has made good progress in tumor segmentation tasks, but challenges remain in areas such as multimodal complementarity, blurred boundary characterization, and small-volume segmentation.
[0005] In PET / CT tumor segmentation, the key lies in fully leveraging the complementary advantages of the two modalities. Currently, three common fusion strategies exist: input-level fusion (channel stitching of PET / CT images), feature-level fusion (fusion after feature extraction by dual encoders), and output-level fusion (integration of results after segmentation by independent networks). These strategies have driven task development, but the exploration of modal complementarity remains insufficient. For example, Zou et al., in "DGCBG-Net: A dual-branch network with global cross-modal interaction and boundary guidance for tumor segmentation in PET / CT images," achieved complementary learning through global cross-modal interaction and shared downsampling. However, the encoder still used a general structure to extract PET / CT features, failing to incorporate clinical experience to explicitly model PET's metabolic hotspot capture capabilities and CT's anatomical boundary delineation capabilities. This resulted in insufficient utilization of modal advantages, insufficient targeting of extracted features for downstream segmentation tasks, and feature fusion performed at the decoding end, leading to the dilution of early fine-grained clues and missed detection of small lesions or blurred boundaries. In "Head and Neck Tumor Segmentation from [18F]F-FDG PET / CT Images Based on 3DDiffusion Model", Dong et al. extended diffusion to the 3D PET / CT head and neck tumor segmentation task and assessed uncertainty. However, they still lacked explicit modal decoupling and modal perception supervision, and could not effectively utilize PET metabolic priors and CT anatomical priors. In low-contrast and blurred boundary scenarios, the segmentation accuracy and robustness were limited. Summary of the Invention
[0006] In view of the shortcomings of the prior art, the purpose of this invention is to provide an automated segmentation method for PET / CT head and neck tumors based on a fusion diffusion model, which aims to solve the problems of insufficient modal prior modeling, improper fusion timing and path in the prior art, which lead to dilution of fine-grained information and easy missed detection of small-volume and blurred-boundary lesions.
[0007] To achieve the above objectives, the present invention adopts the following technical solution:
[0008] An automated segmentation method for head and neck tumors based on a fusion diffusion model in PET / CT, the automated segmentation method includes:
[0009] Step S1: Construct a dataset based on publicly available PET / CT head and neck tumor slice data, and preprocess the data to obtain standardized PET and CT images;
[0010] Step S2: Perform customized feature extraction on the preprocessed standardized PET and CT images. The PET images are extracted using a high-metabolism feature extractor, and the CT images are extracted using an edge feature extractor. Adaptive feature modeling is achieved for the imaging characteristics of different modal medical images.
[0011] Step S3: Based on customized feature extraction, a task-oriented auxiliary supervision mechanism (TAS) is constructed to enhance and correct PET and CT image features through lightweight prediction head constrained feature learning.
[0012] Step S4: Construct the DDPM backbone network. Using the DDPM backbone network as the segmentation backbone, combine the features and supervision signals of customized feature extraction and the task-oriented auxiliary supervision mechanism TAS to achieve conditional denoising segmentation.
[0013] Step S5: Optimize and train the model constructed in steps S1-S4 to improve the accuracy of the model and achieve automated segmentation of PET / CT head and neck tumor slices.
[0014] Furthermore, in step S1, the data preprocessing to obtain standardized PET and CT images includes:
[0015] Step S11: Use a publicly available PET / CT head and neck tumor dataset. The dataset contains registered PET slices, CT slices, and tumor outlines drawn by radiation oncologists as binary ground truth labels. The labels are aligned with the images on the same voxel grid.
[0016] Step S12: Save the data for each patient in three array files and name them accordingly;
[0017] Step S13: Only patients with three CT / PET / LABEL files that are completely identical in shape are included. Patients with mismatched files or obvious truncation are excluded. Random sampling is used to visually overlay images to verify image-label alignment.
[0018] Step S14: Let the patient set P be divided into five disjoint subsets according to a fixed random seed. The kth fold The test set is used as the test set, and 10% of the remaining patients are selected as the validation set, while the rest are used as the training set. The final report is the mean ± standard deviation at 50% discount.
[0019] Step S15: Read the three array files (CT / PET / LABEL) respectively to obtain three sets of arrays and examine each case, verifying the shape, label and PET / CT slice data of the three.
[0020] Step S16: To reduce cross-case resolution differences, each slice is uniformly resampled to... CT and PET images use bilinear interpolation , Labels use nearest neighbor interpolation To avoid blurred boundaries; for the entire patient case, the output shape is uniformly set to... ;
[0021] Step S17: After converting CT and PET to float32, perform z-score normalization on each slice:
[0022]
[0023] in, This is a slice after resampling. This is the average pixel value of the slice. The standard deviation is denoted as ; when At max( Replace the denominator with a scalar to avoid division by zero, and obtain a standardized slice. , .
[0024] Furthermore, in step S2, feature extraction of the PET image using a high-metabolic feature extractor includes:
[0025] Step S211. In the high-metabolic feature extractor, firstly utilize... Convolutional kernels extract large-scale features from the input image to capture overall metabolic distribution information, and combine batch normalization and nonlinear activation functions to enhance the stability and nonlinear expressive power of the network.
[0026] Step S212. Then pass through The convolutional kernel further models local details and also combines batch normalization and nonlinear activation functions to improve the representation ability of PET features in tumor region identification and localization.
[0027] Step S213. Subsequently, the Sobel operator is introduced to extract edge features from the PET image, highlighting the edge contours of metabolic hotspot areas and enhancing the structural information of lesion boundaries;
[0028] Finally, in step S214, the feature extraction results are concatenated to obtain the final PET image feature representation:
[0029]
[0030] in, This represents the extraction of large-scale and local convolutional features. For Sobel edge extraction functions, For channel cascading, To merge convolutional layers.
[0031] Furthermore, in step S2, the CT image is used to extract features using an edge feature extractor, including:
[0032] Step S221. First, use two consecutive... Convolutional kernels extract features from the input image and combine batch normalization and nonlinear activation functions to enhance nonlinear expressive power and training stability, highlighting the role of CT features in anatomical structure depiction and edge contour modeling.
[0033] Step S222. Subsequently, Discrete Wavelet Transform (DWT) is introduced to decompose the CT image at multiple scales, while preserving the overall structural information in the low-frequency components and the edge and detail features in the high-frequency components.
[0034] Step S223. Finally, the feature extraction results are concatenated to obtain the final CT image feature representation:
[0035]
[0036] in, For continuous convolution extraction functions, is the wavelet decomposition function.
[0037] Furthermore, in step S3, based on the task-oriented supervised learning (TAS) mechanism, the enhancement and correction of PET image features includes:
[0038] Step S311. The input PET feature map is sequentially subjected to two layers of convolution, normalization, and non-linear activation functions using a lightweight prediction head to enhance non-linear representation capabilities and training stability. A Dropout layer is then added to suppress overfitting. Finally, a... Convolutional layers and sigmoid activation functions output a single-channel lesion probability map. ;
[0039] Step S312. Align the predicted map pixel-by-pixel with the true mask GT at spatial resolution, and calculate the region loss. We employ a weighted combination of binary cross-entropy (BCE) and Dice loss:
[0040] ;
[0041] , These are the weights for BCE and Dice loss, which are dynamically adjusted during training. This is a probability map of lesions;
[0042] Step S313. Subsequently, the loss is backpropagated to the high metabolic feature extractor, guiding the network to learn discriminative features that are highly correlated with the lesion area location in the early stage, thereby achieving PET image feature enhancement and correction.
[0043] Furthermore, in step S3, based on the task-oriented assisted supervision mechanism (TAS), the enhancement and correction of CT image features includes:
[0044] Step S321. The input CT feature map is sequentially subjected to two layers of convolution, normalization, and nonlinear activation functions using a lightweight prediction head to enhance nonlinear representation capabilities and training stability. A Dropout layer is then added to suppress overfitting. Finally, a... Convolutional layers and a sigmoid activation function generate CT edge probability maps. ;
[0045] Step S322. Perform "dilation and erosion" morphological operations on the real mask GT, and subtract the two to generate edge labels. ;
[0046] Step S323. Prediction results and Compare and calculate edge loss We employ position-weighted BCE loss:
[0047]
[0048] Where pos_weight is the weight for enhancing edge pixel supervision;
[0049] Step S324. Backpropagate the loss to the edge feature extractor to improve the model's accuracy and robustness in boundary characterization and anatomical structure preservation.
[0050] Furthermore, the denoising segmentation in step S4 includes:
[0051] Step S41, Forward Diffusion Process: Segmenting True Value Labels Perform T-step Gaussian noise addition to gradually approximate the normal distribution;
[0052] Step S42, Backdiffusion Process: The neural network learns the backdiffusion process to gradually denoise, thereby reconstructing a result close to the original data. Its form can be expressed as:
[0053]
[0054] Here It is the parameter set of the reverse process, starting from the Gaussian noise distribution, that is:
[0055]
[0056] in Representing the original image, the reverse process will show the distribution of latent variables. Transform into data distribution In order to be symmetrical with the forward process, the reverse process will gradually recover the noisy image and obtain the final clear segmentation result;
[0057] Step S43, Overall Loss Function: Combining diffused noise loss, region loss, and edge loss, a linear attenuation scheduler is used to adjust the weights.
[0058]
[0059] in For DDPM noise prediction loss, , The weights are dynamically adjusted.
[0060] Furthermore, in the backpropagation process of step S42, following the standard implementation of DDPM, a U-Net network is used as the learning network, specifically including:
[0061] Step S421. Construct a four-channel conditional tensor: extract the features from the customized feature extraction... , Compared with the original PET / CT images , Channel cascading yields a four-channel conditional tensor. This tensor will serve as the input to the conditional path:
[0062] ;
[0063] Step S422. Temporal Embedding and Feature Encoding: Temporal embedding is generated using a sine-cosine lookup table. via condition encoder With segment encoder To each With the Noise segmentation Encoding yields:
[0064] and ;
[0065] Step S423. Conditional Fusion and Denoising Prediction: At each downsampling scale, a Conditioning module is introduced to perform frequency domain modulation and normalized multiplicative fusion on the backbone features. These features are then combined with the conditional features and fed into the residual block and attention layer, forming a scale-cascaded conditional injection. After passing through the intermediate Transformer and upsampling path, noise is predicted at the end using final_conv.
[0066] ;
[0067] in , and This represents the set of latent variables after multi-scale fusion of two coding paths. For decoder, with To ensure that the modulated signal can adaptively balance the denoising and segmentation reconstruction process at each diffusion step, the dynamic fusion of multimodal conditions is completed during the diffusion process.
[0068] The network objective uses predicted noise, pred_noise, aligned with objective='pred_noise'. Based on DDPM derivation, the mean of the backward step is expressed as:
[0069] ;
[0070] in, For the noisy mask at step t, It is a four-channel cascade, which only uses the noise predicted by the network. Enter, Let be the noise intensity at step t. , ;
[0071] Step S424. Two-level mechanism for conditional fusion: highlight potential target areas based on pixel domain gating, and improve denoising stability and boundary fidelity based on frequency domain analysis.
[0072] Furthermore, in step S424, the two-level mechanism for conditional fusion specifically includes:
[0073] 4.1 Pixel-domain "attention-like" gating:
[0074] The current segmentation result is obtained in each sampling step. The features are integrated into the encoded features of the original image to achieve complementarity. In the middle two layers of the original image encoder, the conditional feature map... Will be compared with segmentation features of the same scale The fusion method employs an attention-like mechanism. To achieve this, firstly, layer normalization is performed on the two sets of feature maps, then element-wise multiplication is performed to obtain the affinity map. Subsequently, this affinity map is used to weight the conditional features to highlight potential target regions. Its functional form is shown below:
[0075] ;
[0076] 4.2 Frequency Domain Analysis: Introducing a Feature Frequency Parser (FF-Parser) to suppress high-frequency noise and preserve structural features in the Fourier domain, thereby improving denoising stability and boundary fidelity.
[0077] Introducing FF-Parser into the feature integration path solves the problem. The high-frequency noise problem caused by embedded integration is addressed by FF-Parser design to constrain it. For noise-related components in the features, learn a trainable frequency interest map, apply it to the Fourier space features, given a decoder feature map. First, perform a two-dimensional FFT along the spatial dimension, which can be represented as:
[0078] ;
[0079] in Representing a two-dimensional FFT, we obtain the representation of the features in the frequency domain. .
[0080] Next, by using a parameterized attention map multiplied by To modulate Spectrum:
[0081] ;
[0082] in This indicates element-wise multiplication, which is equivalent to weighting specific frequency components in the frequency domain, thereby enhancing useful components and suppressing noise components.
[0083] Finally, this study uses inverse FFT (IFFT) to... Reverse spatial domain:
[0084]
[0085] Spatial features after obtaining constraints , as input for subsequent fusion and decoding.
[0086] Furthermore, the model training and inference in step S5 specifically include:
[0087] Step S51, Training Strategy: In the initial stage of training, optimize only... This allows the main branch to converge; later in the training process, it is gradually introduced... and This allows PET / CT-assisted supervision to be gradually introduced. By adaptively correcting customized feature extraction through gradient backflow of conditional tensors, the separability and boundary quality of the target area are improved, the auxiliary task is avoided from interfering with the main task, and the network is helped to learn discriminative features better.
[0088] Step S52, Inference Process: Only retain the feature extraction and conditional tensor of the customized feature extraction. The construction steps involve inputting test set images and then performing T-step conditional denoising to generate... The segmentation results are then backmapped to the original resolution for evaluation.
[0089] The technical solution adopted in this invention has the following beneficial effects:
[0090] 1. Enhanced modal feature representation capability: Customized feature extraction strategies are designed specifically for PET metabolism and CT anatomy characteristics, which makes up for the modal adaptation defects of single encoders and significantly enhances the relevance and expressiveness of cross-modal features;
[0091] 2. Enhanced segmentation accuracy and robustness: Task-oriented assisted supervision improves the accuracy of tumor region localization and the fineness of boundary characterization through dual constraints of region and edge loss;
[0092] 3. Enhanced early information fidelity and small-volume lesion detection capability: By cascading "PET / CT original image + customized features" in the early channel to form a conditional tensor as the diffusion backbone input, fine-grained clues are continuously transmitted in the diffusion link with high fidelity, reducing the dilution and missed detection of small target information caused by late fusion, which is conducive to improving the recall rate of small-volume and weakly boundaryed lesions. Attached Figure Description
[0093] Figure 1 A schematic diagram of the segmentation process structure of an automated segmentation method for head and neck tumors based on a fusion diffusion model provided by the present invention;
[0094] Figure 2 A customized feature extraction architecture diagram for an automated segmentation method for PET / CT head and neck tumors based on a fusion diffusion model provided by this invention;
[0095] Figure 3 A task-oriented auxiliary supervision structure diagram for an automated segmentation method for head and neck tumors based on a fusion diffusion model provided by this invention;
[0096] Figure 4 This is a comparative visualization of embodiments of the present invention. Detailed Implementation
[0097] To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0098] Based on the above-described invention, this invention proposes an automated segmentation method for head and neck tumors using PET / CT based on a fusion diffusion model. The segmentation process structure is as follows: Figure 1 As shown, the core components are: Customized Feature Extraction Module (CFE), Task-Oriented Supervisory Module (TAS), and DDPM backbone network;
[0099] Workflow diagram: PET / CT images are processed by CFE to extract specific features. TAS corrects the features through region / edge loss. The corrected features and the original image are used to construct a conditional tensor, which is then input into DDPM to perform conditional denoising and generate the final segmentation result.
[0100] Specifically as follows:
[0101] Step 1: Preprocess the input data;
[0102] Step 2: Input the PET / CT images into the customized feature extraction module to achieve more adaptable feature modeling for the imaging characteristics of different modal medical images;
[0103] Step 3: Feature extraction. Through a task-oriented auxiliary supervision module, additional supervision signals are used to enhance the discriminativeness and effectiveness of the features, thereby further improving the accuracy of tumor segmentation.
[0104] Step 4: DDPM serves as the segmentation backbone network, cascading the extracted features and the original image into the backbone network. During the training phase, noise is progressively added to the ground truth and the reverse denoising process is learned. During the inference phase, a segmentation mask is generated through iterative denoising. This method, in conjunction with the conditional representation and auxiliary supervision provided by the customized feature extraction module and the task-oriented auxiliary supervision module, effectively guides the denoising process with metabolic and anatomical priors.
[0105] Step 5: Model training strategy and inference process.
[0106] The technical solution of the present invention will be further described below with reference to specific embodiments and accompanying drawings.
[0107] Step S1: Dataset construction and data preprocessing;
[0108] Furthermore, step S1 specifically includes:
[0109] Step S11: Use a publicly available PET / CT head and neck tumor dataset, which includes registered PET slices, CT slices, and tumor outlines drawn by radiation oncologists as binary ground truth labels. The labels are aligned with the images on the same voxel grid.
[0110] Step S12: Data for each patient is saved as three .npy array files, named according to the following rules (taking patient ID "HN-CHUM-001" as an example): HN-CHUM-001_CT.npy: CT axial slice stack, array shape ( ), data type int16, the value is limited to [-250, 250] after window width / window level preprocessing; HN-CHUM-001_PET.npy: PET axial slice stack, array shape ( ), data type uint8, strength range [0,255]; HN-CHUM-001_LABEL.npy: split tag stack, array shape ( The data type is uint8, and the value set is {0,1}. Here, S represents the number of axial slices for the patient (different for different patients), and H=W=512. All three are aligned slice by slice on the same voxel grid, in the same axial order (from bottom to top).
[0111] Step S13: Only include patients who have all three files (CT / PET / LABEL) and whose shapes are completely identical, exclude patients with mismatched files or obvious truncation, and randomly sample visual overlays to verify image-label alignment.
[0112] Step S14: Let the patient set P be divided into five disjoint subsets according to a fixed random seed. The kth fold The test set is used as the test set, and 10% of the remaining patients are selected as the validation set and the rest as the training set. The final report is the mean ± standard deviation of 50%.
[0113] Step S15: Read *_CT.npy, *_PET.npy, and *_LABEL.npy respectively to obtain three sets of arrays and check each case: whether the shapes of the three are consistent, whether the label is binary, and whether PET / CT contains NaN / Inf (if so, discard or repair with local statistics).
[0114] Step S16: To reduce cross-case resolution differences, each slice is uniformly resampled to... CT and PET images use bilinear (or cubic spline) interpolation. , Labels use nearest neighbor interpolation To avoid blurred boundaries; for the entire patient case, the output shape is uniformly set to ( ).
[0115] Step S17: After converting CT and PET to float32, perform z-score normalization on each slice:
[0116]
[0117] in, This is a slice after resampling. This is the average pixel value of the slice. The standard deviation is denoted as ; when At max( Replace the denominator with 0 to avoid division by zero. This yields a standardized slice. , .
[0118] Step S2: Input the preprocessed PET / CT images into the customized feature extraction module (CFE) to achieve more adaptable feature modeling based on the imaging characteristics of different modalities of medical images, such as... Figure 2 As shown in Figure 2, this is a diagram of the CFE (Customized Feature Extraction) architecture:
[0119] Core components: The left side is the edge feature extractor, and the right side is the high-metabolism feature extractor;
[0120] Workflow diagram: CT images are processed through "continuous convolutional layers + batch normalization + nonlinear activation function" and "DWT multi-scale decomposition module" to output anatomical features. PET images are processed using a combination of multi-scale convolutional layers, batch normalization, and non-linear activation functions, along with a Sobel edge detection module, to output metabolic features. ;
[0121] Furthermore, step S2 specifically includes:
[0122] Step S21: The PET image is fed into the high-metabolism feature extractor (right branch) for feature extraction.
[0123] 1. In a high-metabolic feature extractor, this invention first utilizes Convolutional kernels extract large-scale features from the input image to capture overall metabolic distribution information, and combine batch normalization and nonlinear activation functions to enhance the stability and nonlinear expressive power of the network.
[0124] 2. Then through The convolutional kernel further models local details and also combines batch normalization and nonlinear activation functions to improve the representation ability of PET features in tumor region identification and localization.
[0125] 3. Subsequently, the Sobel operator was introduced to extract edge features from PET images, highlighting the edge contours of metabolic hotspots and enhancing the structural information of lesion boundaries, thereby compensating for the boundary blurring problem of PET under insufficient spatial resolution.
[0126] 4. Finally, the feature extraction results are concatenated to obtain the final PET image feature representation:
[0127]
[0128] in, This represents the extraction of large-scale and local convolutional features. For Sobel edge extraction functions, For channel cascading, To merge convolutional layers.
[0129] Step S22: The CT image is processed by the edge feature extractor (left branch) for feature extraction.
[0130] 1. First use two consecutive ones Convolutional kernels extract features from the input image and combine batch normalization and nonlinear activation functions to enhance nonlinear expressive power and training stability, highlighting the role of CT features in anatomical structure depiction and edge contour modeling.
[0131] 2. Subsequently, Discrete Wavelet Transform (DWT) is introduced to perform multi-scale decomposition of CT images, while preserving the overall structural information in low-frequency components and the edge and detail features in high-frequency components. In this way, the model can further enhance its sensitivity to tumor boundaries and subtle anatomical structures while preserving the global anatomical context, thereby overcoming the limitations of single convolution extraction in cross-scale modeling;
[0132] 3. Finally, the feature extraction results are concatenated to obtain the final CT image feature representation:
[0133]
[0134] in, For continuous convolution extraction functions, is the wavelet decomposition function.
[0135] Step S3: Construct a task-oriented auxiliary supervision mechanism (TAS) based on the CFE module, and learn constrained features through lightweight prediction head, such as... Figure 3 As shown, Figure 3 TAS (Task-Oriented Assistant Supervision) architecture diagram:
[0136] Core components: The left side is the edge feature correction branch, and the right side is the high metabolic feature correction branch.
[0137] Flowchart: A lightweight prediction head using "continuous convolution + normalization + non-linear activation function + Dropout + convolutional layer + activation function" outputs a lesion probability map. , and the Ground Truth Mask Pixel-by-pixel alignment at spatial resolution is used to calculate region loss. ; Marginal probability maps are generated using structurally similar auxiliary prediction heads. We use morphological operations to dilate and erode the ground plane (GT), and subtract the two to obtain the edge label. Calculate edge loss ;
[0138] Furthermore, step S3 specifically includes:
[0139] Step S31: High-metabolic feature correction branch (corresponding to PET features):
[0140] 1. This prediction head first applies two layers of convolution, normalization, and non-linear activation functions sequentially to the input PET feature map to enhance non-linear representation capabilities and training stability. Then, a Dropout layer is added to suppress overfitting. Finally, it passes through a... Convolutional layers and sigmoid activation functions output a single-channel lesion probability map. ;
[0141] 2. The predicted map is aligned pixel-by-pixel with the ground truth (GT) mask in spatial resolution, and the region loss is calculated. We employ a weighted combination of binary cross-entropy (BCE) and Dice loss:
[0142] ;
[0143] 3. Subsequently, the loss is backpropagated to the high metabolic feature extractor, guiding the network to learn discriminative features that are highly correlated with the localization of lesion areas in the early stages.
[0144] Step S32: Edge feature correction branch (corresponding to CT features):
[0145] 1. Using a prediction head with the same structure as S31, generate a CT edge probability map. ;
[0146] 2. Perform the "dilation and erosion" morphological operation on GT, and subtract the two to generate edge tags. ;
[0147] 3. Prediction results and Compare and calculate edge loss We employ position-weighted BCE loss:
[0148]
[0149] Where pos_weight is the weight for enhancing edge pixel supervision;
[0150] 4. The loss is backpropagated to the edge feature extractor, which improves the accuracy and robustness of the model in boundary characterization and anatomical structure preservation.
[0151] Step S4: Construction and conditional fusion of the DDPM backbone network. Using DDPM as the segmentation backbone, and combining the features and supervision signals of CFE and TAS, conditional denoising segmentation is achieved.
[0152] Furthermore, step S4 specifically includes:
[0153] Step S41: Forward diffusion process: splitting the true value into labels Perform T-step Gaussian noise addition to gradually approximate the normal distribution;
[0154] Step S42: Reverse diffusion process (core):
[0155] Neural networks learn the inverse diffusion process to gradually denoise, thereby reconstructing a result that closely approximates the original data. Its form can be expressed as:
[0156]
[0157] Here It is the parameter set for the reverse process. Starting from the Gaussian noise distribution, that is:
[0158]
[0159] in Representing the original image, the reverse process will show the distribution of latent variables. Transform into data distribution To achieve symmetry with the forward process, the reverse process gradually recovers the noisy image, resulting in a final, clear segmentation result.
[0160] Following the standard implementation of DDPM, this method still uses U-Net as the learning network.
[0161] 1. Construct a four-channel conditional tensor: extract the CFE... , Compared with the original PET / CT images , Channel cascading yields a four-channel conditional tensor. This tensor will serve as the input to the conditional path:
[0162] ;
[0163] 2. Temporal Embedding and Feature Encoding: Temporal embeddings are generated using a sine-cosine lookup table. via condition encoder (Conditional Path) and Segment Encoder (Main path) respectively for With the Noise segmentation Encoding yields:
[0164] and ;
[0165] 3. Conditional Fusion and Denoising Prediction: At each downsampling scale, a Conditioning module is introduced to perform frequency domain modulation and normalized multiplicative fusion of the backbone features. These features are then combined with the conditional features and fed into the residual block and attention layer, forming a scale-cascaded conditional injection. After passing through the intermediate Transformer and upsampling path, noise is predicted at the end using final_conv.
[0166] ;
[0167] in , and This represents the set of latent variables after multi-scale fusion of two coding paths. For decoder, with To modulate the signal, the model adaptively balances denoising and segmentation / reconstruction at each diffusion step, thus achieving dynamic fusion of multimodal conditions during diffusion. The network objective uses predicted noise (pred_noise), aligned with objective='pred_noise'. Based on the DDPM derivation, the mean of the backward step can be written as:
[0168] ;
[0169] in, For the noisy mask at step t, It is a four-channel cascade, which only uses the noise predicted by the network. Enter, Let be the noise intensity at step t. , .
[0170] 4. Two-level mechanism for conditional fusion: pixel-domain gating and frequency-domain parsing (FF-Parser)
[0171] 4.1 Pixel-domain "attention-like" gating
[0172] The current segmentation result is obtained in each sampling step. The features are integrated into the encoded features of the original image, achieving complementarity between the two. Specifically, in the middle two layers of the original image encoder, the conditional feature map... Will be compared with segmentation features of the same scale Fusion. The fusion method employs an attention-like mechanism. The implementation first performs layer normalization on the two sets of feature maps, then multiplies them element-wise to obtain the affinity map. This affinity map is then used to weight the conditional features to highlight potential target regions. Its functional form is shown below:
[0173]
[0174] 4.2 Frequency Domain Resolution: A Feature Frequency Parser (FF-Parser) is introduced to suppress high-frequency noise and preserve structural features in the Fourier domain, thereby improving denoising stability and boundary fidelity.
[0175] The model incorporates FF-Parser into the feature integration path to address... High-frequency noise issues arising from embedded integration. FF-Parser design is used for constraint. The features are noise-related components. The main idea is to learn a trainable frequency-of-interest map and apply it to Fourier space features. Given a decoder feature map... First, perform a two-dimensional FFT along the spatial dimension, which can be represented as:
[0176]
[0177] in Representing a two-dimensional FFT, we obtain the representation of the features in the frequency domain. .
[0178] Next, by using a parameterized attention map multiplied by To modulate Spectrum:
[0179]
[0180] in This represents element-wise multiplication. This operation is equivalent to weighting specific frequency components in the frequency domain, thereby enhancing useful components and suppressing noise components.
[0181] Finally, this study uses inverse FFT (IFFT) to... Reverse spatial domain:
[0182]
[0183] Spatial features after obtaining constraints , as input for subsequent fusion and decoding.
[0184] Unlike spatial attention mechanisms, FF-Parser does not operate within local pixel neighborhoods but models globally in the frequency dimension, enabling it to more effectively capture and adjust the overall frequency composition. Therefore, it can adaptively suppress high-frequency noise during training, thereby improving the clarity of segmentation boundaries and the stability of the results.
[0185] Step S43: Overall Loss Function: Combining diffused noise loss, region loss, and edge loss, a linear attenuation scheduler is used to adjust the weights.
[0186]
[0187] in For DDPM noise prediction loss, , The weights are dynamically adjusted.
[0188] Step S5: Model training and inference;
[0189] Step S51: Training Strategy: In the initial stage of training, optimize only... This allows the main branch to converge; later in the training process, it is gradually introduced... and This allows PET / CT-assisted supervision to be gradually introduced. By adaptively correcting the CFE through gradient backflow of the conditional tensor, the separability of the target region and the quality of the boundary are improved, the auxiliary task is prevented from interfering with the main task, and the network is helped to learn discriminative features better.
[0190] Step S52: Inference process: Only CFE feature extraction and conditional tensor are retained The construction steps involve inputting test set images and then performing T-step conditional denoising to generate... The segmentation results are then backmapped to the original resolution for evaluation.
[0191] The following is the verification experiment of this embodiment:
[0192] The implementation details are as follows:
[0193] The overall network architecture is implemented using PyTorch, with software including Python 3.8 and PyTorch 2.0.1, and hardware configuration including an NVIDIA RTX 4090 and an Intel Xeon Gold 5418Y (10 cores). For hyperparameters, this invention uses 150 epochs, a batch size of 4, a learning rate of 1e-4, and 100 timesteps. For CED-Diff, during the training phase, only the diffusion loss of the backbone is optimized until convergence. Subsequently, region loss and edge loss are introduced to implement task-oriented auxiliary supervision, and the weights of each loss are automatically adjusted through a linear decay scheduler. The preprocessing flow and scale settings are consistent between the inference and training phases; the inference results... The mask is then remapped back to its original resolution for evaluation.
[0194] Seven representative models were evaluated on the HeadNeck dataset: U-Net, DiSegNet, Kumar's CNN, MSAM, DGCBG-Net, T-CADiff, and TransDiff, and compared with the proposed method, CED-Diff. Table 1 lists the metrics obtained on this dataset, including Dice, Sensitivity, Precision, and Specificity. Overall, CED-Diff outperforms the other methods in almost all metrics: it has the best Dice value, significantly higher Sensitivity, and a slight improvement in Precision; while the Specificity of each method is basically in the saturation range, with limited discriminative power. Taking Dice as an example, CED-Diff improves by 0.80 and 0.97 percentage points compared to DGCBG-Net and MSAM, respectively; it improves by 1.72 and 2.01 percentage points compared to T-CADiff and TransDiff, respectively; and its improvement is more significant compared to DiSegNet and U-Net.
[0195] Table 1 Experimental Comparison
[0196]
[0197] To facilitate a more intuitive analysis and comparison of different methods, this invention selected representative slice images from the HeadNeck dataset and visualized their segmentation results using different methods. Visual comparison images are shown below. Figure 1As shown in the figure. Here, (A), (B), (C), (D), and (E) depict five different slices. The yellow curves in the figure delineate the tumor boundaries in the PET and CT images. In the segmentation results display, green areas represent true positives, red areas represent false positives, and yellow areas represent false negatives.
[0198] Comparison visualizations, such as Figure 4As shown, in slice (A), the tumor is clearly visible on the CT image but difficult to discern on the PET image. The blurred boundaries of the PET image may lead to false positive segmentation. However, the CED-Diff network proposed in this invention minimizes the occurrence of false positive tumor regions. In slice (B), the tumor is clearly visible on the PET image but not on the CT image. The segmentation results of DiSegNet, Kumar's CNN, MSAM, DGCBG-Net, T-CADiff, and TransDiff show many false negative regions, indicating inadequate segmentation. Meanwhile, U-Net also exhibits false positives, showing a preference for CT features. In slice (C), the tumor shows blurred boundaries and an irregular shape on the PET image. In this case, the blurred boundaries of the PET image may lead to false positive segmentation. The visualization results show false negatives in DiSegNet, DGCBG-Net, T-CADiff, and TransDiff methods, indicating a failure to fully identify the lesion, and false positives in U-Net, Kumar's CNN, and MSAM methods, showing a preference for CT features. For lesions with irregular shapes, CED-Diff can achieve smoother, anatomically consistent boundaries while preserving details, resulting in better segmentation results. In slice (D), the PET image shows highly metabolically normal tissue, leading to numerous false-positive regions in the segmentation results of DiSegNet, Kumar's CNN, and MSAM, indicating a tendency to misdetect adjacent tissues under such interference. Although the DGCBG-Net method introduces a boundary prior branch, its boundary localization accuracy remains insufficient. In contrast, CED-Diff effectively suppresses the interference of highly metabolically normal tissue under dual supervision of region prediction and edge prediction, balancing boundary accuracy and region integrity. In slice (E), small-volume lesions with low signal-to-noise ratios are observed, resulting in false-negative regions in the segmentation results of U-Net and Kumar's CNN, indicating insufficient detection capability for small lesions. DiSegNet, MSAM, T-CADiff, and TransDiff show many false-positive regions, indicating poor boundary handling. Although the DGCBG-Net method effectively reduces false positives, it still suffers from under-segmentation, and false-negative control needs further improvement. CED-Diff maintains both a low false positive rate and a high detection rate in this scenario, which intuitively demonstrates the superiority of the method of this invention.
[0199] In summary, the present invention has the following beneficial effects:
[0200] 1. Customized feature extraction module: In view of the differential characteristics of PET metabolic activity and CT anatomical boundary, feature extraction strategies of "multi-scale convolution + edge operator" (PET) and "continuous convolution + frequency domain multi-scale decomposition" (CT) are designed respectively to realize explicit modeling of modal advantages and solve the problem of insufficient modal feature expression in existing methods;
[0201] 2. Task-oriented auxiliary supervision mechanism: A lightweight prediction head is designed to generate intermediate results of regions / edges. "BCE+Dice weighted loss" (PET region prediction) and "weighted BCE loss" (CT edge prediction) are adopted. The main and auxiliary tasks are coordinated and trained through linear decay weight scheduling to enhance feature discriminativeness.
[0202] 3. Early-channel cascaded multimodal conditional DDPM fusion architecture: The "PET / CT original image + customized features" are cascaded in the early channels to form a conditional tensor. Combined with time embedding and frequency domain analysis modules for optimization, fine-grained clues are continuously and faithfully transmitted in the diffusion link, realizing the dynamic fusion of multimodal information in the diffusion process, improving boundary fidelity and small lesion detection rate.
[0203] In summary, CED-Diff effectively improves the collaborative characterization of metabolic hotspots and anatomical boundaries through differential feature extraction and task-oriented assisted supervision, thus achieving better overall performance on the HeadNeck dataset and demonstrating stronger scalability and robustness.
Claims
1. An automated segmentation method for head and neck tumors based on a fusion diffusion model in PET / CT, characterized in that, Automated segmentation methods include: Step S1: Construct a dataset based on publicly available PET / CT head and neck tumor slice data, and preprocess the data to obtain standardized PET and CT images; Step S2: Perform customized feature extraction on the preprocessed standardized PET and CT images. The PET images are extracted using a high-metabolism feature extractor, and the CT images are extracted using an edge feature extractor. Adaptive feature modeling is achieved for the imaging characteristics of different modal medical images. Step S3: Based on customized feature extraction, a task-oriented auxiliary supervision mechanism (TAS) is constructed. Through lightweight prediction head constrained feature learning, the enhancement and correction of PET image features and CT image features are achieved. Step S4: Construct the DDPM backbone network. Using the DDPM backbone network as the segmentation backbone, combine the features and supervision signals of customized feature extraction and the task-oriented auxiliary supervision mechanism TAS to achieve conditional denoising segmentation. Step S5: Optimize and train the model constructed in steps S1-S4 to improve the accuracy of the model and achieve automated segmentation of PET / CT head and neck tumor slices; In step S3, based on the task-oriented supervised learning (TAS) mechanism, the enhancement and correction of PET image features includes: Step S311. The input PET feature map is sequentially subjected to two layers of convolution, normalization, and non-linear activation functions using a lightweight prediction head to enhance non-linear representation capabilities and training stability. A Dropout layer is then added to suppress overfitting. Finally, a... Convolutional layers and sigmoid activation functions output a single-channel lesion probability map. ; Step S312. Align the predicted map pixel-by-pixel with the true mask GT at spatial resolution, and calculate the region loss. We employ a weighted combination of binary cross-entropy (BCE) and Dice loss: ; , These are the weights for BCE and Dice loss, which are dynamically adjusted during training. This is a probability map of lesions; Step S313. Subsequently, the loss is backpropagated to the high metabolic feature extractor, guiding the network to learn discriminative features that are highly correlated with the localization of the lesion area in the early stage, thereby achieving feature enhancement and correction of PET images; In step S3, based on the task-oriented supervised learning (TAS) mechanism, the enhancement and correction of CT image features includes: Step S321. The input CT feature map is sequentially subjected to two layers of convolution, normalization, and nonlinear activation functions using a lightweight prediction head to enhance nonlinear representation capabilities and training stability. A Dropout layer is then added to suppress overfitting. Finally, a... Convolutional layers and a sigmoid activation function generate CT edge probability maps. ; Step S322. Perform "dilation and erosion" morphological operations on the real mask GT, and subtract the two to generate edge labels. ; Step S323. Prediction results and Compare and calculate edge loss We employ position-weighted BCE loss: ; Where pos_weight is the weight for enhancing edge pixel supervision; Step S324. Backpropagate the loss to the edge feature extractor to improve the model's accuracy and robustness in boundary characterization and anatomical structure preservation; Step S41, Forward Diffusion Process: Segmenting True Value Labels Perform T-step Gaussian noise addition to gradually approximate the normal distribution; Step S42, Backdiffusion Process: The neural network learns the backdiffusion process to gradually denoise, thereby reconstructing a result close to the original data. Its form can be expressed as: ; Here It is the parameter set of the reverse process, starting from the Gaussian noise distribution, that is: ; in Representing the original image, the reverse process will show the distribution of latent variables. Transform into data distribution In order to be symmetrical with the forward process, the reverse process will gradually recover the noisy image and obtain the final clear segmentation result; Step S43, Overall Loss Function: Combining diffused noise loss, region loss, and edge loss, a linear attenuation scheduler is used to adjust the weights. ; in For DDPM noise prediction loss, , To dynamically adjust the weights; In the backpropagation process of step S42, following the standard implementation of DDPM, a U-Net network is used as the learning network, specifically including: Step S421. Construct a four-channel conditional tensor: extract the features from the customized feature extraction... , Compared with the original PET / CT images , Channel cascading yields a four-channel conditional tensor. This tensor will serve as the input to the conditional path: ; Step S422. Temporal Embedding and Feature Encoding: Temporal embedding is generated using a sine-cosine lookup table. via condition encoder With segment encoder To each With the Noise segmentation Encoding yields: and ; Step S423. Conditional Fusion and Denoising Prediction: At each downsampling scale, a Conditioning module is introduced to perform frequency domain modulation and normalized multiplicative fusion on the backbone features. These features are then combined with the conditional features and fed into the residual block and attention layer, forming a scale-cascaded conditional injection. After passing through the intermediate Transformer and upsampling path, noise is predicted at the end using final_conv. ; in , and This represents the set of latent variables after multi-scale fusion of two coding paths. For decoder, with To ensure that the modulated signal can adaptively balance the denoising and segmentation reconstruction process at each diffusion step, the dynamic fusion of multimodal conditions is completed during the diffusion process. The network objective uses predicted noise, pred_noise, aligned with objective='pred_noise'. Based on DDPM derivation, the mean of the backward step is expressed as: ; in, For the noisy mask at step t, It is a four-channel cascade, which only uses the noise predicted by the network. Enter, Let be the noise intensity at step t. , ; Step S424. Two-level mechanism for conditional fusion: highlight potential target areas based on pixel domain gating, and improve denoising stability and boundary fidelity based on frequency domain analysis.
2. The automated segmentation method for head and neck tumors based on a fusion diffusion model according to claim 1, characterized in that, In step S1, data preprocessing to obtain standardized PET and CT images includes: Step S11: Use a publicly available PET / CT head and neck tumor dataset. The dataset includes registered PET slices, CT slices, and tumor outlines drawn by radiation oncologists as binary ground truth labels. The labels are aligned with the images on the same voxel grid. Step S12: Save the data for each patient in three array files and name them accordingly; Step S13: Only include patients who have three CT / PET / LABEL files at the same time and whose shapes are completely identical. Exclude patients with mismatched files or obvious truncation. Randomly sample and visualize the overlay to verify image-label alignment. Step S14: Let the patient set P be divided into five disjoint subsets according to a fixed random seed. The kth fold The test set is used as the test set, and 10% of the remaining patients are selected as the validation set, while the rest are used as the training set. The final report is the mean ± standard deviation at 50% discount. Step S15: Read the three array files (CT / PET / LABEL) respectively to obtain three sets of arrays and examine each case, verifying the shape, label and PET / CT slice data of the three. Step S16: To reduce cross-case resolution differences, each slide is uniformly resampled to... CT and PET images use bilinear interpolation , Labels use nearest neighbor interpolation To avoid blurred boundaries; for the entire patient case, the output shape is uniformly set to... ; Step S17: After converting CT and PET to float32, perform z-score normalization on each slice: ; in, This is a slice after resampling. This is the average pixel value of the slice. The standard deviation is denoted as ; when At max( Replace the denominator with a slash to avoid division by zero, and obtain a standardized slice. , .
3. The automated segmentation method for PET / CT head and neck tumors based on a fusion diffusion model according to claim 1, characterized in that, In step S2, feature extraction of the PET images using a high-metabolic feature extractor includes: Step S211. In the high-metabolic feature extractor, firstly utilize... Convolutional kernels extract large-scale features from the input image to capture overall metabolic distribution information, and combine batch normalization and nonlinear activation functions to enhance the stability and nonlinear expressive power of the network. Step S212. Then pass through The convolutional kernel further models local details and also combines batch normalization and nonlinear activation functions to improve the representation ability of PET features in tumor region identification and localization. Step S213. Subsequently, the Sobel operator is introduced to extract edge features from the PET image, highlighting the edge contours of metabolic hotspot areas and enhancing the structural information of lesion boundaries; Finally, in step S214, the feature extraction results are concatenated to obtain the final PET image feature representation: ; in, This represents the extraction of large-scale and local convolutional features. For Sobel edge extraction functions, For channel cascading, To merge convolutional layers.
4. The automated segmentation method for head and neck tumors based on a fusion diffusion model according to claim 1, characterized in that, In step S2, feature extraction of the CT image using an edge feature extractor includes: Step S221. First, use two consecutive... Convolutional kernels extract features from the input image and combine batch normalization and nonlinear activation functions to enhance nonlinear expressive power and training stability, highlighting the role of CT features in anatomical structure depiction and edge contour modeling. Step S222. Subsequently, Discrete Wavelet Transform (DWT) is introduced to decompose the CT image at multiple scales, while preserving the overall structural information in the low-frequency components and the edge and detail features in the high-frequency components. Step S223. Finally, the feature extraction results are concatenated to obtain the final CT image feature representation: ; in, For continuous convolution extraction functions, is the wavelet decomposition function.
5. The automated segmentation method for PET / CT head and neck tumors based on a fusion diffusion model according to claim 1, characterized in that, In step S424, the two-level mechanism for conditional fusion specifically includes: 4.1 Pixel-domain "attention-like" gating: The current segmentation result is obtained in each sampling step. The features are integrated into the encoded features of the original image to achieve complementarity. In the middle two layers of the original image encoder, the conditional feature map... Will be compared with segmentation features of the same scale The fusion method employs an attention-like mechanism. To achieve this, firstly, layer normalization is performed on the two sets of feature maps, then element-wise multiplication is performed to obtain the affinity map. Subsequently, this affinity map is used to weight the conditional features to highlight potential target regions. Its functional form is shown below: ; 4.2 Frequency Domain Analysis: Introducing the FF-Parser (Feature Frequency Parser) to suppress high-frequency noise and preserve structural features in the Fourier domain, thereby improving denoising stability and boundary fidelity. Introducing FF-Parser into the feature integration path solves the problem. The high-frequency noise problem caused by embedded integration is addressed by the FF-Parser design for constraint. For noise-related components in the features, learn a trainable frequency interest map, apply it to the Fourier space features, given a decoder feature map. First, perform a two-dimensional FFT along the spatial dimension, which can be represented as: ; in Representing a two-dimensional FFT, we obtain the representation of the features in the frequency domain. ; Next, by using a parameterized attention map multiplied by To modulate Spectrum: ; in This indicates element-wise multiplication, which is equivalent to weighting specific frequency components in the frequency domain, thereby enhancing useful components and suppressing noise components. Finally, the inverse FFT is used to... Reverse spatial domain: ; Spatial features after obtaining constraints , as input for subsequent fusion and decoding.
6. The automated segmentation method for head and neck tumors based on a fusion diffusion model according to claim 1, characterized in that, The model training and inference in step S5 specifically include: Step S51, Training Strategy: In the initial stage of training, optimize only... This allows the main branch to converge; later in the training process, it is gradually introduced... and This allows PET / CT-assisted supervision to be gradually introduced. By adaptively correcting customized feature extraction through gradient backflow of conditional tensors, the separability and boundary quality of the target area are improved, the auxiliary task is avoided from interfering with the main task, and the network is helped to learn discriminative features better. Step S52, Inference Process: Only retain the feature extraction and conditional tensor of the customized feature extraction. The construction steps involve inputting test set images and then performing T-step conditional denoising to generate... The segmentation results are then backmapped to the original resolution for evaluation.
Citation Information
Patent Citations
Esophageal cancer tumor target segmentation method based on PET / CT image cross-modal feature fusion
CN116258732A
Multimodal image fusion method based on diffusion model-convolutional neural network
CN119495002A