A medical image segmentation method capable of learning a spectral prior
Patent Information
- Application Number
- CN202610774023.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-01
- Publication Date
- 2026-08-18
AI Technical Summary
但现有频域辅助分割方法通常直接将输入图像计算得到的解析频谱作为固定先验或辅助特征,该类解析频谱中不仅包含与目标结构相关的信息,还可能包含噪声、伪影以及采集条件变化引起的扰动
1)将频谱先验设计为由频谱扩散分支生成的可学习频谱先验,降低固定解析频谱中噪声、伪影和采集扰动对分割结果的影响。
Smart Images

Figure CN122597442A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical image processing technology, and in particular to a medical image segmentation method with learnable spectral prior. Background Technology
[0002] Medical image segmentation is a crucial task in medical image analysis, aiming to automatically or semi-automatically segment lesion regions, organ regions, or other clinically significant target structures from medical images. Accurate medical image segmentation results can provide important evidence for disease diagnosis, quantitative analysis of lesions, surgical planning, efficacy evaluation, and adjuvant therapy.
[0003] In recent years, with the development of convolutional neural networks, Transformer networks, and diffusion models, deep learning-based medical image segmentation methods have made significant progress. Existing medical image segmentation methods typically extract spatial domain features of images through encoder-decoder structures and utilize skip connections, multi-scale feature fusion, or attention mechanisms to enhance segmentation performance. However, medical images often suffer from problems such as low contrast, noise interference, artifacts, blurred boundaries, large variations in lesion morphology, and significant distribution differences across devices, centers, and datasets. Relying solely on spatial domain grayscale, texture, and local morphological features can easily lead to discontinuous boundaries, missed segmentation of small structures, missegmentation of background regions, and insufficient cross-domain generalization ability.
[0004] To improve the robustness of medical image segmentation, some existing techniques incorporate frequency domain information, extracting spectral features from medical images through Fourier transform, wavelet transform, or frequency domain filtering. Frequency domain information can provide supplementary information for image segmentation from aspects such as global structure, boundary details, and texture variations. However, existing frequency-assisted segmentation methods typically use the analytical spectrum calculated from the input image as a fixed prior or auxiliary feature. This type of analytical spectrum not only contains information related to the target structure but may also contain noise, artifacts, and disturbances caused by changes in acquisition conditions. If a fixed analytical spectrum is directly used for segmentation guidance, irrelevant frequency components are easily introduced into the model, thus affecting the stability of the segmentation results.
[0005] Furthermore, existing diffusion-based medical image segmentation methods mostly focus on single-branch generation or denoising of the segmentation mask, with additional priors typically added to the model as external conditions, auxiliary inputs, or regularization constraints. These methods do not model the spectral prior as a learnable variable and struggle to ensure that the spectral representation is updated synchronously with the mask prediction process. Therefore, during iterative denoising, a fixed spectral prior may dynamically mismatch with the current mask prediction state, limiting the model's ability to model complex medical image structures. Summary of the Invention
[0006] The purpose of this invention is to provide a medical image segmentation method with learnable spectral priors, which transforms spectral information from fixed analytical priors into learnable spectral priors, and enables the learnable spectral priors and segmentation masks to be updated collaboratively during backdiffusion to solve the problems described in the background art.
[0007] To achieve the above objectives, the present invention provides a medical image segmentation method with learnable spectral priors, including a model training phase and a model inference phase; The model training phase includes the following steps: S1. Obtain medical image samples and their corresponding segmentation labels, and perform normalization preprocessing on the medical image samples while extracting image conditional features. The segmentation labels are then converted into continuous mask representations. ; S2. Perform two-dimensional Fourier transform and spectral centering on the preprocessed image to obtain the analytical logarithmic amplitude spectrum; S3. Normalize and adjust the size of the analytical logarithmic amplitude spectrum to obtain the spectral anchor point. And construct a structured spectral statistical target based on the analytical logarithmic amplitude spectrum; S4. At the same diffusion time step, the continuous mask representation and the aforementioned spectral anchor point Noise is added separately to obtain the mask noise state during the training phase. and spectral noise status ; S5. Adjust the mask noise state during the training phase. Spectrum noise status Image conditional features The time-step embedded input dual-branch diffusion model is used to jointly predict mask noise and spectral noise; the dual-branch diffusion model includes a mask diffusion branch and a spectral diffusion branch; S6. Optimize the bi-branch diffusion model jointly based on the segmentation loss, spectral supervision loss, and diffusion denoising loss; The model inference phase includes the following steps: S7. Acquire medical images and perform preprocessing to obtain the medical images to be segmented; S8. Input the medical image to be segmented into the image condition feature extraction unit in the trained bi-branch diffusion model, and extract the image condition features of the medical image to be segmented according to step S1. S9. Initialize the mask noise state and the spectrum noise state; S10. Based on the image condition features, the dual-branch diffusion model performs joint back-diffusion denoising on the mask noise state and the spectral noise state; the mask diffusion branch generates mask prediction results from the mask noise state during the back-diffusion process, and the spectral diffusion branch generates learnable spectral priors from the spectral noise state during the back-diffusion process. S11. In the joint backdiffusion denoising process, the mask diffusion branch and the spectral diffusion branch are spatially and spectrally co-updated so that the learnable spectral prior dynamically guides the generation of mask prediction results during the backdiffusion process. S12. Based on the mask prediction results obtained after back diffusion denoising, generate a medical image segmentation probability map, and obtain the final medical image segmentation result based on the medical image segmentation probability map.
[0008] Preferably, in S1, image conditional features The extraction process is as follows: S101. Input the preprocessed medical image samples into the feature extraction network to extract multi-scale image features: ;in, Indicates the number of feature scales. Indicates the first Image features at various scales; S102. Perform channel alignment and spatial size alignment on image features at each scale: ; in, This represents a 1×1 convolution operation. Indicates an upsampling or downsampling operation. Indicates the aligned first Individual scale features; Perform global average pooling on each aligned scale feature: ; in, This indicates a global average pooling operation. Indicates the first A global description vector of each scale feature; S103. Concatenate multiple scale description vectors and input them into a multilayer perceptron. Obtain the scale weights through the Softmax function. ; in, This represents vector concatenation. This represents a multilayer perceptron. , Indicates the first The weights of each scale feature satisfy: ; S104. Weighted fusion of multi-scale features based on scale weights: ; Further processing using convolution, normalization, and activation functions yields image conditional features: ; in, This represents a 3×3 convolution operation. This indicates a normalization operation. Represents a non-linear activation function. Represents image conditional features.
[0009] Preferably, in S2, the formula for calculating the analytical logarithmic amplitude spectrum is: ; in, This represents the preprocessed image; Represents a two-dimensional Fourier transform. This indicates a spectrum centering operation. Indicates the coordinate position in the spectrum plane; Representing coordinates The analytic logarithmic amplitude spectrum at that location.
[0010] Preferably, in S3, the structured spectral statistics target is constructed in the following manner: S31. Divide the centered spectral plane into... Radial frequency band and Each angular frequency band forms One radial-angular sub-band; S32, for the first The radial frequency band and the first Sub-bands formed by the angular frequency bands The average energy of the analytic logarithmic amplitude spectrum within this sub-band is calculated using the following formula: ; in, Indicates sub-band The number of spectral points included; S33. Combining the average energy values corresponding to all sub-bands into a structured spectral statistical objective: ; in, This represents the structured spectral statistics objective.
[0011] Preferably, in S4, the formula for adding noise to the continuous mask representation and spectral anchor points is: ; in, This indicates the un-noiseed diffusion start point, which includes a continuous mask representation. or spectral anchor ; Indicates diffusion time step The noise state under the following conditions Indicates noise scheduling parameters, Indicates Gaussian noise; In S5, the formula for jointly predicting mask noise and spectral noise is: ; in, This represents a two-branch diffusion model. Indicates the mask noise state. Indicates the spectral noise state. Represents image conditional features. Indicates the diffusion time step. This represents the predicted mask noise. This represents the predicted spectral noise. Based on the prediction results of the noise recovery mask and the learnable spectrum prior: ; ; in, This indicates the mask prediction result. This represents a learnable spectral prior.
[0012] Preferably, in S6, the total loss function of the dual-branch diffusion model is: ; in, Indicates the loss from partitioning. Indicates spectrum monitoring loss, This represents the diffusion denoising loss. , , These represent the weighting coefficients of the corresponding loss terms; The spectral supervision loss is: ; in, This represents the pixel-level spectral anchoring loss between the learnable spectral prior and the spectral anchor point. Indicates radial-angular energy band monitoring loss; and These represent the corresponding weights; The pixel-level spectral anchoring loss is: ; in, This represents a learnable spectral prior. Indicates the spectral anchor point; For learnable spectral priors Energy statistics are performed using the same radial-angular partitioning method as the analytical logarithmic amplitude spectrum to obtain the predicted structured spectral statistical representation. ; The radial-angular energy band supervision loss is: ; in, This represents the predicted structured spectral statistical representation obtained from learnable spectral priors. Represents the structured spectral statistics objective; The diffusion denoising loss is: ; in, and These represent the actual noise in the mask diffusion branch and the spectral diffusion branch, respectively. and These represent the noise predicted by the mask diffusion branch and the spectral diffusion branch, respectively. and These represent the weights of the mask diffusion denoising loss and the spectral diffusion denoising loss, respectively.
[0013] Preferably, in S5, the dual-branch diffusion model further includes an image conditional feature extraction unit and an uncertainty calibration bidirectional modulation unit; The image condition feature extraction unit is used to extract image condition features from the medical image to be segmented; the uncertainty calibration bidirectional modulation unit is used to adjust the information interaction between the mask diffusion branch and the spectral diffusion branch during the back diffusion process.
[0014] Preferably, in S10, the learnable spectral prior is the denoised spectral representation generated by the spectral diffusion branch during the back-diffusion denoising process; The learnable spectral prior is not a fixed analytical spectrum obtained by directly performing a Fourier transform on the medical image to be segmented, but is updated synchronously with the mask prediction results during the backdiffusion process and is used to provide frequency domain structural constraints to the mask diffusion branch.
[0015] Preferably, in S10, the joint back-diffusion denoising process includes: S1001, The mask diffusion branch and the spectral diffusion branch are updated synchronously at the same or corresponding diffusion time steps; S1002. At each diffusion time step, the dual-branch diffusion model predicts mask noise and spectral noise respectively based on the current mask noise state, the current spectral noise state, image condition features, and time step embedding. S1003. Recover the mask prediction state at the current diffusion time step based on the predicted mask noise, and recover the learnable spectral prior at the current diffusion time step based on the predicted spectral noise.
[0016] Preferably, in S11, the spatial-spectrum collaborative update includes: S1101. Obtain the mask branch characteristics and spectral branch characteristics of the current back diffusion stage; S1102. Calculate the mask branch reliability index based on the mask branch features, and calculate the spectrum branch reliability index based on the spectrum branch features; the mask branch reliability index and the spectrum branch reliability index are obtained by performing context encoding, global pooling, and nonlinear mapping on the mask branch features and the spectrum branch features, respectively; S1103. Based on the mask branch reliability index and the spectrum branch reliability index, calculate the first gate coefficient of the mask diffusion branch modulated by the spectrum diffusion branch, and the second gate coefficient of the spectrum diffusion branch modulated by the mask diffusion branch; the first gate coefficient is used to control the strength of the frequency domain structure information transmitted from the spectrum diffusion branch to the mask diffusion branch; the second gate coefficient is used to control the strength of the spatial semantic information transmitted from the mask diffusion branch to the spectrum diffusion branch. S1104. Based on the first gating coefficient, the spectral branch features are modulated to the mask diffusion branch, and based on the second gating coefficient, the mask branch features are modulated to the spectral diffusion branch, so as to achieve bidirectional modulation of uncertainty calibration of the mask diffusion branch and the spectral diffusion branch.
[0017] Therefore, the medical image segmentation method with learnable spectral prior described above has the following beneficial effects: 1) Design the spectral prior as a learnable spectral prior generated by the spectral diffusion branch to reduce the impact of noise, artifacts and acquisition disturbances in the fixed analytical spectrum on the segmentation results.
[0018] 2) By constructing a dual-branch diffusion model through mask diffusion branch and spectrum diffusion branch, the mask prediction results and learnable spectrum priors are updated synchronously in a unified back diffusion denoising process, which alleviates the mismatch between fixed spectrum priors and dynamic mask prediction states.
[0019] 3) During the training phase, analytical spectral anchors and structured spectral statistics are used to constrain the spectral diffusion branch, so that the learnable spectral priors can maintain consistency with the frequency domain structure of medical images while avoiding over-reliance on local noise responses in the original spectrum.
[0020] 4) By constructing a structured spectral statistical target through radial-angular spectral energy statistics, the model can learn frequency domain structure information at different scales and in different directions, thereby improving its ability to express the overall shape, boundary changes and detailed structure of the target.
[0021] 5) By using a spatial-spectral uncertainty calibration bidirectional modulation mechanism, the mask diffusion branch and the spectral diffusion branch can dynamically adjust the information interaction intensity according to their respective reliability. When the boundary information of the mask branch is insufficient, frequency domain structure constraints are introduced, and when the spectral branch is affected by noise, spatial semantic information is used for adjustment, thereby improving the stability of dual-branch collaborative denoising.
[0022] 6) Improves the boundary preservation, structural consistency and robustness of medical image segmentation results, and can be applied to automatic segmentation tasks of skin lesions, tumors, ultrasound lesions, fundus structures, organ regions and other medical image target regions.
[0023] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0024] Figure 1 This is an overall flowchart of a medical image segmentation method with learnable spectral priors according to the present invention; Figure 2 This is a schematic diagram of the overall structure of the dual-branch diffusion model in this invention; Figure 3 This is a schematic diagram of LSP learning under structured spectrum supervision in this invention; Figure 4 This is a schematic diagram of the spatial-spectral uncertainty calibration bidirectional modulation module in this invention; Figure 5 This is a schematic diagram showing the skin lesion segmentation results of the present invention and other methods on the ISIC2018 dataset; Figure 6 This is a schematic diagram of brain tumor segmentation results on the BraTS2021 dataset using the present invention and the remaining methods. Detailed Implementation
[0025] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments. Unless otherwise defined, the technical or scientific terms used in this invention should have the ordinary meaning understood by those skilled in the art. It should be understood that the following embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the present invention. For those skilled in the art, equivalent substitutions made to the network structure, number of parameters, loss function weights, number of diffusion sampling steps, and feature extraction network type without departing from the concept of the present invention should be included within the scope of protection of the present invention.
[0026] Example 1 like Figure 1-4 As shown, this invention provides a medical image segmentation method with learnable spectral priors. The method includes a model training phase and a model inference phase. The medical image segmentation process is modeled as a joint backdiffusion denoising process of mask state and spectral state. A learnable spectral prior is generated through a spectral diffusion branch, and the learnable spectral prior is used to dynamically guide the generation of mask prediction results during the backdiffusion process.
[0027] The specific process of the model training phase is as follows: S1. Obtain medical image samples and their corresponding segmentation labels, denoted as: ; in, Indicates the number of segmentation categories. Indicates the height of the medical image. Indicates the width of the medical image.
[0028] Perform size normalization, grayscale normalization, or intensity normalization on medical images. For example, resize medical images to a preset size and normalize pixel values to a preset range. For multi-channel medical images, convert them to normalized grayscale images, denoted as... ;in, Used for subsequent spectral analysis. For single-channel medical images, they can be directly normalized to obtain the results. For multi-channel medical images, channel averaging, weighted averaging, or channel selection methods can be used to obtain the final image. .
[0029] Image conditional features The extraction process is as follows: S101. Input the preprocessed medical image samples into the feature extraction network to extract multi-scale image features: ;in, Indicates the number of feature scales. Indicates the first Image features at various scales. The extraction method is as follows: Feature extraction networks can be convolutional neural networks, Transformer networks, or hybrid convolutional-Transformer networks. Shallow features are used to preserve boundary, texture, and local detail information, while deep features are used to express semantic structure and global contextual information.
[0030] S102. To obtain stable image conditional features, channel alignment and spatial size alignment are performed on image features at each scale: ; in, This represents a 1×1 convolution operation. Indicates an upsampling or downsampling operation. Indicates the aligned first Individual scale features.
[0031] Perform global average pooling on each aligned scale feature: ; in, This indicates a global average pooling operation. Indicates the first A global description vector of features at each scale.
[0032] S103. Concatenate multiple scale description vectors and input them into a multilayer perceptron. Obtain the scale weights through the Softmax function. ; in, This represents vector concatenation. This represents a multilayer perceptron. , Indicates the first The weights of each scale feature satisfy: .
[0033] S104. Weighted fusion of multi-scale features based on scale weights: ; Further processing using convolution, normalization, and activation functions yields image conditional features: ; in, This represents a 3×3 convolution operation. This indicates a normalization operation. Represents a non-linear activation function. Represents image conditional features. Image conditional features Simultaneously inputting the mask diffusion branch and the spectral diffusion branch from the dual-branch diffusion model guides the reverse diffusion denoising process of the two branches.
[0034] Convert the segmentation labels into continuous mask representations. For binary classification and segmentation tasks, It can be a single-channel binary mask; for multi-class segmentation tasks... It can be used as a multi-channel one-hot encoded mask.
[0035] S2. Process the preprocessed image Performing a two-dimensional Fourier transform and spectral centering, the analytical logarithmic amplitude spectrum is obtained; such as Figure 3 As shown, analytical spectrum anchors are constructed based on medical image samples. These analytical spectrum anchors serve only as non-learning spectrum constraints during the training phase and are not used as fixed spectrum priors directly in the inference phase.
[0036] For normalized images A two-dimensional Fourier transform is performed, followed by spectral centering to obtain the centered frequency domain representation. The analytic logarithmic amplitude spectrum is calculated using the following formula: ; in, Represents a two-dimensional Fourier transform. This indicates a spectrum centering operation. Indicates the coordinate position in the spectrum plane; Representing coordinates The analytic logarithmic amplitude spectrum at that location.
[0037] S3, Analysis of Logarithmic Amplitude Spectrum Normalization and size adjustment are performed to obtain the spectral anchor point during spectral spread branch training. It should be noted that, The spectral anchor points are obtained during the training phase through analytical logarithmic amplitude spectrum normalization and size adjustment. However, the learnable spectral priors used in this method are learned spectral priors generated during the back-diffusion process of the spectral diffusion branch; these are different. To avoid the spectral diffusion branch relying excessively on local noise responses in the analytical spectrum, a structured spectral statistical objective is further constructed. The specific process is as follows: S31. Divide the centered spectral plane into... Radial frequency band and Each angular frequency band forms The radial-angular sub-band. The first... The radial frequency band and the first The sub-bands formed by the angular frequency bands are denoted as: ;in, , For example, take R =5, D =4, forming 20 radial and angular spectral energy bands; the radial band is used to express the frequency distribution at different scales, and the angular band is used to express the structural changes in different directions.
[0038] S32. Perform average energy statistics on the analytical logarithmic amplitude spectrum within each radial-angular sub-band: ; in, Indicates sub-band The number of spectral points included; Indicates the first The radial frequency band and the first The average spectral energy corresponding to each angular frequency band.
[0039] S33. Combining the average energy values corresponding to all sub-bands into a structured spectral statistical objective: ; in, This represents the structured spectral statistics objective.
[0040] S4. At the same diffusion time step, for continuous mask representation and spectral anchor Noise is added separately to obtain the mask noise state during the training phase. and spectral noise status The specific process is as follows: At the same diffusion time step, for continuous mask representation and spectral anchor Noise is added separately to obtain the mask noise state during the training phase. and spectral noise status ; For the un-noiseed diffusion initiation point The forward diffusion noise addition process is represented as: ; in, This indicates the un-noiseed diffusion start point, which includes the continuous mask representation. or spectral anchor ; Indicates diffusion time step The noise state under the following conditions Indicates noise scheduling parameters, This represents Gaussian noise.
[0041] Specifically, for mask diffusion branches, the continuous mask is represented... As the starting point for diffusion of the un-noiseed mask, the mask noise state is obtained: ; For the spectral diffusion branch, the spectral anchor point As the starting point for the spread of the un-noiseed spectrum, the spectral noise state is obtained: ; in, This represents the Gaussian noise corresponding to the mask branch. This represents the Gaussian noise corresponding to the spectral branch.
[0042] S5, Change the mask noise status Spectrum noise status Image conditional features and diffusion time step Input a bi-branch diffusion model to jointly predict mask noise and spectral noise: ; in, This represents a two-branch diffusion model. This represents the predicted mask noise. This represents the predicted spectral noise.
[0043] Based on the prediction results of the noise recovery mask and the learnable spectrum prior: ; ; in, This indicates the mask prediction result. This represents a learnable spectral prior.
[0044] in, This is the learnable spectral prior in this method. The learnable spectral prior is generated by the spectral diffusion branch during the backdiffusion process, rather than by a fixed analytical spectrum obtained by directly performing a Fourier transform on the input medical image.
[0045] S6. Jointly optimize the bi-branch diffusion model based on the segmentation loss, spectral supervision loss, and diffusion denoising loss; the total loss function of the bi-branch diffusion model is: ; in, Indicates the loss from partitioning. Indicates spectrum monitoring loss, This represents the diffusion denoising loss. , , These represent the weighting coefficients of the corresponding loss terms.
[0046] Segmentation loss is used to constrain the probabilistic map of medical image segmentation. With split labels Consistency between them: ; in, This indicates Dice's loss. Represents cross-entropy loss, This represents the weights of the cross-entropy loss.
[0047] The spectrum supervision loss is: ; in, This represents the pixel-level spectral anchoring loss between the learnable spectral prior and the spectral anchor point. Indicates radial-angular energy band monitoring loss; and These represent the corresponding weights.
[0048] The pixel-level spectral anchoring loss is: ; in, This represents a learnable spectral prior. Indicates the spectral anchor point.
[0049] For learnable spectral priors Energy statistics are performed using the same radial-angular partitioning method as the analytical logarithmic amplitude spectrum to obtain the predicted structured spectral statistical representation. .
[0050] The radial-angular energy band supervision loss is: ; in, This represents the predicted structured spectral statistical representation obtained from learnable spectral priors. This represents the structured spectral statistics objective.
[0051] The diffusion denoising loss is: ; in, and These represent the actual noise in the mask diffusion branch and the spectral diffusion branch, respectively. and These represent the noise predicted by the mask diffusion branch and the spectral diffusion branch, respectively. and These represent the weights of the mask diffusion denoising loss and the spectral diffusion denoising loss, respectively.
[0052] The model inference process only requires the input of the medical image to be segmented, without the need for segmentation labels. The specific process is as follows: S7. Obtain the medical image and perform preprocessing. Obtain the medical image to be segmented, denoted as: ;in, Indicates the height of the medical image to be segmented. Indicates the width of the medical image to be segmented. This indicates the number of channels in the medical image to be segmented.
[0053] S8. Input the medical image to be segmented into the image conditional feature extraction unit in the trained bi-branch diffusion model, and extract the image conditional features of the medical image to be segmented according to step S1.
[0054] S9. Initialize mask noise state and spectrum noise state: ; in, Indicates the initial mask noise state. This indicates the initial spectral noise state. This represents a standard Gaussian distribution.
[0055] S10. Based on image conditional features, a dual-branch diffusion model performs joint back-diffusion denoising on both mask noise state and spectral noise state; for example... Figure 2 As shown, the mask diffusion branch generates mask prediction results from the mask noise state during backdiffusion, while the spectral diffusion branch generates learnable spectral priors from the spectral noise state during backdiffusion; the specific process is as follows: S1001, the mask diffusion branch and the spectral diffusion branch are updated synchronously at the same or corresponding diffusion time steps; S1002. At each diffusion time step, the dual-branch diffusion model predicts mask noise and spectral noise respectively based on the current mask noise state, the current spectral noise state, image condition features, and time step embedding. S1003. Recover the mask prediction state at the current diffusion time step based on the predicted mask noise, and recover the learnable spectral prior at the current diffusion time step based on the predicted spectral noise.
[0056] S11, such as Figure 4 As shown, in the joint backdiffusion denoising process, the mask diffusion branch and the spectral diffusion branch are spatially and spectrally co-updated to enable the learnable spectral prior to dynamically guide the generation of mask prediction results during backdiffusion; the spatial-spectral co-updating includes: S1101, in the Each decoding stage and diffusion time step Next, obtain mask branch features. and spectral branching features .
[0057] S1102. Perform context encoding on the mask branch features and the spectral branch features respectively: ; in, and This represents a context encoding operation, which can be a self-attention operation, a convolutional attention operation, or another context modeling operation. Represents mask context features, This represents spectral context features.
[0058] Calculate the branch reliability index based on contextual features: ; ; in, This indicates the reliability index of the mask branch. This represents the reliability index of the spectrum branch. This represents the Sigmoid activation function. and This represents a multilayer perceptron.
[0059] S1103. Calculate the bidirectional gating coefficient based on the mask branch reliability index and the spectrum branch reliability index: ; ; in, This represents the first gating coefficient of the mask branch modulated by the spectral branch. This represents the second gating coefficient of the spectral branch modulated by the mask branch. and This represents the learnable mapping parameters. The first gating coefficient controls the strength of the frequency domain structure information transmitted from the spectral diffusion branch to the mask diffusion branch; the second gating coefficient controls the strength of the spatial semantic information transmitted from the mask diffusion branch to the spectral diffusion branch.
[0060] S1104. Generate cross-branch residuals based on the contextual features of the opposite branch: ; ; in, and This represents a learnable mapping used for channel alignment, preferably a 1×1 convolution.
[0061] Bidirectional residual modulation is performed based on the gating coefficient: ; ; in, This represents element-wise multiplication. This indicates the branching characteristics of the modulated mask. This indicates the branching characteristics of the modulated spectrum.
[0062] In this way, when the mask branch is not reliable enough in the region with blurred boundaries or low contrast, the spectral branch can transmit frequency domain structure information to the mask branch through the first gating coefficient; when the spectral branch is affected by noise or artifacts, the mask branch can transmit spatial semantic information to the spectral branch through the second gating coefficient, thereby realizing the coordinated updating of spatial information and spectral information.
[0063] After a preset number of backdiffusion sampling steps, the mask prediction results are obtained. and learnable spectral priors .
[0064] S12. Based on the mask prediction results obtained after back-diffusion denoising, generate a medical image segmentation probability map, and obtain the final medical image segmentation result based on the medical image segmentation probability map: For binary classification segmentation tasks, the medical image segmentation probability map is binarized according to a preset threshold to obtain the final medical image segmentation result; for multi-class segmentation tasks, the maximum channel value is selected from the medical image segmentation probability map to obtain the final medical image segmentation result.
[0065] During the inference phase, spectral priors can be learned. Instead of being directly output as the final medical image segmentation result, it serves as a dynamic frequency domain structure constraint during the backdiffusion denoising process, guiding the mask diffusion branches to generate more accurate mask prediction results.
[0066] Further testing was conducted on the model constructed using this method. Table 1 shows the performance validation results of different methods on the ISIC2018, BUSI, BraTS2021, and REFUGE2 datasets. This method achieved superior performance on multiple datasets, including ISIC2018, BUSI, BraTS2021, REFUGE2-Disc, and REFUGE2-Cup, demonstrating its good adaptability to different imaging modalities and target structures. On ISIC2018, this method achieved the highest Dice and IoU, reaching 91.6% and 84.5%, respectively. It also achieved optimal results on REFUGE2-Disc and REFUGE2-Cup datasets, with Dice and IoU of 95.5% and 88.1% for Disc segmentation and 86.2% and 76.6% for Cup segmentation, respectively. While our method doesn't achieve the highest scores across all metrics on the BUSI and BraTS2021 datasets, its overall performance remains competitive. Notably, it achieves the lowest HD95 value of 10.32 on BraTS2021, demonstrating advantages in boundary error control and spatial structure preservation. These results indicate that our method maintains stable performance across various medical image segmentation tasks and achieves good overall performance in terms of boundary accuracy and region overlap.
[0067] Table 1 Comparison of performance verification results of different methods ; Visual segmentation results as follows Figure 5 As shown, existing methods are prone to oversegmentation, undersegmentation, and boundary shifts to varying degrees when the lesion boundaries are complex, the target morphology is irregular, or the image contrast is low. For example, some methods misclassify the surrounding background region as the target region in skin lesion images, resulting in a prediction range that is significantly larger than the mask; some diffusion-based methods, while able to capture the main body region, have insufficiently smooth boundary contours, exhibiting local dilation or missing phenomena. In contrast, the segmented region obtained by this method is more complete, and the boundary fits the mask better, effectively suppressing background false detections and reducing target region missed detections.
[0068] In the BraTS2021 visualization results, such as Figure 6 As shown, existing methods have certain limitations in characterizing the spatial structure of tumor regions. Some results exhibit problems such as an increase in red false positive regions and insufficient coverage of green true regions, indicating that they are prone to oversegmentation or undersegmentation in complex brain structures and weakly defined regions. Our proposed method shows a higher degree of overlap between the prediction results and the mask, more stable boundary contours, and relatively fewer false positive regions, demonstrating good preservation of both the main tumor region and edge details. The results indicate that our method, through learnable spectral priors and bibranch diffusion co-modeling, enhances the model's ability to represent global structure and boundary details, thereby improving the completeness, boundary fit, and robustness of medical image segmentation.
[0069] Therefore, this invention employs a learnable spectral prior medical image segmentation method, replacing the fixed analytical spectrum with a learnable spectral prior to filter noise, artifacts, and acquisition disturbances, thus improving segmentation robustness. A mask-spectral bi-branch diffusion model is constructed to achieve synchronous evolution of the spectral prior and mask prediction, resolving the dynamic state mismatch problem. Spatial-spectral bidirectional modulation is introduced to dynamically fuse spatial details and frequency domain structure, strengthening boundary preservation and structural consistency. Structured spectral statistical constraints suppress the model's excessive dependence on local noise. Overall, this significantly improves the accuracy, stability, and generalization ability of medical image segmentation, adapting to various clinical segmentation scenarios.
[0070] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A medical image segmentation method with learnable spectral priors, characterized in that, This includes the model training phase and the model inference phase; The model training phase includes the following steps: S1. Obtain medical image samples and their corresponding segmentation labels, and perform normalization preprocessing on the medical image samples while extracting image conditional features. The segmentation labels are then converted into continuous mask representations. ; S2. Perform two-dimensional Fourier transform and spectral centering on the preprocessed image to obtain the analytical logarithmic amplitude spectrum; S3. Normalize and adjust the size of the analytical logarithmic amplitude spectrum to obtain the spectral anchor point. And construct a structured spectral statistical target based on the analytical logarithmic amplitude spectrum; S4. At the same diffusion time step, the continuous mask representation and the aforementioned spectral anchor point Noise is added separately to obtain the mask noise state during the training phase. and spectral noise status ; S5. Adjust the mask noise state during the training phase. Spectrum noise status Image conditional features The time-step embedded input dual-branch diffusion model is used to jointly predict mask noise and spectral noise; the dual-branch diffusion model includes a mask diffusion branch and a spectral diffusion branch; S6. Optimize the bi-branch diffusion model jointly based on the segmentation loss, spectral supervision loss, and diffusion denoising loss; The model inference phase includes the following steps: S7. Acquire medical images and perform preprocessing to obtain the medical images to be segmented; S8. Input the medical image to be segmented into the image condition feature extraction unit in the trained bi-branch diffusion model, and extract the image condition features of the medical image to be segmented according to step S1. S9. Initialize the mask noise state and the spectrum noise state; S10. Based on the image condition features, the dual-branch diffusion model performs joint back-diffusion denoising on the mask noise state and the spectral noise state; the mask diffusion branch generates mask prediction results from the mask noise state during the back-diffusion process, and the spectral diffusion branch generates learnable spectral priors from the spectral noise state during the back-diffusion process. S11. In the joint backdiffusion denoising process, the mask diffusion branch and the spectral diffusion branch are spatially and spectrally co-updated so that the learnable spectral prior dynamically guides the generation of mask prediction results during the backdiffusion process. S12. Based on the mask prediction results obtained after back diffusion denoising, generate a medical image segmentation probability map, and obtain the final medical image segmentation result based on the medical image segmentation probability map.
2. The medical image segmentation method with learnable spectral prior according to claim 1, characterized in that, In S1, image conditional features The extraction process is as follows: S101. Input the preprocessed medical image samples into the feature extraction network to extract multi-scale image features: ;in, Indicates the number of feature scales. Indicates the first Image features at various scales; S102. Perform channel alignment and spatial size alignment on image features at each scale: ; in, This represents a 1×1 convolution operation. Indicates an upsampling or downsampling operation. Indicates the aligned first Individual scale features; Perform global average pooling on each aligned scale feature: ; in, This indicates a global average pooling operation. Indicates the first A global description vector of each scale feature; S103. Concatenate multiple scale description vectors and input them into a multilayer perceptron. Obtain the scale weights through the Softmax function. ; in, This represents vector concatenation. This represents a multilayer perceptron. , Indicates the first The weights of each scale feature satisfy: ; S104. Weighted fusion of multi-scale features based on scale weights: ; Further processing using convolution, normalization, and activation functions yields image conditional features: ; in, This represents a 3×3 convolution operation. This indicates a normalization operation. Represents a non-linear activation function. Represents image conditional features.
3. The medical image segmentation method with learnable spectral prior according to claim 1, characterized in that, In S2, the formula for calculating the analytic logarithmic amplitude spectrum is: ; in, This represents the preprocessed image; Represents a two-dimensional Fourier transform. This indicates a spectrum centering operation. Indicates the coordinate position in the spectrum plane; Representing coordinates The analytic logarithmic amplitude spectrum at that location.
4. The medical image segmentation method with learnable spectral prior according to claim 3, characterized in that, In S3, the structured spectral statistics target is constructed in the following way: S31. Divide the centered spectral plane into... Radial frequency band and Each angular frequency band forms One radial-angular sub-band; S32, for the first The radial frequency band and the first Sub-bands formed by the angular frequency bands The average energy of the analytic logarithmic amplitude spectrum within this sub-band is calculated using the following formula: ; in, Indicates sub-band The number of spectral points included; S33. Combining the average energy values corresponding to all sub-bands into a structured spectral statistical objective: ; in, This represents the structured spectral statistics objective.
5. The medical image segmentation method with learnable spectral prior according to claim 4, characterized in that, In S4, the formula for adding noise to the continuous mask representation and spectral anchor points is: ; in, This indicates the un-noiseed diffusion start point, which includes a continuous mask representation. or spectral anchor ; Indicates diffusion time step The noise state under the following conditions Indicates noise scheduling parameters, Indicates Gaussian noise; In S5, the formula for jointly predicting mask noise and spectral noise is: ; in, This represents a two-branch diffusion model. Indicates the mask noise state. Indicates the spectral noise state. Represents image conditional features. Indicates the diffusion time step. This represents the predicted mask noise. This represents the predicted spectral noise. Based on the prediction results of the noise recovery mask and the learnable spectrum prior: ; ; in, This indicates the mask prediction result. This represents a learnable spectral prior.
6. The medical image segmentation method with learnable spectral prior according to claim 5, characterized in that, In S6, the total loss function of the dual-branch diffusion model is: ; in, Indicates the loss from partitioning. Indicates spectrum monitoring loss, This represents the diffusion denoising loss. , , These represent the weighting coefficients of the corresponding loss terms; The spectral supervision loss is: ; in, This represents the pixel-level spectral anchoring loss between the learnable spectral prior and the spectral anchor point. Indicates radial-angular energy band monitoring loss; and These represent the corresponding weights; The pixel-level spectral anchoring loss is: ; in, This represents a learnable spectral prior. Indicates the spectral anchor point; For learnable spectral priors Energy statistics are performed using the same radial-angular partitioning method as the analytical logarithmic amplitude spectrum to obtain the predicted structured spectral statistical representation. ; The radial-angular energy band supervision loss is: ; in, This represents the predicted structured spectral statistical representation obtained from learnable spectral priors. Represents the structured spectral statistics objective; The diffusion denoising loss is: ; in, and These represent the actual noise in the mask diffusion branch and the spectral diffusion branch, respectively. and These represent the noise predicted by the mask diffusion branch and the spectral diffusion branch, respectively. and These represent the weights of the mask diffusion denoising loss and the spectral diffusion denoising loss, respectively.
7. The medical image segmentation method with learnable spectral prior according to claim 1, characterized in that, In S5, the dual-branch diffusion model further includes an image conditional feature extraction unit and an uncertainty calibration bidirectional modulation unit; The image conditional feature extraction unit is used to extract image conditional features from the medical image to be segmented; The uncertainty calibration bidirectional modulation unit is used to adjust the information interaction between the mask diffusion branch and the spectral diffusion branch during the back diffusion process.
8. The medical image segmentation method with learnable spectral prior according to claim 1, characterized in that, In S10, the learnable spectral prior is the denoised spectral representation generated by the spectral diffusion branch during the back-diffusion denoising process; The learnable spectral prior is not a fixed analytical spectrum obtained by directly performing a Fourier transform on the medical image to be segmented, but is updated synchronously with the mask prediction results during the backdiffusion process and is used to provide frequency domain structural constraints to the mask diffusion branch.
9. The medical image segmentation method with learnable spectral prior according to claim 1, characterized in that, In S10, the joint back-diffusion denoising process includes: S1001, The mask diffusion branch and the spectral diffusion branch are updated synchronously at the same or corresponding diffusion time steps; S1002. At each diffusion time step, the dual-branch diffusion model predicts mask noise and spectral noise respectively based on the current mask noise state, the current spectral noise state, image condition features, and time step embedding. S1003. Recover the mask prediction state at the current diffusion time step based on the predicted mask noise, and recover the learnable spectral prior at the current diffusion time step based on the predicted spectral noise.
10. The medical image segmentation method with learnable spectral prior according to claim 1, characterized in that, In S11, the spatial-spectrum collaborative update includes: S1101. Obtain the mask branch characteristics and spectral branch characteristics of the current back diffusion stage; S1102. Calculate the mask branch reliability index based on the mask branch features, and calculate the spectrum branch reliability index based on the spectrum branch features; the mask branch reliability index and the spectrum branch reliability index are obtained by performing context encoding, global pooling, and nonlinear mapping on the mask branch features and the spectrum branch features, respectively; S1103. Based on the mask branch reliability index and the spectrum branch reliability index, calculate the first gate coefficient of the mask diffusion branch modulated by the spectrum diffusion branch, and the second gate coefficient of the spectrum diffusion branch modulated by the mask diffusion branch; the first gate coefficient is used to control the strength of the frequency domain structure information transmitted from the spectrum diffusion branch to the mask diffusion branch; the second gate coefficient is used to control the strength of the spatial semantic information transmitted from the mask diffusion branch to the spectrum diffusion branch. S1104. Based on the first gating coefficient, the spectral branch features are modulated to the mask diffusion branch, and based on the second gating coefficient, the mask branch features are modulated to the spectral diffusion branch, so as to achieve bidirectional modulation of uncertainty calibration of the mask diffusion branch and the spectral diffusion branch.