Blurred image segmentation method based on mixed diffusion model

By introducing a hybrid diffusion model and a wavelet convolution encoder into image segmentation, the problems of the prior art in blurred images and small-objective segmentation are solved, and efficient and accurate image segmentation effect is achieved.

CN120163829APending Publication Date: 2025-06-17CHONGQING UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510217541.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Existing image segmentation algorithms have difficulties in processing blurred images, background noise influence, large-target segmentation and small-target segmentation, especially in subtle features and small-target segmentation, which makes it difficult to achieve high-precision segmentation.

Method used

A fuzzy image segmentation method based on a hybrid diffusion model is proposed, using the HAD-Net model, which consists of a diffusion model and a segmentation model. Through the diffusion model, the underlying distribution of data is learned and denoised, and combined with a wavelet convolution encoder to extract high and low frequency information to achieve accurate image segmentation.

Benefits of technology

This method can effectively process blurred images, reduce training costs, and improve segmentation accuracy, especially when dealing with segmentation targets with large morphological distribution and size gaps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163829A_ABST
    Figure CN120163829A_ABST
Patent Text Reader

Abstract

The invention discloses a blurred image segmentation method based on a mixed diffusion model, and relates to the technical field of artificial intelligence image processing and computer vision. The invention provides a novel fast, efficient and accurate hybrid denoising segmentation model HAD-Net. The HAD-Net is an end-to-end architecture composed of a diffusion model and a segmentation model; wherein underlying distribution of data is learned by using a diffusion model, and noise distribution is fully captured through a Markov chain, so that an image is denoised in a reasoning stage; a clear denoised image is used as prior input to a segmentation model, a feature encoder containing wavelet convolution is designed, and high and low frequency information is fully extracted through wavelet convolution to enable a prior image to realize guide segmentation; reasonable construction of the overall architecture greatly reduces training cost, and reasonable network depth and internal module design solve the problem of low segmentation precision caused by too large difference between morphological distribution and size of segmentation targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing and computer vision in artificial intelligence, and specifically to a fuzzy image segmentation method based on a hybrid diffusion model. Background Art

[0002] Image processing is a collection of techniques and methods for operating on images to enhance their visibility or extract useful information, and has wide applications, including medical imaging, remote sensing, computer vision, photography, entertainment and other fields. Its process usually includes image acquisition, preprocessing, feature extraction, representation and description, recognition and classification, and postprocessing. Subsequently, a specific task is completed through the designed model structure.

[0003] Semantic segmentation is a technique in image processing and computer vision, aiming to divide an image into several regions and assign a semantic label to each pixel to represent the object category to which the pixel belongs. It has wide applications, including autonomous driving, medical image analysis, remote sensing image processing, and intelligent monitoring. Its process usually includes steps such as image preprocessing, feature extraction, classification and annotation. Using a deep learning training model can extract multi-scale features while maintaining high resolution, achieving fine pixel-level classification. Semantic segmentation technology plays an important role in applications such as road scene understanding, tumor detection, and ground object classification.

[0004] In recent years, the rapid development of deep learning has shown powerful performance in the field of image segmentation. Deep learning models can achieve high-precision segmentation while taking into account time and space efficiency by virtue of automatic feature learning, convenient operation, and high efficiency. The classic fully convolutional network (FCN) was proposed by Long et al. in 2015. Its main feature is to accept input images of any size and generate output images of the same size, assigning semantic labels to each pixel, thus achieving end-to-end pixel-level semantic segmentation. However, due to multiple pooling operations, FCN loses attention to image details. Subsequently, to solve the problem of accuracy loss, U-Net was proposed.

[0005] Due to its symmetric encoder-decoder structure, which provides a connection path for the original features to each layer of the decoder, the overall network can fully retain the original detailed information, resulting in a more accurate segmentation result. However, the U-Net network still has deficiencies. Due to its own structural limitations, it performs poorly in special scenarios or for special tasks. Therefore, by leveraging the global feature fitting ability of the Transformer structure and combining it with U-Net, TransUNet was designed. By combining the ability of convolutional neural networks to efficiently extract local features with the global attention perception ability of the Transformer, a more accurate segmentation effect is achieved. For different scenarios, using different convolutional kernels to adapt to the distribution patterns of different objects can also effectively improve the segmentation effect. For example, CE-Net combines dilated convolution as a context feature extraction module, aiming to fully focus on the context feature information between objects, enabling the model to learn deeper features. While TP-Net uses dilated convolution as the processing module for the final image output. By using different-sized dilated convolution kernels to integrate information from multi-scale features, the robustness of segmentation is enhanced.

[0006] In recent years, deep learning has made remarkable progress in the field of image segmentation. Deep learning models can automatically learn features, are easy to operate, and have high performance. They can obtain high-precision segmentation results while taking into account both time and space efficiency. The classic U-Net network was designed by Ronneberger et al., adopting a symmetric encoder-decoder structure and using skip connections for the transfer of original features, thus improving the network's attention to fine features. These characteristics enable it to perform well in the task of retinal vessel segmentation. However, due to the extensive distribution of thin vessels and small target vessels in fundus images, accurately capturing vessel features remains difficult. In addition, the information loss in the pooling operation exacerbates the difficulty of fine vessel segmentation. To address these problems, Gu et al. designed a multi-kernel pooling fusion multi-attribute convolution module to fully integrate context information for vessel feature extraction, thus solving the problem of information loss caused by pooling. Yin et al. used multi-source image inputs to ensure the transfer of retinal vessel features, utilized feature fusion at different scales to provide the original feature information of vessels for each layer of the network, and to some extent supplemented the information lost in the pooling operation. In addition, they used the Hessian matrix to obtain weak information in vessels, which although generates some noise, also enhances the attention to vessels. Li et al. used dual CNN and RNN encoders for feature extraction and fusion of fine vessel features, and utilized multiple context information fusion modules to fully extract vessel features, alleviating the problem of information loss. In summary, despite many improvements, image segmentation, especially in the segmentation of fine features and small targets, still faces many challenges and technical bottlenecks.

[0007] (1) Most segmentation algorithms ignore the influence of background noise, and the images obtained in most scenarios cannot clearly display the target. As the size of the target decreases, the segmentation difficulty increases continuously, resulting in poor segmentation effects in complex scenarios.

[0008] (2) Since there are many small target objects in the actual scenario, and there are usually phenomena such as blurred boundaries. The existing methods at present cannot effectively extract object features to achieve accurate segmentation. For example, the targets in large forests, the target images taken by remote sensing satellites, etc.

[0009] (3) Frequency features are an important information that cannot be ignored. Effectively capturing the information in different frequency domains to provide effective guidance for the model is an effective means to improve the segmentation accuracy. However, the existing methods are prone to ignoring the connection between different frequency information, and a large amount of information such as texture details is not fully explored. This results in insufficient extraction of detail information by the model.

[0010] Therefore, a new solution needs to be proposed for the above problems. Summary of the Invention

[0011] The purpose of the present invention is to provide a fuzzy image segmentation method based on a hybrid diffusion model, which is used to segment fuzzy images, train a model based on deep learning and can achieve accurate segmentation effects, so as to solve the existing technical problems proposed in the background technology.

[0012] To achieve the above purpose, the present invention provides the following technical solutions: A fuzzy image segmentation method based on a hybrid diffusion model, at least including the following steps:

[0013] S1: Batch preprocess the collected image data;

[0014] S2: Divide the data into a dataset without cross;

[0015] S3: Design a hybrid loss function for accurate segmentation;

[0016] S4: Design a hybrid diffusion model, the hybrid diffusion model includes but is not limited to a diffusion model and a segmentation model, and the hybrid diffusion model is HAD-Net;

[0017] S5: Train HAD-Net and verify the segmentation metrics. In this process, train through the joint loss function formulated in S3, use the gradient backpropagation of the neural network in the training process to update the model weights, judge whether the model is better according to the segmentation effect of the model on the validation set after training, update the saved model weights, and finally, after the model training is completed, evaluate the segmentation effect on the test set;

[0018] S6: Use the trained HAD-Net for fuzzy image segmentation.

[0019] Furthermore, the image data in S1 includes but is not limited to OCTA images in bmp format, and the data preprocessing includes but is not limited to normalizing the image data for input into the model.

[0020] Furthermore, S2 at least includes the following steps: randomly shuffle and divide the image data, and divide it into training, validation, and test sets in a ratio of 8:1:1, where no image data can cross between these sets.

[0021] Furthermore, the loss function in S3 is set as follows:

[0022] Select the mixed loss of DiceLoss and BCELoss as the final loss function;

[0023] The weight hyperparameters of the DiceLoss and BCELoss are λ D and λ M , and their sum is 1;

[0024] MSE Loss:

[0025]

[0026] where N represents the total number of pixel points; y i is the true label, the value of the i-th pixel point, is the value predicted by the model;

[0027] DiceLoss:

[0028]

[0029] BCELoss:

[0030]

[0031] Final Loss:

[0032] L Final = λL MSE + L BCE + L clDice (4)

[0033] where C represents the total number of categories; represents the c-th category in the binary GroundTruth, and the i-th pixel corresponds to the category value; Represents the predicted probability value of the corresponding category; ε = 1×10-5, which represents the smoothing exponent. ε is used to prevent the denominator prediction from being 0, thus avoiding some extreme situations. In addition, it can also play a role in smoothing the loss and gradient.

[0034] Furthermore, the application process of the SHAD-Net at least includes the following steps:

[0035] Let Represents the vascular segmentation mask, with a spatial resolution of where the value of M′ is restricted within the set {0,1}; Given a random Gaussian noise image at the specified time step t, this image follows the standard normal distribution The diffusion model predicts the noise level of images with different degrees of damage, and calculates the mean square error between the predicted noise and the Gaussian noise;

[0036] When the training of the diffusion model is basically stable, activate the segmentation model. Let Represents the original OCTA image, with a spatial resolution of H×W and containing C channels. The pixel values in M are distributed within the range [0,1]. Use M as the original conditional image at time step T, and perform T steps of reverse denoising. In order to generate a more stable denoised image and reduce the impact of a single outlier, perform five complete T-step denoisings on the original image, and calculate the average image as the final prior image;

[0037] The segmentation model uses the clear vascular image generated from the diffusion model as a prior, combines the information of the original OCTA image, and performs the vascular segmentation task through supervised learning. The goal of the segmentation model is to compare the ground truth label with the segmentation result predicted by the model, thereby guiding the network for efficient training;

[0038] At the same time, the gradient of the forward noise feature learning in the diffusion model will be suppressed and gradually decreased until the training ends, so as to obtain a stable output.

[0039] Furthermore, using HAD-Net for image denoising and point repair at least includes the following steps:

[0040] Use the original image as the generation prior, and redefine the segmentation task as a denoising problem;

[0041] The forward process of the diffusion model gradually adds random Gaussian noise within t time steps, and gradually converts the original segmentation mask into noise according to the Markov chain, which results in a mixed data distribution. Define the forward process of the diffusion model as:

[0042]

[0043] α tis a scheduling factor that controls the variance of the noise, and as the time step t increases, the noise gradually increases; the goal of the diffusion model is to gradually transform the image into noise through this forward process, thereby achieving variational inference x0, …, x in the Markov chain at time step t T ;

[0044] α t = 1 - β t , where β t is a parameter that controls the noise intensity at each step. By reparameterizing β t , the noise addition process can be controlled more flexibly;

[0045] x t is the image at the current time step t; (1 - α t )I θ is the noise variance;

[0046] Use the improved U-Net network D(·) to generate the prior distribution layer D(x). Therefore, the entire forward process is defined as:

[0047]

[0048] where q(x 1:T |x0, D(x)) represents the conditional probability density of the entire diffusion process, and x 1:T represents the state sequence from t = 1 to t = T, that is, the image sequence from the initial image x0 to the image at the final time step T;

[0049] This part shows that the joint probability q(x 1:T |x0, D(x)) of the diffusion process is decomposed into a product of conditional probabilities;

[0050] This means that the noise addition at each step depends not only on the image state x t-1 at the previous moment but also on the generated prior distribution D(x), thereby more effectively capturing structural features such as blood vessels. It should be noted that the improved U-Net network designs a parametric Fourier transform control unit in the U-Net encoder;

[0051] The diffusion model D(·) learns the relationship between the data distribution of the segmentation mask and the noise during the entire training process, effectively capturing the features of blood vessels while modeling the noise features;

[0052] The denoising process learns the inverse distribution p(x t-1 |x t ) by minimizing the KL divergence between the forward and backward distributions at all time steps through the neural network θ, defined as:

[0053]

[0054] Among them, μ θ (x t , t) and ∑ θ (x t , t) are parameters learned by the neural network θ, used to predict the recovery process from x t to x t-1 ;

[0055] Use the parameterized diffusion model D(·) to predict the noise in the image at time x t and recover the image x t-1 at that time. The denoising process is defined as:

[0056]

[0057] Use the original OCTA image as the data before the T-step of denoising, and achieve denoising by adjusting the length of the T-step. The complete reverse process is defined as:

[0058]

[0059] Among them, p(M) represents the prior probability;

[0060] After multiple time-step transformations, the diffusion model removes the noise from the original image, effectively highlighting the vascular structure for recovery. The specific quantification of the diffusion model used in the present invention is a U-Net neural network, and this model is used to learn the underlying data distribution.

[0061] Furthermore, in order to enable the model to better perceive frequency information, a parametric Fourier transform control unit (PFT) is designed to accurately capture potential vascular features, highlight the gray-scale changes in the OCTA image, and enhance the overall dynamic range;

[0062] The parametric Fourier transform control unit is used to perform the fast Fourier transform and add learnable parameters to adjust the frequency-domain features of the image;

[0063] Perform the fast Fourier transform F(.) on the domain of the image to obtain the real part R and the imaginary part I, which respectively contain the high-frequency domain information, low-frequency domain information, and edge detail information of the image;

[0064] In order to better control the edge texture information and contrast information, learnable parametric factors α and β are introduced to adjust the amplitude and phase after the transformation, respectively;

[0065] Subsequently, x' is obtained by performing an inverse transformation using an imaginary number i, expressed as:

[0066] R, I = F(x) ∈ RC×H×W (10)

[0067] x′ = R·α + i·(I·β) (11)

[0068] Similarly, after the inverse transformation Embed a parametric Fourier transform control unit in the encoder of the diffusion model of HAD-Net to adjust the frequency domain information of the extracted data features, thereby enhancing the overall contrast of the image;

[0069] First, apply the convolutional layer Conv in to increase the number of channel features. After transformation, use Conv out layer to smooth the output features. The entire encoder process is defined as:

[0070] F out = Conv out (PFT(Conv in (F in ))) (12)

[0071] where Conv represents a 3×3 convolutional layer with a normalization layer and ReLU activation, F in represents the input features, and F out represents the output features.

[0072] Furthermore, in some regions, thin blood vessels are still mixed with background noise. Instead of directly applying a filtering method to eliminate these noise effects, the spatial representation of small blood vessels is retained;

[0073] To separate small blood vessels from the mixed background, a WTEncoder guided by wavelet convolution is designed. In addition, to retain complex blood vessel details, the use of downsampling layers is minimized to maintain the spatial scale;

[0074] Given an input image Use four filters to extract high-frequency and low-frequency channels. The filters are defined as:

[0075]

[0076]

[0077] f LL represents a low-pass filter for extracting low-frequency information;

[0078] while the other three filters form a group of high-pass filters for extracting high-frequency information;

[0079] The four filter banks generate four different filter channel responses through a convolutional Conv operation with a stride of 2, denoted as

[0080] Subsequently, the deep convolution DC is applied to the four generated frequency-domain channels to extract features and strengthen the interaction between different frequency domains, thereby obtaining the output Y out , and this process is defined as:

[0081] Y out = DC(Conv(X, {f LL , f LH , f HL , f HH})) (13)

[0082] Since the four filters form an orthonormal basis, a transposed convolution with a stride of 2 is applied to Y out to perform the operation and obtain the restored image

[0083] Meanwhile, a deep convolution is performed on the original feature vector to extract its information representation, and the frequency-domain features obtained by the inverse transform are used to enhance it. The final output feature is F out , which is expressed as:

[0084] F out = DC(X) + TC(Y out , {f LL , f LH , f HL , f HH}) (14)

[0085] where TC represents the transposed convolution, and the final output is the sum of the original feature convolution and the wavelet convolution;

[0086] After that, BatchNorm and GELU activation are performed.

[0087] To fully capture the edge and frequency information, a two-layer wavelet convolution encoder is designed;

[0088] First, a 7×7 sliding convolution kernel is used to extract information from the local area, and a 1×1 convolution is used to provide a residual connection for the original information;

[0089] The wavelet transform allows iterative transformation to generate higher-dimensional and lower-scale frequency-domain information;

[0090] However, in view of the existing clear prior vascular information, a two-step low-order wavelet transform convolution is designed to further fit the features instead of calculating redundant frequency-domain information;

[0091] The single-layer wavelet convolution is used for the initial enhancement of low-frequency and texture features;

[0092] After the second - layer wavelet convolution, recalculate the correlation between frequency domains, gradually guiding the model to focus on clear main blood vessels and capillaries;

[0093] Finally, use a trainable enhancement factor to globally guide the feature extraction path and linearly adjust the correlation between the residual path and the main path.

[0094] Compared with the prior art, the beneficial effects of the present invention are:

[0095] Aiming at the problems of high training cost, narrow application range, large shape gap, blurred boundary and small target segmentation objects for different segmentation targets in most segmentation models, the present invention proposes a new fast, efficient and accurate hybrid denoising segmentation model HAD - Net; HAD - Net is an end - to - end architecture composed of a diffusion model and a segmentation model; among them, the diffusion model is used to learn the underlying distribution of data, and the noise distribution is fully captured through the Markov chain to denoise the image in the inference stage; use the clear denoised image as the prior input into the segmentation model, design a feature encoder (EncoderWT) containing wavelet convolution, and fully extract high - and low - frequency information through wavelet convolution to enable the prior image to achieve guided segmentation; the reasonable construction of the overall architecture greatly reduces the training cost, and the reasonable network depth and internal module design solve the problem of low segmentation accuracy caused by large morphological distribution and size gap of segmentation targets;

[0096] To demonstrate the excellent performance of the proposed segmentation model in different segmentation tasks, especially in tubular segmentation tasks, multiple datasets are used for training, expanding the application range of the model;

[0097] Compared with other segmentation algorithms, HAD - Net is an innovative network. It re - defines the diffusion model generation task, converts the generation of segmentation masks into image denoising by controlling the time step, which enables it to effectively process blurred images. Integrate this diffusion model and finally adopt a lightweight segmentation model to fully fuse conditional encoding information, so as to achieve the accurate segmentation of target objects. The entire model has strong feature fitting ability and feature scale reduction ability. Training on this basis can effectively back - propagate gradients to reduce the loss value and can be effectively applied to various scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0098] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0099] Figure 1Flowchart of the overall image segmentation algorithm provided by the present invention;

[0100] Figure 2 Schematic diagram of the HDA-Net network model provided by the present invention;

[0101] Figure 3 Schematic diagram of the EncoderWT module provided by the present invention;

[0102] Figure 4 Schematic diagram of the segmentation model framework provided by the present invention;

[0103] Figure 5 Schematic diagram of the denoising framework provided by the present invention;

[0104] Figure 6 Visual comparison chart of segmentation masks of different models provided by the present invention on three datasets;

[0105] Figure 7 Visual comparison chart of segmentation masks of different models provided by the present invention on three datasets;

[0106] Figure 8 Visualization diagram of denoising of the diffusion model provided by the present invention. Detailed implementation manners

[0107] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.

[0108] Please refer to Figures 1 - 5 , a fuzzy image segmentation method based on a hybrid diffusion model, at least including the following steps:

[0109] S1: Batch process the collected image data for preprocessing;

[0110] S2: Divide the data into non-overlapping datasets;

[0111] S3: Design a hybrid loss function for accurate segmentation;

[0112] S4: Design a hybrid diffusion model, the hybrid diffusion model includes but is not limited to a diffusion model and a segmentation model, and the hybrid diffusion model is HAD-Net;

[0113] S5: Train HAD-Net and conduct segmentation metric verification. During this process, train using the combined loss function defined in S3, utilize the gradient backpropagation of the neural network during training to update the model weights, determine whether the model is better based on the segmentation effect of the model on the validation set after training, update the saved model weights, and finally, after the model training is completed, evaluate the segmentation effect on the test set;

[0114] To objectively evaluate the proposed HDArch, seven commonly used evaluation metrics are adopted, including Dice coefficient (Dice), Jaccard index (JAC), accuracy (Acc), sensitivity (Sen), specificity (Spe), balanced accuracy (BACC), and precision (Pre). The probability map generated by the network has a threshold of 0.5, and pixels greater than 0.5 are predicted as blood vessels, and vice versa are predicted as the background. The metric definitions are as follows:

[0115]

[0116] Taking the ground truth as the benchmark, correctly classified retinal blood vessels are classified as the TP class, those that are actually retinal blood vessels but are classified as the background by the network are classified as the FN class, correctly classified background is classified as the TN class, and misclassified background classified as retinal blood vessels is classified as the FP class.

[0117] S6: Use the trained HAD-Net for fuzzy image segmentation.

[0118] The image data in S1 includes but is not limited to OCTA images in bmp format, and data preprocessing includes but is not limited to normalizing the image data for input into the model.

[0119] S2 at least includes the following steps: randomly shuffle and divide the image data, and divide it into training, validation, and test sets in a ratio of 8:1:1, where no image data can cross between these sets.

[0120] The loss function in S3 is set as follows:

[0121] To deal with the characteristics of segmentation targets with large scale variations in different tasks (such as houses and pedestrians in remote sensing images), and the existence of small targets in some tasks (such as biomedical lesions), the target area varies greatly and is complex; therefore, a mixed loss of DiceLoss and BCELoss is selected as the final loss function;

[0122] The weight hyperparameters of the DiceLoss and BCELoss are λ D and λ M, the sum is 1, and it is set to [0.5, 0.5] in the experiment to make the model training more balanced and stable;

[0123] MSE Loss:

[0124]

[0125] Among them, N represents the total number of pixel points; y i is the true label, the value of the i-th pixel point, is the value predicted by the model;

[0126] DiceLoss:

[0127]

[0128] BCELoss:

[0129]

[0130] Final Loss:

[0131] L Final = λL MSE + L BCE + L clDice (4)

[0132] Among them, C represents the total number of categories; represents the value of the c-th category corresponding to the i-th pixel in the binary GroundTruth; represents the predicted probability value of the corresponding category; ε = 1×10-5, which represents the smoothing exponent. ε is used to prevent the denominator prediction from being 0, thus avoiding some extreme situations. In addition, it can also smooth the loss and gradient.

[0133] The application process of SHAD-Net includes at least the following steps:

[0134] Let represent the vascular segmentation mask, and the spatial resolution is where the value of M′ is restricted within the set {0, 1}; given a random Gaussian noise image at the specified time step t, this image follows the standard normal distribution The diffusion model predicts the noise level of images with different degrees of damage, and calculates the mean square error between the predicted noise and the Gaussian noise;

[0135] When the diffusion model training is basically stable, activate the segmentation model, and let Denote the original OCTA image with a spatial resolution of H×W and containing C channels, where the pixel values in M are distributed in the range [0, 1]. Use M as the original conditional image at time step T and perform T steps of reverse denoising. To generate a more stable denoised image and reduce the impact of a single outlier, perform five complete T-step denoisings on the original image and calculate the average image as the final prior image;

[0136] The segmentation model uses the clear vascular image generated from the diffusion model as a prior, combines the original OCTA image information, and performs the vascular segmentation task through supervised learning. The goal of the segmentation model is to compare the ground truth labels with the segmentation results predicted by the model, thereby guiding the network for efficient training;

[0137] Meanwhile, the gradients of the forward noise feature learning in the diffusion model will be suppressed and gradually reduced until the end of training, thus obtaining a stable output.

[0138] The image denoising and point repair using HAD-Net include at least the following steps:

[0139] Mainstream methods refine the diffusion model by prior encoding of the image to generate the final segmentation mask. In contrast, the present invention does not predict the segmentation mask through the diffusion model, but uses the original image as the generation prior and redefines the segmentation task as a denoising problem;

[0140] The forward process of the diffusion model gradually adds random Gaussian noise within t time steps, and gradually converts the original segmentation mask into noise according to the Markov chain, which results in a mixed data distribution. Define the forward process of the diffusion model as:

[0141]

[0142] α t is a scheduling factor that controls the variance of the noise. As the time step t increases, the noise gradually increases; the goal of the diffusion model is to gradually turn the image into noise through this forward process, thereby realizing variational inference x0,…,x in the Markov chain of t time steps T ;

[0143] α t = 1 - β t where β t is a parameter that controls the noise intensity at each step. By reparameterizing β t , the process of adding noise can be controlled more flexibly;

[0144] x t is the image at the current time step t; (1 - α t )I θ is the noise variance;

[0145] Use the improved U-Net network D(·) to generate the prior distribution layer D(x), so the entire forward process is defined as:

[0146]

[0147] where q(x 1:T |x0, D(x)) represents the conditional probability density of the entire diffusion process, and x 1:T represents the state sequence from t = 1 to t = T, that is, the image sequence from the initial image x0 to the image at the final time step T;

[0148] This part shows that the joint probability q(x 1:T |x0, D(x)) of the diffusion process is decomposed into a product of conditional probabilities;

[0149] This means that the noise addition at each step (in the forward process) depends not only on the image state x t-1 at the previous moment, but also on the generated prior distribution D(x) to more effectively capture structural features such as blood vessels. It should be noted that the improved U-Net network is to design a parametric Fourier transform control unit in the U-Net encoder;

[0150] The diffusion model D(·) learns the relationship between the data distribution of the segmentation mask and the noise during the entire training process, and effectively captures the features of blood vessels while modeling the noise features;

[0151] The denoising process learns the inverse distribution p(x t-1 |x t ) by minimizing the KL divergence of the forward and backward distributions at all time steps through the neural network θ, defined as:

[0152]

[0153] where μ θ (x t , t) and ∑ θ (x t , t) are the parameters learned by the neural network θ, used to predict the recovery process from x t to x t-1 ;

[0154] Use the parametric diffusion model D(·) to predict the noise in the image at time x t and recover the image x t-1 at that moment. The denoising process is defined as:

[0155]

[0156] Use the original OCTA image as the data before the denoising T step, and achieve denoising by adjusting the length of the T step. The complete reverse process is defined as:

[0157]

[0158] where p(M) represents the prior probability;

[0159] After multiple time-step transformations, the diffusion model removes noise from the original image, effectively highlighting the vascular structure for recovery.

[0160] Texture information and edge changes are crucial for accurate vascular segmentation. However, OCTA images usually exhibit a relatively concentrated gray-value distribution and low overall contrast, making accurate vascular segmentation challenging. The Fast Fourier Transform (FFT) updates global information through spectral-domain points, where the real and imaginary parts contain rich texture and edge structure elements. To enable the model to better perceive frequency information, a parametric Fourier transform control unit (PFT) is designed to accurately capture potential vascular features, highlight gray-scale changes in OCTA images, and enhance the overall dynamic range;

[0161] The parametric Fourier transform control unit is used to perform the Fast Fourier Transform and add learnable parameters to adjust the frequency-domain characteristics of the image;

[0162] For the image perform the Fast Fourier Transform F(.) in the domain to obtain the real part R and the imaginary part I, which contain the high-frequency domain information, low-frequency domain information, and edge detail information of the image respectively;

[0163] To better control the edge texture information and contrast information, learnable parametric factors α and β are introduced to adjust the amplitude and phase after the transformation respectively;

[0164] Subsequently, x′ is obtained by performing the inverse transform using the imaginary number i, expressed as:

[0165] R, I = F(x) ∈ R C×H×W (10)

[0166] x′ = R·α + i·(I·β) (11)

[0167] Similarly, after the inverse transform Embed the parametric Fourier transform control unit in the encoder of the diffusion model of HAD-Net to adjust the frequency-domain information of the extracted data features, thereby enhancing the overall contrast of the image;

[0168] First, apply the convolutional layer Conv in to increase the number of channel features. After transformation, use Convout A layer is used to smooth the output features, and the entire encoder process is defined as:

[0169] F out = Conv out (PFT(Conv in (F in ))) (12)

[0170] where Conv represents a 3×3 convolutional layer with a normalization layer and ReLU activation, F in represents the input features, and F out represents the output features.

[0171] The denoised OCTA image effectively highlights blood vessels, and the background appears as low-frequency and discrete noise. However, in some areas, thin blood vessels are still mixed with background noise. Instead of directly applying a filtering method to eliminate these noise effects, the spatial representation of small blood vessels is retained;

[0172] To separate small blood vessels from the mixed background, a WTEncoder guided by wavelet convolution is designed. In addition, to retain complex blood vessel details, the use of downsampling layers is minimized to maintain the spatial scale;

[0173] Traditional convolution is constrained by spatial invariance. For example, when using a 3×3 convolutional kernel for sliding window feature extraction, the receptive field is limited to a local range. However, due to the complex topology and slender structure of blood vessels, a smaller convolutional kernel cannot capture discrete features, while a larger convolutional kernel will result in computational redundancy and accuracy loss. To solve the problem of the changing morphological distribution of blood vessels, some methods increase the receptive field in the feature extraction layer by designing dilated convolutions with different dilation rates, while extracting multi-scale feature information. This method ignores the continuity of blood vessel topology and is difficult to capture complex morphological changes. Linearly adjusting the relative position of the convolutional kernel represents a significant advancement in retinal blood vessel segmentation, effectively capturing the branch features of different blood vessels. However, a large amount of noise in the OCTA image will cause the convolutional kernel to shift.

[0174] The denoised image sampled from the diffusion model retains a large amount of low-frequency details, but some blood vessel details are suppressed by the reverse process. To more effectively extract the fine blood vessel details retained in the low-frequency components and enable the discriminator to distinguish more effective blood vessel pixels, we introduce wavelet convolution.

[0175] Given an input image Four filters are used to extract high-frequency and low-frequency channels, and the filters are defined as:

[0176]

[0177] fLL represents a low - pass filter for extracting low - frequency information;

[0178] while the other three filters form a set of high - pass filters for extracting high - frequency information;

[0179] The four filter banks generate four different filter channel responses through a convolution Conv operation with a stride of 2, denoted as

[0180] Subsequently, depth - wise convolution DC is applied to the four generated frequency - domain channels to extract features and strengthen the interaction between different frequency domains, thereby obtaining the output Y out , and this process is defined as:

[0181] Y out = DC(Conv(X,{f LL ,f LH ,f HL ,f HH})) (13)

[0182] Since the four filters form an orthonormal basis, a transposed convolution with a stride of 2 is applied to Y out to perform an operation to obtain the restored image

[0183] Meanwhile, depth - wise convolution is performed on the original feature vector to extract its information representation, and the frequency - domain features obtained by inverse transformation are used to enhance it. The final output feature is F out , denoted as:

[0184] F out = DC(X)+TC(Y out ,{f LL ,f LH ,f HL ,f HH}) (14)

[0185] where TC represents transposed convolution, and the final output is the sum of the original feature convolution and the wavelet convolution;

[0186] After that, BatchNorm and GELU activation are performed.

[0187] To fully capture edge and frequency information, a two - layer wavelet convolution encoder is designed;

[0188] First, a 7×7 sliding convolution kernel is used to extract information from the local area, and 1×1 convolution is used to provide a residual connection for the original information;

[0189] Wavelet transform allows iterative transformation to generate higher - dimensional and lower - scale frequency - domain information;

[0190] However, in view of the existing clear prior vascular information, a two-step low-order wavelet transform convolution is designed to further fit the features instead of calculating redundant frequency-domain information;

[0191] The single-layer wavelet convolution is used for the initial enhancement of low-frequency and texture features;

[0192] After the second-layer wavelet convolution, the correlation between frequency domains is recalculated, gradually guiding the model to focus on clear main blood vessels and capillaries;

[0193] Finally, a trainable enhancement factor is used to globally guide the feature extraction path and linearly adjust the correlation between the residual path and the main path.

[0194] Please refer to Figures 6 - 8 , where the HDA-Net proposed in the present invention is trained and tested on three public datasets: OCTA-6M, OCTA-3M, and ROSE-1. The performance of HDA-Net is compared with other existing models.

[0195] OCTA6M: OCTA6M has a 6mm field of view and is used for macular imaging using a commercial 70kHz spectral-domain OCT system with a central wavelength of 840nm. The data we used is the maximum projection of OCTA between the inner limiting membrane (ILM) and Bruch's membrane (BM). OCTA-6M contains 300 subjects (NO.10001 - NO.10300), among which 180 are in the training group (NO.1000 - NO.10300), 20 are in the validation group (NO.1181 - NO.10200), and 100 are in the test group (NO.1201 - NO.10300).

[0196] OCTA3M: OCTA3M has a 3mm field of view and is used for macular imaging using a commercial 70kHz spectral-domain OCT system with a central wavelength of 840nm. The data we used is the maximum projection of OCTA between the inner limiting membrane (ILM) and Bruch's membrane (BM). OCTA-3M contains 200 subjects (NO.10301 - NO.10500), including 140 in the training set (NO.10301 - NO.10440), 10 in the validation set (NO.01441 - NO.10450), and 50 in the test set (NO.10451 - NO.10500)

[0197] ROSE-1: The ROSE-1 dataset comes from 39 subjects provided by the Cixi Institute of Biomedical Engineering, Ningbo Institute of Materials Technology and Engineering, including 26 Alzheimer's disease patients and 13 healthy individuals. It contains 117 images obtained by an RTVue XR Avanti SD-OCT system equipped with AngioVue software. SVC, DCV, and SVC+DVC angiographies were performed in a 3×3 area centered on the fovea with a diameter of 0.6 mm - 2.5 mm. We used a total of 39 superficial vascular complex (SVC) angiography images, with 27 in the training set, 3 in the validation set, and 9 in the test set.

[0198] Comparisons with state-of-the-art methods, including 7 specialized retinal vessel segmentation models and 3 diffusion segmentation models, are shown in Tables 1, 2, and 3 for the three datasets. There are 7 evaluation metrics: Dice, JAC, Acc, Sen, Spe, BACC, and Pre.

[0199] Table 1 Comparison results on the OCTA-6M dataset

[0200]

[0201]

[0202] Table 2 Comparison results on the OCTA-3M dataset

[0203]

[0204] Table 3 Comparison results on the ROSE-1 dataset

[0205]

[0206]

[0207] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, in any aspect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed by the present invention. Any reference signs in the claims should not be construed as limiting the claimed rights.

Claims

1. A fuzzy image segmentation method based on a mixed diffusion model, characterized in that: At least the following steps are included: S1: Batch data preprocessing of the collected image data; S2: Divide the data into non-overlapping data sets; S3: Designing a hybrid loss function for accurate segmentation; S4: designing a hybrid diffusion model, wherein the hybrid diffusion model includes but is not limited to a diffusion model and a segmentation model, and the hybrid diffusion model is HAD-Net; S5: Train HAD-Net and verify the segmentation index. In this process, the joint loss function formulated in S3 is used for training. The gradient back propagation of the neural network is used to update the model weights during the training process. The segmentation effect on the validation set after model training is used to determine whether the model is better. The saved model weights are updated. Finally, the model training is completed and the segmentation effect is evaluated on the test set. S6: Use the trained HAD-Net to perform blurry image segmentation.

2. The fuzzy image segmentation method based on the mixed diffusion model according to claim 1, characterized in that: The image data in S1 includes but is not limited to OCTA images in bmp format, and the data preprocessing includes but is not limited to normalizing the image data for input into the model.

3. The fuzzy image segmentation method based on the mixed diffusion model according to claim 2, characterized in that: The S2 at least The method comprises the following steps: randomly shuffling and dividing the image data into training, verification and test sets in a ratio of 8:1:1, wherein any image data cannot overlap between these sets.

4. The fuzzy image segmentation method based on the mixed diffusion model according to claim 3 is characterized in that: The loss function in S3 is set as follows: Select the mixed loss of DiceLoss and BCELoss as the final loss function; The weight hyperparameters of DiceLoss and BCELoss are λ D and λ M , the sum is 1; MSE Loss: Where N represents the total number of pixels; y i is the true label, the value of the i-th pixel, is the value predicted by the model; DiceLoss: BCELoss: FinalLoss: THE Final =λL MSE +L BCE +L clDice (4) Where C represents the total number of categories; Indicates the c-th class in the binary GroundTruth, and the i-th pixel corresponds to the class value; Represents the predicted probability value of the corresponding category; ε = 1×10-5, which represents the smoothing exponent. ε is used to prevent the denominator from being predicted to be 0, which may lead to some extreme situations. In addition, it can also play a role in smoothing loss and gradient.

5. The fuzzy image segmentation method based on the mixed diffusion model according to claim 4 is characterized in that: The application process of HAD-Net is at least The following steps are involved: set up Represents the vessel segmentation mask with a spatial resolution of where the value of M′ is restricted to the set {0,1}; given a random Gaussian noise image at a specified time step t, the image follows a standard normal distribution The diffusion model predicts the noise level of images with different levels of corruption, and the mean square error between the predicted noise and Gaussian noise is calculated; When the diffusion model training is basically stable, activate the segmentation model and set Represents the original OCTA image with a spatial resolution of H×W and contains C channels, where the pixel values ​​in M ​​are distributed in the range [0, 1]. Use M as the original conditional image at time step T and perform T-step reverse denoising. In order to generate a more stable denoised image and reduce the impact of a single outlier, the original image is denoised five times for a complete T-step denoising, and the average image is calculated as the final prior image; The segmentation model uses the clear vascular images generated from the diffusion model as a priori, combined with the original OCTA image information, and performs the vascular segmentation task through supervised learning. The goal of the segmentation model is to compare the ground truth labels with the segmentation results predicted by the model, thereby guiding the network to perform efficient training; At the same time, the gradient of the forward noise feature learning in the diffusion model will be suppressed and gradually decrease until the end of training, resulting in a stable output.

6. The fuzzy image segmentation method based on the mixed diffusion model according to claim 5, characterized in that: Using HAD-Net to denoise and repair images includes at least the following steps: Using the original image as a generative prior, the segmentation task is redefined as a denoising problem; The forward process of the diffusion model gradually adds random Gaussian noise in t time steps, and gradually transforms the original segmentation mask into noise according to the Markov chain, which leads to a mixed data distribution. The forward process of the diffusion model is defined as: α t is a scheduling factor that controls the variance of the noise. As the time step t increases, the noise gradually increases. The goal of the diffusion model is to gradually transform the image into noise through this forward process, thereby realizing variational inference x0,…,x in the Markov chain of t time steps. T ; α t =1-β t , where β t is a parameter that controls the noise intensity at each step. t Re-parameterization can provide more flexible control over the noise adding process; x t is the graph at the current time step t; (1-α t )I θ is the noise variance; The improved U-Net network D(·) is used to generate the prior distribution layer D(x), so the entire forward process is defined as: Among them, q(x 1:T |||x0,D(x)) represents the conditional probability density of the entire diffusion process, x 1:T represents the state sequence from t=1 to t=T, that is, the image sequence from the initial image x0 to the final time step T; This part shows that the joint probability q(x 1:T |x0,D(x)) is decomposed into a product of conditional probabilities; This means that the noise addition at each step depends not only on the image state x at the previous moment t-1 ,It also relies on the generated prior distribution D(x) to more effectively capture structural features such as blood vessels. It should be noted that the improved U-Net network is to design a parameter Fourier transform control unit in the U-Net encoder; The diffusion model D(·) learns the relationship between the data distribution of the segmentation mask and the noise during the entire training process, effectively capturing the characteristics of blood vessels while modeling the noise characteristics; The denoising process learns the inverse distribution p(x) by minimizing the KL divergence of the forward and backward distributions at all time steps through a neural network θ t-1 |x t ), defined as: Among them, μ θ (x t ,t) and ∑ θ (x t ,t) is the parameter learned by the neural network θ, which is used to predict the transition from xt to x t-1 the recovery process; Predict x using the parameterized diffusion model D(·) t The noise in the image at time instant and restore the image x at that time instant t-1 , the denoising process is defined as: Using the original OCTA image as the data before the denoising T step, and by adjusting the length of the T step to achieve denoising, the complete reverse process is defined as: Among them, p(M) represents the prior probability; After multiple time-step transformations, the diffusion model removes noise from the original image and effectively highlights the vascular structure for restoration.

7. The fuzzy image segmentation method based on the mixed diffusion model according to claim 6, characterized in that: In order to enable the model to better perceive frequency information, a parameter Fourier transform control unit is designed in the U-Net encoder to obtain an improved U-Net network to accurately capture potential vascular features, highlight grayscale changes in OCTA images, and enhance the overall dynamic range; The parameter Fourier transform control unit is used to perform fast Fourier transform and add learnable parameters to adjust the frequency domain characteristics of the image; For images The fast Fourier transform F(.) is performed in the domain to obtain the real part R and the imaginary part I, which respectively contain the high-frequency domain information and low-frequency domain information of the image and the edge detail information; In order to better control the edge texture information and contrast information, learnable parameterized factors α and β are introduced to adjust the transformed amplitude and phase respectively; Subsequently, x′ is obtained by inverse transformation using imaginary numbers, expressed as: R,I=F(x)∈R C×H×W (10) x′=R·α+i·(I·β) (11) Similarly, after the inverse transformation A parameter Fourier transform control unit is embedded in the encoder of the diffusion model of HAD-Net to adjust the frequency domain information of the extracted data features, thereby enhancing the overall contrast of the image; First, apply the convolutional layer Conv in To increase the number of channel features, after conversion, use Conv out Layers are used to smooth the output features. The entire encoder process is defined as: F out =Con out (PFT(Conv in (F in ))) (12) Among them, Conv is represented as a 3×3 convolutional layer with a normalization layer and ReLU activation, and F in represents the input features, and F out Represents the output features.

8. The fuzzy image segmentation method based on the mixed diffusion model according to claim 7, characterized in that: In view of the fact that in some areas, thin blood vessels are still mixed with background noise, no filtering method is directly applied to eliminate these noise effects, but the spatial representation of small blood vessels is preserved; In order to separate small blood vessels from the mixed background, a WTEncoder guided by wavelet convolution is designed. In addition, in order to preserve the complex vascular details, the use of downsampling layers is minimized to maintain the spatial scale; Given an input image Four filters are used to extract the high and low frequency channels, the filters are defined as: f LL represents a low-pass filter, used to extract low-frequency information; While the other three filters form a set of high-pass filters to extract high-frequency information; The four filter groups are subjected to a convolution operation with a stride of 2 to generate four different filter channel responses, expressed as The deep convolution DC is then applied to the four generated frequency domain channels to extract features and strengthen the interaction between different frequency domains, resulting in the output Y out , this process is defined as: Y out =DC(Conv(X,{f LL ,f LH ,f HL ,f HH })) (13) Since the four filters form an orthonormal basis, we apply a transposed convolution with stride 2 to Y out Execute the operation to get the restored image At the same time, the original feature vector is deeply convolved to extract its information representation, and the frequency domain features obtained by inverse transformation are used to enhance it. The final output feature is F out , expressed as: F out =DC(X)+TC(Y out ,{f LL ,f LH ,f HL ,f HH }) (14) Among them, TC represents transposed convolution, and the final output is the sum of the original feature convolution and the wavelet convolution; Afterwards, BatchNorm and GELU activation are performed. In order to fully capture the edge and frequency information, a two-layer wavelet convolutional encoder is designed; First, a 7×7 sliding convolution kernel is used to extract information from the local area, and a 1×1 convolution is used to provide a residual connection for the original information; Wavelet transform allows iterative transformation to generate higher dimensional and lower scale frequency domain information; However, given the existing clear prior vascular information, a two-step low-order wavelet transform convolution was designed to further fit the features instead of calculating redundant frequency domain information; A single layer of wavelet convolution is used for initial enhancement of low-frequency and texture features; After the second layer of wavelet convolution, the correlation between the frequency domains is recalculated, gradually guiding the model to focus on clear main vessels and capillaries; Finally, a trainable augmentation factor is used to globally guide the feature extraction path and linearly regulate the correlation between the residual path and the main path.

Citation Information

Cited By

  • Medical image segmentation system based on multi-architecture fusion and diffusion model

    CN121053147A

  • Channel fingerprint generation and recovery method based on conditional latent diffusion model

    CN121190311A