Unsupervised OCT retina image denoising method based on diffusion model
Through an unsupervised method based on the diffusion model, the high-frequency multi-scale features of OCT images are extracted using wavelet transform and fused into the diffusion model, which solves the problem of the existing methods requiring truth tags and ignoring high-frequency information, and achieves efficient OCT image denoising and edge detail restoration.
Patent Information
- Application Number
- CN202510077164.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-30
AI Technical Summary
The existing OCT retinal image denoising method requires a large number of hard-to-get truth-value tags, and ignores the key information of the OCT image in the high frequency domain, resulting in blurred image edge details after denoising.
Using an unsupervised method based on diffusion model, the OCT images are decomposed into low-frequency and high-frequency domains through wavelet transformation, high-frequency multi-scale features are extracted, and fused into the diffusion model to guide noise removal.
The OCT images can be denoised without the truth tag, effectively utilize high-frequency information, improve the edge detail restoration effect of the image, enhance the usability of the image and the accuracy of medical diagnosis.
Smart Images

Figure CN120070233A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical image processing, and relates to a diffusion model, mainly for noise removal of OCT retinal images. Background Art
[0002] Optical Coherence Tomography (OCT) is a non-invasive imaging technique that can be used to obtain three-dimensional images of biological tissues in real time. It does not require direct contact with organs during operation, and patients do not experience excessive discomfort during the examination, which makes it play a crucial role in the field of organ examination. To prevent motion artifacts, OCT usually captures OCT images at a relatively low sampling rate, resulting in relatively low-resolution and noisy OCT images, which is not conducive to subsequent medical diagnosis using OCT images. Therefore, studying denoising methods for OCT retinal images can contribute to subsequent medical work.
[0003] In recent years, deep learning-based methods, especially Convolutional Neural Networks (CNNs), have been widely used in medical image processing. For denoising retinal OCT images, conventional CNN models mainly consist of convolutional layers and pooling layers. Since convolutional and pooling operations can only extract local features, conventional CNN models lack the ability to perceive and learn global structural features in images. Conditional Generative Adversarial Networks (cGANs) select U-shaped networks (U-Nets) as generators to generate low-noise images, and Visual Geometry Group networks (VGGs) as discriminators to distinguish between ground-truth images and denoised images. Compared with traditional algorithms, deep learning-based denoising methods have achieved better results in preserving image edge details. However, the training of the model requires a large number of noise-free OCT retinal clean images as labels to form noise-noise-free image pairs for the neural network to learn the mapping relationship between the two. However, ground-truth images need to be obtained through the method of multi-frame averaging, which is time-consuming and complex, and is not conducive to the popularization of denoising methods. To solve this problem, in recent years, researchers have proposed unsupervised deep learning denoising strategies, such as using paired noisy images to train deep learning networks to achieve denoising performance similar to supervised learning. Existing methods usually ignore the differences in the low-frequency and high-frequency domains of OCT images and do not process OCT images in the frequency domain, resulting in relatively blurred edge textures in the denoised OCT images, affecting the usability of the images and subsequent precision medical analysis. Summary of the Invention
[0004] To overcome the problem that existing methods require a large number of true value labels which are difficult to obtain, and the traditional denoising means often ignore the key information contained in the high-frequency domain of OCT images during the processing, resulting in the blurred edge details of the finally restored images, the present invention proposes an unsupervised OCT retinal image denoising method based on a diffusion model. The diffusion model can simulate the noise distribution without labels to restore OCT images. The wavelet transform can decompose the OCT retinal image into a low-frequency domain and a high-frequency domain, splice the high-frequency domain components on the channels, then extract high-frequency multi-scale features, and use these features to guide the noise in the diffusion model, aiming to help the network learn the high-frequency detailed structure.
[0005] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0006] An unsupervised OCT retinal image denoising method based on a diffusion model, comprising the following steps:
[0007] 1) Preprocess the original OCT retinal image;
[0008] 2) Perform wavelet transform on the preprocessed OCT retinal image to decompose it into a low-frequency domain and a high-frequency domain;
[0009] 3) For the high-frequency domain components obtained by decomposition, first splice them on the channels, then obtain multi-scale features through a backbone network, and then extract high-frequency multi-scale features containing edge textures and high-frequency noise through a multi-layer convolutional network;
[0010] 4) Construct a diffusion model architecture, use U-Net as the backbone network of the diffusion model, and the original OCT image is gradually added with noise during the forward process and finally degenerates into a pure noise image;
[0011] 5) During the process of U-Net predicting noise, fuse the processed high-frequency multi-scale features into the diffusion model to guide the removal of noise, and finally generate a denoised OCT image.
[0012] Further, in the step 1), the preprocessing process is: adjust the OCT retinal image to a fixed size to ensure that all images input into the subsequent model have a unified size, then normalize the pixel values to the interval [0, 1], and map them to the interval [1, 3] through a linear transformation.
[0013] Still further, in the step 2), the wavelet transform process is: use the Haar wavelet to decompose the image into a low-frequency domain approximation image and high-frequency domain detail subgraphs of different scales. The low-frequency approximation image represents the contour and low-frequency information of the image, and the high-frequency detail subgraphs correspond to the high-frequency detail information in the horizontal, vertical, and diagonal directions.
[0014] Furthermore, in the step 3), the high-frequency information is processed through feature extraction to obtain high-frequency multi-scale features, which includes the following sub-steps:
[0015] 31) Pass through a pyramid feature extraction layer with a backbone network to extract a list of pyramid-shaped features;
[0016] 32) Use a convolutional layer with a convolution kernel size of 3×3 to perform feature extraction on each layer of the feature map. The convolutional layer includes a combination of a convolution function, a ReLu activation function, and a normalization function;
[0017] 33) Input the multi-scale features after the above processing into a multi-layer convolutional network to obtain abstract multi-scale high-frequency features.
[0018] Furthermore, in the step 4), the process of denoising the original OCT image using a diffusion model includes the following sub-steps:
[0019] 41) Forward diffusion process, adding noise: Gradually add Gaussian noise to the training image until the image becomes pure noise. This process is regarded as a Markov process;
[0020] 42) Model training process, learning noise: Adopt a U-Net architecture as the backbone network, which has an encoder-decoder structure and can effectively perform feature extraction and reconstruction. The input of the network is the noisy data and the encoding of the time step. Convert the time step into a vector and embed it into the network input through positional encoding;
[0021] 43) Through the backpropagation algorithm, use the noisy data and the corresponding original data to train the network and adjust the parameters of the network so that the network can learn how to predict data closer to the real denoised data given the time step and the noisy data;
[0022] 44) Reverse generation process, removing noise: The reverse generation process starts from pure noise and uses the trained model to predict the denoising steps to gradually restore the noise-free image.
[0023] In the step 5), the fusion process of the high-frequency multi-scale features includes the following sub-steps:
[0024] 51) Input the latest feature map into the corresponding convolution operation of the current module, and at the same time input the time embedding information for processing to obtain intermediate features;
[0025] 52) In the encoder stage of the U-Net, at each non-final resolution level, feature fusion and downsampling operations are performed on the features of the corresponding level in the high-frequency multi-scale features and the current feature map. The high-frequency detail features are adjusted to the same size as the current feature through the bilinear interpolation function, and after concatenation in dimension 1, they are added to the current feature, so as to integrate the high-frequency detail multi-scale features into the feature representation of the current level, helping the network capture the detail information of the image and enhancing the perception ability of different-scale structures;
[0026] 53) In the decoder stage of the U-Net, feature fusion and upsampling operations are also performed at the non-final resolution level. The adjusted high-frequency detail features are obtained through bilinear interpolation, and after concatenation in dimension, they are added to the current feature, enabling the U-Net to utilize the high-frequency detail information of the image to learn the noise distribution and image structure.
[0027] Preferably, the optimal parameters of the network model are obtained during training. The network training method is as follows: backpropagation optimization is performed on the loss function value between the predicted noise and the true noise, and the mean square error loss is used to calculate the error between the predicted denoised data and the true data.
[0028] An unsupervised OCT retinal image denoising method based on a diffusion model proposed by the present invention is mainly characterized by the extraction and utilization of OCT high-frequency detail information. Due to the difficulty of obtaining ground truth labels, supervised OCT retinal denoising images are difficult to promote. Conventional OCT denoising methods do not analyze the different characteristics of OCT in the low-frequency domain and the high-frequency domain, and cannot well restore the edge details of OCT images after denoising. Therefore, the present invention proposes to use a diffusion model to denoise OCT retinal images. The diffusion model does not require labels. It adds noise through forward propagation, uses the network to simulate the predicted noise, and then removes the noise through backpropagation to achieve the denoising of OCT retinal images. The high-frequency components after wavelet decomposition are processed by a multi-layer convolutional network, and the high-frequency detail information is converted into a multi-scale feature representation. The multi-layer convolutional network for processing high-frequency components starts from a higher convolutional layer number and reduces the convolutional layer number as the network deepens. This structural design helps to learn local detail information of the image in the shallow layer of the network. Furthermore, the present invention proposes a feature fusion module to guide the network predicted noise using multi-scale high-frequency detail features. In the encoder part of the U-Net, the multi-scale features are concatenated with the output features of each layer of the encoder in the channel dimension, so that when the U-Net performs feature extraction, it can simultaneously utilize the original image information and the multi-scale features after high-frequency detail processing. In the decoder part of the U-Net, the multi-scale high-frequency features are inversely transformed and then concatenated with the output features of each output layer, which can be used as prior information to guide the training process of the unsupervised denoising diffusion model. The high-frequency subband features are fused with the noise images at different time steps in the diffusion model. Using the noisy OCT image as the input, by minimizing the loss function, the U-Net network is trained to learn the mapping relationship from the noisy image to the noise, so that the diffusion model can better utilize the high-frequency detail information of the image to learn the noise distribution and image structure.
[0029] The beneficial effects of the present invention are: getting rid of the dependence on a large number of difficult-to-obtain ground truth labels, and effectively avoiding the drawback of blurred edge detail restoration due to the lack of high-frequency information, aiming to improve the quality of OCT images for subsequent medical diagnosis. Description of the Drawings
[0030] Figure 1 It is a flowchart of an unsupervised OCT retinal image denoising method based on a diffusion model.
[0031] Figure 2 It is a flowchart of high-frequency detail processing.
[0032] Figure 3 It is a multi-scale high-frequency feature fusion diagram. Detailed Embodiment
[0033] The present invention will be further described below in conjunction with the flowchart.
[0034] Reference Figures 1 to 3 , an unsupervised OCT retinal image denoising method based on a diffusion model, comprising the following steps:
[0035] 1) Preprocess the original OCT retinal image. The processing process is as follows: adjust the OCT retinal image to a fixed size, such as 512×512, to ensure that all images input into the subsequent model have a unified size, and then normalize the pixel values to the interval [0,1] and map them to the interval [1,3] through a linear transformation;
[0036] 2) Decompose the preprocessed image into a low-frequency domain and three high-frequency domains in the horizontal, vertical, and diagonal directions using Haar wavelets, and splice these three high-frequency domains on the channel;
[0037] 3) Process the spliced high-frequency information, as Figure 2 shown, the processing process includes the following sub-steps:
[0038] 31) Extract a pyramidal feature list through a feature extraction layer with a backbone network, and the sizes of its feature maps are 16×16, 32×32, 64×64, and 128×128 from small to large;
[0039] 32) Use a convolutional layer with a convolutional kernel size of 3×3 to extract features for each layer of the feature map. The convolutional layer includes a combination of a convolution function, a ReLu activation function, and a normalization function;
[0040] 33) Input the multi-scale features after the above processing into a multi-layer convolutional network to obtain abstract multi-scale high-frequency features. The multi-scale high-frequency features are 16×16, 32×32, 64×64, and 128×128 in ascending order of the feature map size.
[0041] Specifically, the first column convolutional module of the multi-layer convolutional network consists of four convolutional layers, the second column convolutional module consists of three convolutional layers, the third column convolutional module consists of two convolutional layers, and the fourth column convolutional module consists of one convolutional layer. In each convolutional module, the input of the latter convolutional layer is formed by splicing the outputs of the previous several convolutional layers on the channel. Among them, the convolutional kernel size of the convolutional layer is 3×3, and it includes a ReLu activation function and a normalization function.
[0042] 4) The process of using the diffusion model to denoise the original OCT image includes the following sub-steps:
[0043] 41) Forward diffusion process, adding noise: Gradually add Gaussian noise to the training image until the image becomes pure noise. This process is regarded as a Markov process;
[0044] 42) Model training process, learning noise: The U-Net architecture is used as the backbone network, which has an encoder-decoder structure and can effectively perform feature extraction and reconstruction. The input to the network is noisy data and the encoding of the time step. The time step is converted into a vector and embedded into the network input through positional encoding;
[0045] 43) Through the backpropagation algorithm, the network is trained using the noisy data and the corresponding original data, and the parameters of the network are adjusted so that the network can learn how to predict data closer to the true denoised data given the time step and the noisy data;
[0046] 44) Inverse generation process, removing noise: The inverse generation process starts from pure noise and uses the trained model to predict the denoising steps to gradually recover the noise-free image.
[0047] In step 5) above, the fusion process of high-frequency multi-scale features, as Figure 3 shown, includes the following sub-steps:
[0048] 51) The latest feature map is passed into the corresponding convolutional operation of the current module, and at the same time, the time embedding information is passed in for processing to obtain intermediate features;
[0049] 52) In the U-Net encoder stage, feature fusion and downsampling operations are performed at each non-final resolution level. For the feature map size H×W×C (H is the height, W is the width, and C is the number of channels) at each resolution level in the encoder, the size of the high-frequency detail features is interpolated to H×C through the bilinear interpolation function, and then the two are added together, so as to integrate the high-frequency detail multi-scale features into the feature representation of the current level;
[0050] 53) In the U-Net decoder stage, feature fusion and upsampling operations are also performed at the non-final resolution level. For the feature map size H×W×C (H is the height, W is the width, and C is the number of channels) at each resolution level in the decoder, the size of the high-frequency detail features is interpolated to H×C through the bilinear interpolation function to obtain the adjusted high-frequency detail features through bilinear interpolation, and then the two are added together;
[0051] 54) After the noise prediction network is trained, it predicts the noise of the current image, and then gradually outputs the noise-free image through the inverse process.
[0052] In this embodiment, the optimal network parameters are obtained during model training, and the network training method is: backpropagation optimization is performed on the loss function value between the predicted noise and the true noise. Specifically, the mean square error loss is used to calculate the error between the predicted denoised data and the true data.
[0053] The content described in the embodiments of this specification is only an enumeration of the implementation forms of the inventive concept and is for illustrative purposes only. The protection scope of the present invention should not be regarded as limited to the specific forms stated in these embodiments, and the protection scope of the present invention also extends to equivalent technical means that can be conceived by those of ordinary skill in the art based on the inventive concept of the present invention.
Claims
1. An unsupervised OCT retinal image denoising method based on a diffusion model, characterized in that: The method comprises the following steps: 1) Preprocessing the original OCT retinal images; 2) Perform wavelet transform on the preprocessed OCT retinal image to decompose the low-frequency domain and high-frequency domain; 3) For the high-frequency domain components obtained by decomposition, they are first spliced on the channel, then multi-scale features are obtained through a backbone network, and then high-frequency multi-scale features including edge texture and high-frequency noise are extracted through a multi-layer convolutional network; 4) Construct a diffusion model architecture and use U-Net as the backbone network of the diffusion model. The original OCT image is gradually added with noise in the forward process and eventually degraded into a pure noise image; 5) In the process of U-Net predicting noise, the high-frequency multi-scale features extracted by the high-frequency detail processing module are fused into the diffusion model to guide the removal of noise and finally generate the denoised OCT image.
2. The unsupervised OCT retinal image denoising method based on a diffusion model as claimed in claim 1, characterized in that: In step 1), the preprocessing process is: adjusting the OCT retinal image to a fixed size to ensure that all images input into the subsequent model have a uniform size, and then normalizing the pixel values to the [0, 1] interval, and mapping them to the [1, 3] interval through linear transformation.
3. The unsupervised OCT retinal image denoising method based on a diffusion model as claimed in claim 1 or 2, characterized in that: In the step 2), the wavelet transform process is: using Haar wavelet to decompose the image into a low-frequency domain approximate image and high-frequency domain detail sub-images of different scales, the low-frequency approximate image represents the contour and low-frequency information of the image, and the high-frequency detail sub-image corresponds to the high-frequency detail information in the horizontal, vertical and diagonal directions.
4. The unsupervised OCT retinal image denoising method based on a diffusion model as claimed in claim 1 or 2, characterized in that: In step 3), extracting high-frequency information to obtain high-frequency multi-scale features includes the following sub-steps: 31) extracting a pyramidal feature list through a pyramid feature extraction layer with a backbone network; 32) A convolution layer with a convolution kernel size of 3×3 is used to extract features from each layer of feature maps. The convolution layer includes a combination of a convolution function, a ReLu activation function, and a normalization function; 33) The multi-scale features after the above processing are input into a multi-layer convolutional network to obtain abstract multi-scale high-frequency features.
5. The unsupervised OCT retinal image denoising method based on a diffusion model as claimed in claim 1 or 2, characterized in that: In step 4), the structural attention decoder module includes the following sub-steps at each level of processing: 41) Forward diffusion process, adding noise: Gaussian noise is gradually added to the training image until the image becomes pure noise. This process is regarded as a Markov process; 42) Model training process, learning noise: The U-Net architecture is used as the backbone network, which has an encoder-decoder structure and can effectively perform feature extraction and reconstruction. The input of the network is the noisy data and the encoding of the time step. The time step is converted into a vector and embedded into the network input through position encoding; 43) Through the back propagation algorithm, the network is trained using noisy data and the corresponding original data, and the parameters of the network are adjusted so that the network can learn how to predict data that is closer to the actual denoised data given a time step and noisy data; 44) Reverse generation process to remove noise: The reverse generation process starts from pure noise and uses the trained model to predict the denoising steps to gradually restore the noise-free image.
6. The unsupervised OCT retinal image denoising method based on a diffusion model as claimed in claim 1 or 2, characterized in that: In step 5), the processing process of the fusion module includes the following sub-steps: 51) The latest feature map is passed into the convolution operation corresponding to the current module, and the time embedding information is passed in for processing to obtain the intermediate features; 52) In the U-Net encoder stage, at each non-final resolution level, the features of the corresponding level in the high-frequency multi-scale features are fused and downsampled with the current feature map, and the high-frequency detail features are adjusted to the same size as the current features through the bilinear interpolation function. The channels of the two are aligned through the channel splicing operation, so as to integrate the high-frequency detail multi-scale features into the feature representation of the current level, helping the network capture the detailed information of the image; 53) In the U-Net decoder stage, feature fusion and upsampling operations are also performed at the non-final resolution level. The high-frequency multi-scale features are inversely transformed, and then the corresponding feature maps are bilinearly interpolated to obtain the same feature map size as the current level of the decoder. The two channels are aligned through channel concatenation and then added to the current feature, so that U-Net can use the high-frequency detail information of the image to learn the noise distribution and image structure.
7. The unsupervised OCT retinal image denoising method based on a diffusion model as claimed in claim 6, characterized in that: The optimal parameters of the network model are obtained during training. The network training method is: back-propagation optimization of the loss function value between the predicted noise and the real noise is performed, and the error between the predicted denoised data and the real data is calculated using the mean square error loss.
Citation Information
Cited By
Clinical AS-OCT image restoration method and system based on conditional diffusion model
CN121685319A
Medical image segmentation method and system based on diffusion difference learning model
CN121811411A