An OCT Image Enhancement Method and Application Based on a Registration Network

By using a registration network-based method in OCT image enhancement, combining the multi-head self-attention transformation module and the multi-scale deformation field fusion module, the problem of image features loss and low accuracy during multi-denoising registration is solved, and efficient and accurate image denoising and enhancement effects are achieved.

CN116188309BActive Publication Date: 2025-06-10SUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310158800.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-24
Publication Date
2025-06-10
Estimated Expiration
2043-02-24

AI Technical Summary

Technical Problem

In the prior art, there is a problem of image features lost and low accuracy when multiple denoising registration is obtained.

Method used

The OCT image enhancement method based on the registration network is adopted, and the image registration network is constructed through the multi-head self-attention transformation module and the multi-scale deformation field fusion module to achieve high-precision registration of fixed images and mobile images, and reduce artifacts and enhance images.

Benefits of technology

The calculation efficiency of image denoising is improved, artifacts are reduced, image quality is enhanced, and image registration with high registration accuracy is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116188309B_ABST
    Figure CN116188309B_ABST
Patent Text Reader

Abstract

The present invention relates to an OCT image enhancement method based on a registration network, which includes selecting any one OCT image in the sample as the fixed image and the remaining OCT images as the moving images; inputting the fixed image into the fixed image encoding module, passing through five fixed encoding blocks, and the fixed encoding blocks input the encoded output to the corresponding decoding blocks; inputting the moving image into the moving image encoding module, passing through five moving encoding blocks, and the moving encoding blocks input the encoded output to the corresponding decoding blocks; the outputs of the third fixed encoding block and the third moving encoding block are output to the first decoding block of the decoding module through the multi-head self-attention transformation module; each decoding block in the decoding module is based on the output of the previous decoding block and the output of the corresponding encoding block, and after restoring the dimensions layer by layer from low resolution to high resolution, it outputs to the multi-scale deformation field fusion module to obtain the registered image; repeating the foregoing steps to obtain multiple registered images, and performing superposition averaging with the fixed image to obtain the denoised image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of OCT imaging, and in particular to an OCT image enhancement method and application based on a registration network. Background Art

[0002] Optical Coherence Tomography (OCT) is a high-resolution, non-contact, non-invasive technology, which is widely used in the structural imaging of living biological tissues. Statistical results show that more than half of the reports on OCT are in the field of ophthalmology for diagnosing various retinal diseases, etc. However, since OCT is based on low-coherence interference, the measured signal will inevitably be affected by speckle noise, resulting in poor OCT image quality.

[0003] In recent years, many methods have been proposed to obtain high-quality OCT images by suppressing speckle noise. The methods for suppressing speckle noise are mainly divided into two categories. One category of methods is to reduce speckle noise by post-processing the two-dimensional image of a single OCT, such as BM3D using traditional algorithms. This method matches with adjacent image blocks, integrates several similar blocks into a three-dimensional matrix, filters the three-dimensional matrix in the three-dimensional space, and then inverse-transforms the filtered result to a two-dimensional image to form a denoised image. The single-image denoising method based on deep learning, such as Edge-cGAN, can greatly improve the computational efficiency. This method proposes an end-to-end conditional generative adversarial network, adds an edge loss function to the objective function, enhances the sensitivity of the model to edge details, enhances the contrast of the image, and suppresses the speckle noise in the image. Another category of methods is to denoise based on multi-frame superposition averaging. Multi-frame superposition averaging is one of the effective methods for reducing speckle noise. However, when the eyeball moves during shooting and there are position differences between too many frames, it is very likely to cause artifacts after averaging. Therefore, before averaging, the misalignment between images must be calibrated by a registration algorithm. Therefore, this category of methods focuses on the technical research of image registration. For example, the regular dynamic programming technique is used to correct the axial offset in the axial direction, and the cross-correlation degree at a series of positions is calculated in the horizontal direction to obtain the optimal registration position; another example is when using the TransMorph model based on deep learning for image registration, the convolutional structure of encoding and decoding is used to capture the relationship between image pixels, and the registration deformation field is formed by the decoded features of the last layer for registration.

[0004] Single-frame denoising reduces speckle noise by post-processing a two-dimensional slice. Traditional single-frame denoising methods such as the BM3D algorithm use filtering for denoising, but there are some drawbacks. For example, this method has a high complexity in implementation and low computational efficiency. In single-frame denoising based on deep learning, such as Edge-cGAN, a method Edge-cGAN for end-to-end removal of image noise based on an improved conditional generative adversarial network cGAN is proposed and has achieved good performance. However, there are still some deficiencies. For example, it is limited by single-slice scanning, resulting in information loss in some areas covered by noise; the complexity of the improved cGAN model increases, and there are too many hyperparameters; and single-frame denoising based on deep learning requires a gold standard for image denoising for training, and the source of this gold standard is also obtained by averaging multiple images. Deep learning approaches this gold standard through different trainings but cannot exceed it.

[0005] Moreover, if multiple-frame denoising is directly averaged, due to the positional differences between frames, serious artifacts will occur. Therefore, registration is required first. In terms of registration, traditional registration methods such as using regular dynamic programming techniques to correct axial offsets axially and calculating the cross-correlation degrees at a series of positions to obtain the optimal registration position horizontally require relatively complex cross-correlation calculations and a large computational cost. Although the deep learning-based registration method in the TransMorph model has greatly improved the computational efficiency and registration accuracy, it will miss some intermediate encoding and decoding layers, resulting in loss of feature information due to errors such as interpolation. Summary of the Invention

[0006] Therefore, the technical problem to be solved by the present invention is to overcome the problems of image feature loss and low accuracy during multi-frame denoising registration in the prior art.

[0007] To solve the above technical problems, the present invention provides an OCT image enhancement method based on a registration network, including:

[0008] Select any OCT image in a certain sample of the training set as the fixed image, and use the remaining OCT images in this sample as the moving images;

[0009] Input the fixed image into the fixed image encoding module of the registration network. After five fixed encoding blocks, the encoded output of each fixed encoding block is input into the corresponding decoding block;

[0010] Input the moving image into the moving image encoding module of the registration network. After five moving encoding blocks, the encoded output of each moving encoding block is input into the corresponding decoding block;

[0011] The outputs of the third fixed encoding block of the fixed image encoding module and the third mobile encoding block of the mobile image encoding module both pass through the multi-head self-attention transformation module; the multi-head self-attention transformation module performs convolution operations and flattening operations on the input fixed image feature map and mobile image feature map respectively, and splices and fuses the fixed image feature map and the mobile image feature map through a concatenation operation, inputs the fused feature map and the image position information into the transformer encoding unit, and then flattens and outputs to the first decoding block of the decoding module through layer normalization and convolution block operations;

[0012] Each decoding block in the decoding module is based on the output of the previous decoding block and the outputs of the corresponding fixed encoding block and mobile encoding block. After restoring the dimension of the input feature image layer by layer from low resolution to high resolution, it outputs to the multi-scale deformation field fusion module to obtain the registered image of the fixed image;

[0013] Repeat to input the fixed image and the unregistered mobile image into the registration network to obtain the remaining registered images of the fixed image;

[0014] Overlay and average the fixed image and multiple registered images to obtain the denoised image of the fixed image.

[0015] In an embodiment of the present invention, the transformer encoding unit includes a multi-head self-attention block and a multi-layer perceptron; the multi-head self-attention block includes a plurality of self-attention blocks connected together in a channel manner in the transformer encoding unit.

[0016] In an embodiment of the present invention, both the fixed image encoding module and the mobile image encoding module include a first encoding block, a second encoding block, a third encoding block, a fourth encoding block, and a fifth encoding block connected in series in the forward propagation direction;

[0017] The first encoding block includes two serially connected convolution units;

[0018] The convolution unit includes a convolution block, a batch normalization layer, and a ReLU layer connected in series in the forward propagation direction;

[0019] Both the second encoding block and the third encoding block include a max pooling unit and two convolution units connected in series in the forward propagation direction;

[0020] Both the fourth encoding block and the fifth encoding block include a max pooling unit.

[0021] In one embodiment of the present invention, the decoding module includes a first decoding block, a second decoding block, a third decoding block, a fourth decoding block, and a fifth decoding block connected in series in sequence; each decoding block includes a multi-scale deformation field fusion module, a spatial transformation module, and a convolution unit connected in series in sequence; the multi-scale deformation field fusion module is used to learn the details of global affine registration; the convolution unit includes a convolution block, a batch normalization layer, and a ReLU layer connected in series in sequence; the spatial transformation module is used to learn the input features, obtain transformation parameters, apply the transformation parameters to the moving image, and perform mapping to obtain a registered image.

[0022] In one embodiment of the present invention, the multi-scale deformation field fusion module in the first decoding block performs convolution on the input features to obtain a first deformation field, and performs upsampling on the first deformation field to obtain a first output deformation field;

[0023] The multi-scale deformation field fusion module in the nth decoding block performs convolution on the input features to obtain an nth deformation field, performs upsampling on the nth deformation field and the (n - 1)th deformation field respectively, and superimposes the upsampled results to obtain an nth output deformation field; where 2 ≤ n ≤ 5.

[0024] In one embodiment of the present invention, before setting any one OCT image in a certain sample of the training set as the fixed image, it includes collecting OCT images of multiple retinal samples, performing resampling, constructing an image data set, and dividing the image data set into a training set and a test set according to a preset ratio.

[0025] In one embodiment of the present invention, obtaining multiple registered images of the fixed image includes:

[0026] According to the deformation field parameters output by the multi-scale deformation field fusion module, use the spatial transformation module to distort the moving image so that it is registered with the fixed image, and obtain the registered image of the fixed image.

[0027] In one embodiment of the present invention, after obtaining the registered image of the fixed image, it further includes:

[0028] Optimize the registration network according to the loss function of the fixed image and the registered image;

[0029] Select any one OCT image in a certain sample of the test set as the fixed image, use the remaining OCT images in this sample as the moving images, and repeatedly input the fixed image and different moving images into the optimized registration network to obtain multiple registered images of the fixed image in this sample of the test set;

[0030] Superimpose and average the fixed image and the multiple registered images of this sample to obtain a denoised image.

[0031] In one embodiment of the present invention, the loss function of the fixed image and the registered image is as follows:

[0032] L joint = L sim + αL smooth ,

[0033]

[0034]

[0035] where L sim is used to constrain the smoothness in the deformation field, L smooth is used to constrain the local spatial variation in the deformation field, and α is a regularization parameter, set to 10; represents the k-th layer label of the segmentation map L f of the fixed image layer, represents the k-th layer label of the segmentation map L w of the registered image layer, k represents the number of layers in the layer segmentation label; φ represents the deformation field, p represents any point in the deformation field, μ(p) represents the value of the coordinates (x, y) of point p in the deformation field, the gradient of p in the x direction of the deformation field the gradient of p in the y direction of the deformation field

[0036] The embodiment of the present invention also provides an application of the OCT image enhancement method based on the registration network in the field of retinal OCT image enhancement.

[0037] The above technical solution of the present invention has the following advantages compared with the prior art:

[0038] The OCT image enhancement method based on the registration network provided by the present invention uses the registered images for multi-frame average denoising, effectively solving the problem of image information masking existing in single-frame denoising; compared with the traditional single-frame denoising method, the calculation efficiency is greatly improved; compared with the single-frame denoising method based on deep learning, it is not limited by the quality of the denoising gold standard in the training data. And the present invention constructs an image registration network by combining the multi-head self-attention conversion module and the multi-scale deformation field fusion module; using the multi-head self-attention conversion module, captures more relationships between the fixed image and the moving image, and realizes high-precision registration of the fixed image and the moving image; using the multi-scale deformation field fusion module to use the multi-resolution strategy to learn the details of the global affine registration, the deformation field based on the features of the previous layer and the feature output have a high-resolution deformation field, enhancing the fine-grained positioning, enhancing the volume overlap between anatomical structures, and reducing the number of folded voxels between the registration deformation fields, avoiding the loss of feature information, so that when performing average denoising, artifacts are reduced, the image is enhanced, high-precision image registration is achieved, and a good image denoising effect is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to make the content of the present invention easier to be clearly understood, the following further describes the present invention in detail according to specific embodiments of the present invention in combination with the accompanying drawings, where

[0040] Figure 1 is the framework diagram of the registration network provided by the present invention;

[0041] Figure 2 is the structural diagram of the multi-head self-attention transformation module of the registration network provided by the present invention;

[0042] Figure 3 is the schematic diagram of the decoding block structure of the registration network provided by the present invention;

[0043] Figure 4 is the flowchart of the retinal OCT image enhancement method based on the registration network provided by the present invention;

[0044] Figure 5 (a) of represents the original image in the first test set, Figure 5 (b), (c), (d), (e), (f), (g), (h) of respectively represent the denoising effect images of denoising using BM3D, KSVD, DP+HC, Edge-cGAN, Mini-cGAN, DHNet, and the method proposed in the embodiment of the present invention;

[0045] Figure 6 (a) of represents the original image in the second test set, Figure 6 (b), (c), (d), (e), (f), (g), (h) of respectively represent the denoising effect images of denoising using BM3D, KSVD, DP+HC, Edge-cGAN, Mini-cGAN, DHNet, and the method proposed in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The following further describes the present invention in combination with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments cited do not limit the present invention.

[0047] Referring to Figure 1 shown, the framework diagram of the registration network provided by the present invention, the OCT image enhancement method based on the registration network includes:

[0048] S1: Select any OCT image in a certain sample of the training set as the fixed image, and use the remaining OCT images in the sample as the moving images;

[0049] S2: Input the fixed image into the fixed image encoding module of the registration network. After passing through five fixed encoding blocks, the encoded output of each fixed encoding block is input into the corresponding decoding block;

[0050] The fixed image encoding module includes a first fixed encoding block, a second fixed encoding block, a third fixed encoding block, a fourth fixed encoding block, and a fifth fixed encoding block connected in series along the forward propagation direction;

[0051] The first fixed encoding block includes two convolutional units connected in series. The convolutional unit includes a convolutional block, a batch normalization layer, and a ReLU layer connected in series along the forward propagation direction. The second fixed encoding block and the third fixed encoding block both include a max pooling unit and two convolutional units connected in series along the forward propagation direction. The fourth fixed encoding block and the fifth fixed encoding block both include a max pooling unit.

[0052] S3: Input the moving image into the moving image encoding module of the registration network. After passing through five moving encoding blocks, the encoded output of each moving encoding block is input into the corresponding decoding block;

[0053] The moving image encoding module includes a first moving encoding block, a second moving encoding block, a third moving encoding block, a fourth moving encoding block, and a fifth moving encoding block connected in series along the forward propagation direction;

[0054] The first moving encoding block includes two convolutional units connected in series. The second moving encoding block and the third moving encoding block both include a max pooling unit and two convolutional units connected in series along the forward propagation direction. The fourth moving encoding block and the fifth moving encoding block both include a max pooling unit.

[0055] S4: The outputs of the third fixed encoding block of the fixed image encoding module and the third moving encoding block of the moving image encoding module both pass through the multi-head self-attention transformation module. The multi-head self-attention transformation module performs convolutional operations and flattening operations on the input fixed image feature map and moving image feature map respectively, and splices and fuses the fixed image feature map and the moving image feature map through a concatenation operation. The fused feature map and the image position information are input into the transformer encoding unit, and then flattened and output to the first decoding block of the decoding module through layer normalization and convolutional block operations;

[0056] Refer to Figure 2As shown in the figure, the multi-head self-attention transformation module provided by the embodiment of the present invention is such that the feature images of the fixed image and the moving image pass through the convolution operation with a convolution kernel of 8 and the flattening operation at the same time, and then the two feature images are fused through the concatenation operation. The fused feature image and the position information of the image are sent to the transformer encoding module together. The transformer encoding module includes a multi-head self-attention block and a multi-layer perceptron. Finally, the output of the multi-head self-attention transformation module is obtained through the normalization operation and the convolution operation.

[0057] The fact that the multi-head self-attention transformation module can successfully model long-range dependencies is mainly attributed to the multi-head self-attention block. The multi-head self-attention block mainly includes a plurality of self-attention blocks SA connected together in a channel manner in the transformer encoding block. In the embodiment of the present invention, the number of SA blocks is set to 12.

[0058] S5: Each decoding block in the decoding module is based on the output of the previous decoding block and the outputs of the corresponding fixed encoding block and moving encoding block. After restoring the dimension of the input feature image layer by layer from low resolution to high resolution, it outputs to the multi-scale deformation field fusion module to obtain the registered image of the fixed image.

[0059] Refer to Figure 3 As shown in the figure, the decoding module provided by the embodiment of the present invention corresponds to the fixed image encoding module and the moving image encoding module, and includes a first decoding block, a second decoding block, a third decoding block, a fourth decoding block and a fifth decoding block connected in series in sequence to realize the dimension restoration from low resolution to high resolution. Each decoding block includes a multi-scale deformation field fusion module, a spatial transformer module and a convolution unit connected in series along the forward propagation direction. The spatial transformation module is used to learn the input features to obtain transformation parameters, apply the transformation parameters to the moving image, and map the moving image features of the same size under the guidance of the deformation field to register with the fixed image to obtain the registered image. The multi-scale fusion module is designed to fuse multi-layer deformation field information to enhance fine-grained positioning. The convolution unit is mainly used to fuse multi-semantic context information.

[0060] The multi-scale deformation field fusion module in the first decoding block convolves the input features to obtain the first deformation field, and upsamples the first deformation field to obtain the first output deformation field. The multi-scale deformation field fusion module in the nth decoding block convolves the input features to obtain the nth deformation field, upsamples the nth deformation field and the (n - 1)th deformation field respectively, and superimposes the upsampled results to obtain the nth output deformation field, where 2 ≤ n ≤ 5.

[0061] Specifically, the outputs of the fifth fixed encoding block of the fixed image encoding module and the fifth mobile encoding block of the mobile image encoding module are both input to the first decoding block. Combining with the output of the multi-head self-attention transformation module, a first feature image is generated and input to the second decoding block; the second decoding block generates a second feature image according to the outputs of the fourth fixed encoding block and the fourth mobile encoding block and the output of the first decoding block; the third decoding block generates a third feature image according to the outputs of the third fixed encoding block and the third mobile encoding block and the output of the second decoding block; the fourth decoding block generates a fourth feature image according to the outputs of the second fixed encoding block and the second mobile encoding block and the output of the third decoding block; the fifth decoding block generates a fifth feature image according to the outputs of the first fixed encoding block and the first mobile encoding block and the output of the fourth decoding block, and inputs it to the multi-scale deformation field fusion module; so that the multi-scale deformation field fusion module can obtain multiple registered images of the fixed image according to the registration deformation field.

[0062] S6: Repeatedly input the fixed image and the unregistered mobile image into the registration network to obtain the registered image of the fixed image;

[0063] S7: Superimpose and average the fixed image and multiple registered images to obtain the denoised image of the fixed image.

[0064] Specifically, referring to Figure 3 As shown, in order to further understand the potential relationship between the fixed image and the image to be registered, a multi-scale deformation field fusion module is proposed. The multi-scale deformation field fusion module uses a multi-resolution strategy to learn the details of global affine registration, and uses the low-resolution and strong semantic features of the previous layer as the input F in , to obtain a high-resolution deformation field φ out , and a high-resolution feature F out . In addition to the first decoding block, the other four decoding blocks also require the deformation field φ in of the upper layer as the input.

[0065] Specifically, each decoding block performs a convolution operation with a convolution kernel of 3 on the input feature F in to obtain the deformation field φ conv :

[0066] φ conv = Conv 3×3 (F in ),

[0067] For the first decoding block, the output deformation field φ out is obtained by upsampling φ conv , that is:

[0068] φ out = φ up-conv = UpSample(φ conv),

[0069] For the second decoding block, the third decoding block, the fourth decoding block, and the fifth decoding block, it is necessary to upsample φ conv to obtain φ up-conv After that, the deformation field φ of the previous layer in is also used as an input, and φ is upsampled in to obtain the deformation field φ up-in The output deformation field φ of the final module out :

[0070] φ up-conv = UpSample(φ conv ), φ up-in = UpSample φ in ), φ out = φ up-conv + φ up-in ;

[0071] The spatial transformation module uses a differentiable spatial transformation module based on a spatial transformation network to construct a distorted image. This module learns to regress the transformation parameters of the input features, then applies the estimated transformation parameters to the input features for mapping, and finally constructs the final distorted image through a sampler. Specifically, by inputting the feature map to be registered and the deformation field parameters, the image obtained by distorting the image to be registered according to the deformation field parameters can be output. This module is model-independent and can perform spatial transformation on each pixel without any additional training supervision and modification to the optimization process.

[0072] Based on the above embodiments, in the embodiments of the present invention, before step S1, it further includes collecting OCT images of multiple retinal samples, resampling them, constructing an image data set, and dividing the image data set into a training set and a test set according to a preset ratio; after step S6, it further includes optimizing the registration network according to the loss function of the fixed image and the registered image; selecting any OCT image in a certain sample of the test set as the fixed image, using the remaining OCT images in the sample as the moving images, repeatedly inputting the fixed image and different moving images into the optimized registration network, and obtaining multiple registered images of the fixed image in the sample of the test set; superimposing and averaging the fixed image and all the registered images of the sample to obtain a denoised image.

[0073] Specifically, by calculating the loss function of the fixed image and the registered moving image, the image registration network is optimized, and the loss function is: L joint = L sim + αL smooth ,

[0074]

[0075]

[0076] Among them, L sim is used to constrain the smoothness in the deformation field, and L smooth is used to constrain the local spatial variation in the deformation field. α is a regularization parameter, which is set to 10; represents the k-th layer label of the segmentation map L of the fixed image layer f ; represents the k-th layer label of the segmentation map L of the registered image layer w , where k represents the number of layers in the layer segmentation label; φ represents the deformation field, p represents any point in the deformation field, μ(p) represents the value of the coordinates (x, y) of point p in the deformation field, the gradient of p in the x direction of the deformation field the gradient of p in the y direction of the deformation field

[0077] The OCT image enhancement method based on the registration network provided by the present invention uses the registered images for multi-frame average denoising, effectively solving the problem of image information masking existing in single-frame denoising; compared with the traditional single-frame denoising method, the calculation efficiency is greatly improved; compared with the single-frame denoising method based on deep learning, it is not limited by the quality of the denoising gold standard in the training data. Moreover, the present invention constructs an image registration network by combining a multi-head self-attention conversion module and a multi-scale deformation field fusion module; using the multi-head self-attention conversion module, captures more relationships between the fixed image and the moving image, and realizes high-precision registration of the fixed image and the moving image; uses the multi-scale deformation field fusion module to learn the details of global affine registration using a multi-resolution strategy, and based on the deformation field of the previous layer feature and the feature output has a high-resolution deformation field, enhancing fine-grained localization, enhancing the volume overlap between anatomical structures, and reducing the number of folded voxels between the registration deformation fields, avoiding the loss of feature information. Therefore, when performing average denoising, artifacts are reduced, the image is enhanced, and high-precision image registration is achieved.

[0078] Based on the above embodiments, the present invention also provides an application of the OCT image enhancement method based on the registration network in the field of retinal OCT image enhancement.

[0079] When performing retinal OCT image enhancement, the fixed layer segmentation map of the selected fixed image and the moving layer segmentation map of the moving image are input into the image registration network; according to the loss function of the fixed layer segmentation map and the moving layer segmentation map, the registration network is optimized; the segmentation layers of the fixed layer segmentation map and the moving layer segmentation map both include a nerve fiber layer, a ganglion cell layer to an outer plexiform layer, an outer nuclear layer, an ellipsoid zone to a retinal pigment epithelium layer, and a background region layer.

[0080] Referring to Figure 4 as shown, the retinal OCT image enhancement method based on the registration network according to the embodiment of the present invention includes:

[0081] Based on the above embodiments, in this embodiment, data is acquired using five commercial OCT scanners. The OCT images are resampled to 1024×1024, and the data is proportionally divided into a training dataset D tr and two test datasets D ts1 and D ts2 . The training dataset D tr is the image data of 248 eyes collected in a single-line scanning manner from a BV1000 scanner. The data size for each eye is 1000×1024×20 (width×height×number of slices). The first test dataset D ts1 is the data of 62 eyes collected in the same manner as D tr . The second test dataset D ts2 is the image data of a total of 48 eyes obtained by four other scanners including Topcon DRI-1, Topcon1000, Topcon 2000, and Zeiss in an area scanning mode. The specific distributions of the training dataset, the first test dataset, and the second test dataset are shown in Table 1

[0082] Table 1: Specific distributions of the training dataset, the first test dataset, and the second test dataset

[0083]

[0084] Online data augmentation is performed on the data in the above datasets, including horizontal flipping, horizontal translation with a width range of 2.5%, vertical translation with a height range of 20%, rotation with a range of 15°, random blurring, and local adaptive histogram equalization. The training and testing of the model are completed based on the integrated environment of Pytorch and an NVIDIA RTX 3060 GPU with 12GB of storage space. The model is trained by minimizing the cross-entropy loss through the backpropagation algorithm, and the optimizer Adam is used to minimize the cost function. The basic learning rate and weight decay are both set to 0.00004. The batch size is set to 1, and the number of epochs is set to 40000

[0085] To quantitatively evaluate the performance of the image registration network provided by the present invention, 4 commonly used classification evaluation metrics are used, including the signal-to-noise ratio SNR, the contrast-to-noise ratio CNR, the speckle suppression index SSI, and the edge preservation index EPI. Their definitions are as follows

[0086] Signal-to-noise ratio:

[0087] Contrast-to-noise ratio:

[0088] Noise suppression index:

[0089] Edge-preserving coefficient:

[0090] where σ s and σ b are the standard deviations of the signal region and the background region respectively. In CNR, S is the number of regions of interest, u i and σ i are the mean and the standard deviation of the i-th region of interest respectively. In SSI measures the speckle intensity. In EPI, I 0 and I d represent the noisy image and the denoised image respectively, and i and j represent the vertical and horizontal coordinates in the image.

[0091] To test whether the present invention helps to reduce the speckle noise prominent in retinal OCT, the present invention was tested on two test data sets. First, the network was registered and trained. For each sample in the training set, any one was selected as the fixed image, and the other images of the sample were used as the moving images. By calculating the loss function, the model was optimized and saved. During the test, the fixed image in the test sample was selected, the entire sample was fed into the saved network structure, and the saved network parameters were used to obtain the fixed image of each sample and the registered moving image; all the images were superimposed and averaged to obtain the final denoised image.

[0092] The performance of the OCT image enhancement method based on the registration network proposed by the present invention was compared with other methods, including block-matching and 3D filtering (BM3D), KSVD, multi-frame averaging denoising methods based on dynamic programming and hill climbing (DP+HC), Edge-cGAN and Mini-cGAN and DHNet. In these experiments, the parameters of each method were set to achieve the best results.

[0093] Example denoising result figures are shown in Figure 5 In the first test data set D ts1Above, through the registration process, motion artifacts generated by physiological characteristics such as tremors, drifts, saccades, and the scanning pattern of the system optical path during shooting are removed. Thus, we can average multiple registered frames acquired at the same position to obtain a denoised image on the test set. Compared with the original image, the visual quality is significantly improved, and the overall structure and details of the retina can be well preserved, achieving the best signal-to-noise ratio. In terms of measurement metrics, the signal-to-noise ratio of this method has increased from 7.32 dB to 34.85 dB. Although the average values of CNR, SSI, and EPI do not reach the best results, they are very close to the best values. From the image details, it can be seen that the method we proposed can well display information such as posterior vitreous detachment and the outer limiting membrane of the retina, and make the retinal boundary clearer. The test index results are shown in Table 2:

[0094] Table 2: Test evaluation metrics of different denoising methods on the first test dataset

[0095] SNR (dB) CNR (dB) SSI EPI Original Image 0.53±0.43 5.15±0.57 1.00±0.00 1.00±0.00 BM3D 7.32±3.26 13.56±1.54 1.24±0.03 0.37±0.02 KSVD 4.31±1.59 12.71±0.74 0.47±0.03 0.45±0.04 DP + HC 13.60±4.09 10.78±2.15 0.17±0.13 1.17±0.06 Edge - cGAN 17.89±4.33 12.45±1.20 0.11±0.01 0.73±0.06 Mini - cGAN 21.64±1.73 13.47±0.90 0.09±0.01 1.03±0.05 DHNet 14.76±6.40 12.40±1.64 0.12±0.01 0.95±0.07 The Method of the Present Invention 34.85±4.15 11.80±0.96 0.10±0.01 1.00±0.03

[0096] Refer to the example denoising result diagram Figure 6 As shown, on the second test dataset D ts2 Although the model is obtained from the training data collected by the device corresponding to the first test dataset D ts1 excellent results are still obtained on the second test dataset D ts2 Visually, BM3D has a weak contrast and blurred edges, which is also reflected in the lower EPI metric and the highest SSI metric; the image of KSVD is more blurred, with the lowest SNR, and ghosting can be seen from the multi-frame averaging based on dynamic programming and hill climbing; Edge-cGAN, Mini-cGAN, and DHNet are overly smoothed in the retinal layer, and there is adhesion between layers; while our method shows the best overall performance and also obtains the best signal-to-noise ratio result in terms of metric measurement. The test index results are shown in Table 3:

[0097] Table 3: Test evaluation metrics of different denoising methods on the second test dataset

[0098]

[0099]

[0100] Based on the above charts, it can be seen that the OCT image enhancement method based on the registration network proposed by the present invention uses a registration network (MsFTMorph) based on multi-scale fusion and attention mechanism to guide the speckle noise denoising of retinal OCT images, has a high registration accuracy, and obtains the best signal-to-noise ratio. In terms of details, the image registration network provided by the present invention adds two new modules, the multi-head self-attention transformation module MST and the multi-scale deformation field fusion module MsDFF, in order to capture global context information, enhance the volume overlap between anatomical structures, and reduce the number of folded voxels between registration deformation fields. In addition, according to the experimental results of speckle noise suppression on datasets collected by different OCT scanners using different acquisition methods, the method proposed by the present invention retains more retinal structure information and obtains better contrast enhancement. Even for data of different retinal diseases, the present invention also has better performance than other denoising methods, demonstrating the robustness, generalization, and effectiveness of the present invention for speckle denoising.

[0101] The present invention combines the multi-head self-attention mechanism and the convolutional neural network method for the first time in the registration of retinal OCT images, and for the first time proposes a registration method based on multi-scale fusion and attention mechanism to guide the speckle noise denoising of retinal OCT images, which is applicable to the denoising of optical coherence tomography (OCT) images and achieves good denoising effects. The present invention constructs an image registration network by combining the multi-head self-attention transformation module and the multi-scale deformation field fusion module; uses the multi-head self-attention transformation module to capture more relationships between the fixed image and the moving image, and realizes the high-precision registration of the fixed image and the moving image; uses the multi-scale deformation field fusion module to use the multi-resolution strategy to learn the details of global affine registration, and the deformation field based on the features of the previous layer and the feature output have a high-resolution deformation field, which enhances the fine-grained localization, enhances the volume overlap between anatomical structures, and reduces the number of folded voxels between registration deformation fields, avoiding the loss of feature information. Therefore, when averaging denoising, artifacts are reduced, the image is enhanced, and high-precision image registration is achieved, which has better robustness, generalization, and effectiveness for speckle denoising.

[0102] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0103] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one or more of the flows Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0104] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means for implementing the functions specified in one or more of the flows Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0105] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0106] Obviously, the above embodiments are merely examples for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or variations derived therefrom are still within the protection scope of the present invention.

Claims

1. An OCT image enhancement method based on a registration network, characterized in that, it includes: Select any OCT image in a certain sample of the training set as the fixed image, and use the remaining OCT images in this sample as the moving images; Input the fixed image into the fixed image encoding module of the registration network. After five fixed encoding blocks, the encoded output of each fixed encoding block is input into the corresponding decoding block; Input the moving image into the moving image encoding module of the registration network. After five moving encoding blocks, the encoded output of each moving encoding block is input into the corresponding decoding block; The outputs of the third fixed encoding block of the fixed image encoding module and the third moving encoding block of the moving image encoding module both pass through the multi-head self-attention transformation module; the multi-head self-attention transformation module performs convolution operations and flattening operations on the input fixed image feature map and moving image feature map respectively, and splices and fuses the fixed image feature map and the moving image feature map through a concatenation operation. The fused feature map and the image position information are input into the transformer encoding unit, and then flattened and output to the first decoding block of the decoding module through layer normalization and convolution block operations; Each decoding block in the decoding module is based on the output of the previous decoding block and the outputs of the corresponding fixed encoding block and moving encoding block. After restoring the dimension of the input feature image layer by layer from low resolution to high resolution, it outputs to the multi-scale deformation field fusion module to obtain the registered image of the fixed image; Repeat to input the fixed image and the unregistered moving image into the registration network to obtain the remaining registered images of the fixed image; Overlay and average the fixed image and multiple registered images to obtain the denoised image of the fixed image.

2. The OCT image enhancement method based on a registration network according to claim 1, characterized in that, the transformer encoding unit includes a multi-head self-attention block and a multi-layer perceptron; the multi-head self-attention block includes multiple self-attention blocks connected together in a channel manner in the transformer encoding unit.

3. The OCT image enhancement method based on a registration network according to claim 1, characterized in that, both the fixed image encoding module and the moving image encoding module include a first encoding block, a second encoding block, a third encoding block, a fourth encoding block, and a fifth encoding block connected in series in the forward propagation direction; the first encoding block includes two serially connected convolutional units; the convolutional unit includes a convolutional block, a batch normalization layer, and a ReLU layer connected in series in the forward propagation direction; both the second encoding block and the third encoding block include a max pooling unit and two convolutional units connected in series in the forward propagation direction; both the fourth encoding block and the fifth encoding block include a max pooling unit.

4. The OCT image enhancement method based on a registration network according to claim 1, characterized in that, The decoding module includes a first decoding block, a second decoding block, a third decoding block, a fourth decoding block, and a fifth decoding block connected in series in sequence; each decoding block includes a multi-scale deformation field fusion module, a spatial transformation module, and a convolutional unit connected in series in sequence; the multi-scale deformation field fusion module is used to learn the details of global affine registration; the convolutional unit includes a convolutional block, a batch normalization layer, and a ReLU layer connected in series in sequence; the spatial transformation module is used to learn the input features to obtain transformation parameters, apply the transformation parameters to the moving image, and perform mapping to obtain a registered image.

5. The OCT image enhancement method based on a registration network according to claim 4, wherein, the multi-scale deformation field fusion module in the first decoding block performs convolution on the input features to obtain a first deformation field, and performs upsampling on the first deformation field to obtain a first output deformation field; the multi-scale deformation field fusion module in the nth decoding block performs convolution on the input features to obtain an nth deformation field, performs upsampling on the nth deformation field and the (n - 1)th deformation field respectively, and superimposes the upsampled results to obtain an nth output deformation field; where 2 ≤ n ≤ 5.

6. The OCT image enhancement method based on a registration network according to claim 1, wherein, before determining any one OCT image in a certain sample of the training set as the fixed image, it includes collecting OCT images of multiple retinal samples, performing resampling, constructing an image data set, and dividing the image data set into a training set and a test set according to a preset ratio.

7. The OCT image enhancement method based on a registration network according to claim 1, wherein, the obtaining of multiple registered images of the fixed image includes: According to the deformation field parameters output by the multi-scale deformation field fusion module, using the spatial transformation module to distort the moving image to register it with the fixed image, and obtaining the registered image of the fixed image.

8. The OCT image enhancement method based on a registration network according to claim 1, wherein, after obtaining the registered image of the fixed image, it further includes: Optimizing the registration network according to the loss function of the fixed image and the registered image; selecting any one OCT image in a certain sample of the test set as the fixed image, using the remaining OCT images in the sample as the moving image, and repeatedly inputting the fixed image and different moving images into the optimized registration network to obtain multiple registered images of the fixed image in the sample of the test set; superimposing and averaging the fixed image and the multiple registered images of the sample to obtain a denoised image.

9. The OCT image enhancement method based on a registration network according to claim 8, wherein, the loss function of the fixed image and the registered image is: L joint = L sim + αL smooth , Among them, L sim is used to constrain the smoothness in the deformation field, and L smooth is used to constrain the local spatial variation in the deformation field. α is the regularization parameter, which is set to 10; represents the k-th layer label of the segmentation map L f of the fixed image layer, represents the k-th layer label of the segmentation map L w of the registered image layer. k represents the number of layers in the layer segmentation label; φ represents the deformation field, p represents any point in the deformation field, μ(p) represents the value of the coordinates (x, y) of point p in the deformation field, the gradient of p in the x direction of the deformation field the gradient of p in the y direction of the deformation field 10. Application of the OCT image enhancement method based on a registration network according to any one of claims 1 to 9 in the field of retinal OCT image enhancement.

Citation Information

Patent Citations

  • Medical image registration method, electronic equipment and storage medium

    CN112419378A

  • Skin disease image segmentation method and system based on joint attention convolutional neural network

    CN115457021A