Medical image defogging method based on transfer learning and optical flow estimation

Through transfer learning and optical flow estimation technology, combined with the dual-domain feature fusion mechanism, the problem of insufficient feature extraction and fusion of existing medical imaging defogging methods is solved, and more efficient defogging performance and stronger generalization capabilities are achieved, which significantly improves the clarity and safety of the surgical field of view.

CN120047358APending Publication Date: 2025-05-27FUJIAN MEDICAL UNIV UNION HOSPITAL +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510122848.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing medical image defogging method has insufficient feature extraction and fusion in the spatial and frequency domains, and it is difficult to utilize the timing correlation and continuous features in the video sequence, resulting in limited defogging performance.

Method used

The transfer learning strategy is adopted, based on pre-trained Ef-RAFT and DFFNet networks, combined with optical flow estimation technology, and the precise alignment of video frame sequences is achieved, relatively foggy frames are used to assist fogging frames for defogging, and an efficient dual-domain feature fusion mechanism is designed.

Benefits of technology

It significantly improves the performance of medical imaging defog removal, improves the clarity of the surgical field, reduces the risk of surgery, and has strong generalization capabilities. It is suitable for other medical imaging processing scenarios that require precise visual guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047358A_ABST
    Figure CN120047358A_ABST
Patent Text Reader

Abstract

The invention provides a medical image defogging method based on transfer learning and optical flow estimation, and the method comprises the steps: inputting a foggy medical image for defogging in a transfer learning mode based on an Ef-RAFT network obtained through the pre-training of an optical flow estimation data set and a DFFNet network obtained through the pre-training of a foggy image data set; reconstructing the defogged frame to obtain a defogged medical image; the DFFNet network comprises an encoder, a frequency domain feature fusion module, a decoder and a soft reconstruction module. In combination with an optical flow estimation technology, accurate alignment of video frame sequences is realized, and a relatively fogless frame is utilized to assist a foggy frame in defogging.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of medical image processing, image dehazing, etc., and particularly relates to a medical image dehazing method based on transfer learning and optical flow estimation. Background Technique

[0002] Eliminating the fog during the operation can ensure a clear view for the surgeon, improve the surgical accuracy, reduce the operation errors caused by fog occlusion, and reduce the complication risk. Therefore, developing an efficient and reliable medical image dehazing method is of great significance for improving the surgical safety and intelligent level.

[0003] In the field of image dehazing, traditional methods are based on priors, such as Dark Channel Prior (DCP), Color Attenuation Prior (CAP), etc. Usually, the Atmospheric Scattering Model (ASM) is used to model the degradation process of hazy images. These priors are based on empirical statistics. Therefore, when the scene does not meet the assumptions of these priors, these methods often perform poorly. Learning-based methods rely on datasets to restore the potential haze-free images. With the emergence of large-scale hazy image datasets such as RESIDE, recent methods tend to directly restore haze-free images. Some existing works attempt to combine the advantages of Transformer and CNN to design dehazing models. DeHamer proposes to modulate CNN features under the condition of Transformer features by learning a modulation matrix, which has the global context modeling ability of Transformer and the local representation ability of CNN. DehazeFormer modifies the key designs in SwinTransformer that are not suitable for image dehazing and proposes the improved DehazeFormer. MixDehazeNet proposes a Transformer-style hybrid structure block, which introduces a large receptive field and multi-scale features through a multi-scale parallel large convolutional kernel module, and processes the uneven fog distribution through an enhanced parallel attention module.

[0004] The prior art also attempts to use the differences between clear / degraded image pairs in the frequency domain for dehazing. The differences between haze-free / hazy image pairs in the frequency domain are mostly in the low frequency and a small part in the high frequency. Since the low frequency is mainly restored based on the spatial domain, the research focus of frequency-domain based image dehazing is on how to extract the high-frequency features in the input features and how to fuse the high-frequency features to better restore the high-frequency subbands in the image. FocalNet proposes a frequency selection module, which obtains the low-frequency features in the input features by using global average pooling, subtracts the low-frequency features from the input features to obtain the high-frequency features, and uses the element-wise multiplication of the high-frequency features and the input features and residual connections to obtain the output features. FSNet proposes a multi-branch dynamic frequency selection module and a multi-branch compact frequency selection module, which decouple the features into different frequency features by using multi-scale convolution and pooling respectively, and use SKFusion[2] and addition to reconstruct the output features. ChaIR introduces an implicit frequency domain channel attention module that uses convolution to amplify high-frequency features and uses SKFusion [2] to fuse high-frequency features to reconstruct the output features.

[0005] Most existing defogging methods are limited to single-frame image processing and fail to fully utilize the temporal correlation and continuous features between frames in a video sequence. In recent years, with the development of video processing technology, optical flow-based methods have provided new ideas. RAFT draws inspiration from variational methods, combines multi-scale recurrent gate units and correlation matrices, and improves flow estimation through iterative refinement. Flow1D introduces an attention-based method to estimate the similarity between pixel pairs and enhances the estimation effect by combining vertical and horizontal correlations. Flowformer innovatively replaces the correlation matrix with a correlation memory, constructs a storage structure using pixel labels, and calculates correlation values based on attention. Craft uses an attention mechanism to address the impact of noise on the correlation matrix and extracts features with greater global consistency and semantic stability. To address the challenges posed by occlusion, GMA generalizes information to occluded regions through an attention mechanism to obtain a more comprehensive understanding of the scene. When dealing with repetitive patterns, due to the difficulty of pixel correspondence caused by similar feature vectors, GMFlowNet uses attention to identify possible correct correspondences and shows obvious advantages in large displacement scenarios. Ef-RAFT [3] proposes an amorphous search operator as a new method to solve the large displacement problem of optical flow estimation methods, and at the same time proposes an attention-based feature locator as a new method to solve the repetitive pattern problem of optical flow estimation methods.

[0006] At the same time, there are two main problems with existing single-frame defogging methods: insufficient spatial domain feature extraction and fusion, and poor frequency domain feature fusion effect, which to some extent limits the further improvement of defogging performance. In addition, it is difficult to produce a foggy medical image training set, which requires huge human and time costs, and it is difficult to obtain a fog-free ground truth image corresponding to the foggy medical image.

[0007] References:

[0008] [1]Wang W,Xie E,Li X,et al.Pvt v2:Improved baselines with pyramidvision transformer[J].Computational Visual Media,2022,8(3):415-424.

[0009] [2]Song Y, He Z, Qian H, et al. Vision transformers for single image dehazing[J]. IEEE Transactions on Image Processing, 2023, 32: 1927 - 1941.

[0010] [3]Eslami N, Arefi F, Mansourian A M, et al. Rethinking RAFT for efficient optical flow[C] / / 2024 13th Iranian / 3rd International Machine Vision and Image Processing Conference (MVIP). IEEE, 2024: 1 - 7。 Summary of the Invention

[0011] In the context of the deep integration of technology and the medical field, computer vision technology is reshaping the development pattern of modern medicine. In clinical scenarios such as surgeries, when using high - temperature devices such as electrosurgical knives and lasers for tissue cutting, the water and fat in the tissue will quickly vaporize to form smoke. This smoke not only affects the doctor's vision but also seriously interferes with the real - time detection and tracking of surgical targets by intelligent assistance systems based on computer vision. Therefore, developing efficient and reliable medical image de - hazing methods is of great significance for improving surgical safety and the level of intelligence. Limited by the scarcity of medical image datasets and annotation costs, it is difficult to directly train a dedicated de - hazing model. Most existing de - hazing methods are limited to single - frame image processing and fail to fully utilize the temporal correlation and continuous features between frames in video sequences. At the same time, existing de - hazing methods do not fully extract and fuse features in the spatial domain and frequency domain, which to a certain extent limits the further improvement of de - hazing performance. Therefore, this invention adopts a transfer learning strategy. Based on the pre - training of the dataset, combined with optical flow estimation technology to achieve precise alignment of video frame sequences, and uses relatively haze - free frames to assist haze - filled frames for de - hazing. At the same time, this invention proposes an efficient dual - domain feature fusion mechanism for more sufficient extraction and fusion, significantly improving the de - hazing performance of the model. Experimental results show that the model proposed in this invention demonstrates excellent performance in surgical video processing, significantly improving the clarity of the surgical field of view, providing strong support for the precise operation of surgeons, and thus greatly reducing surgical risks. In addition, the technical solution of this invention has strong generalization ability and can be extended to other medical image processing scenarios that require precise visual guidance, showing important clinical application value in aspects such as surgical safety guarantee and intelligent auxiliary decision - making.

[0012] The technical solution specifically adopted by the present invention to solve its technical problems is as follows:

[0013] A medical image dehazing method based on transfer learning and optical flow estimation: Based on the Ef-RAFT network pre-trained through an optical flow estimation dataset, and the DFFNet network pre-trained through a hazy image dataset, the hazy medical image is input for dehazing in a transfer learning manner, and the dehazed medical image is obtained by reconstructing the frame after dehazing processing; the DFFNet network includes an encoder, a frequency domain feature fusion module, a decoder, and a soft reconstruction module.

[0014] Further, the encoder specifically is: The input hazy image h ∈ R 3×H×W is subjected to block embedding operation to obtain the block-embedded feature ψ ∈ R 24×H×W ;

[0015] The feature ψ is input into the first-scale encoder composed of N spatially domain feature fusion modules connected in series;

[0016] In the multi-head self-attention part of the spatially domain feature fusion module, first, the feature ψ is subjected to batch normalization, 1×1Conv convolution, and Gaussian error linear unit operation to obtain the intermediate feature ψ 1 ; The large kernel attention and pixel attention are respectively used to selectively focus on the global feature and local feature of ψ 1 , and the features enhanced by the large kernel attention and pixel attention are concatenated and input into the convolutional feed-forward neural network CFFN [1] ; The output feature is connected with the input feature ψ by residual connection to obtain the output ψ 2 of the multi-head self-attention part; In the feed-forward neural network part of the spatially domain feature fusion module, first, ψ 2 is subjected to batch normalization, and then the convolutional feed-forward neural network CFFN [1] is used for feature mapping and the output feature is connected with ψ 2 by residual connection to obtain the final output ψ 3 of the spatially domain feature fusion module; The output h 1 of the first-scale encoder is obtained through N spatially domain feature fusion modules connected in series;

[0017] h 1 is downsampled and then input into the second-scale encoder composed of N spatially domain feature fusion modules to obtain the output h 2 of the second-scale encoder;

[0018] h 2 is downsampled and then input into the third-scale encoder composed of N spatially domain feature fusion modules as the final output h o of the encoder.

[0019] Furthermore, in the frequency domain feature fusion module, three convolutional blocks are stacked to amplify the high-frequency features in the encoder output h o and enrich the diversity of high-frequency features, obtaining three kinds of high-frequency features {u 1 , u 2 , u 3}; each of the convolutional blocks consists of 3×3Conv and GELU;

[0020] The three kinds of high-frequency features are input into the multi-branch channel attention MCA to calculate the three corresponding channel attention weights v i for each high-frequency feature, where i = 1, 2, 3; i

[0021] Each high-frequency feature is weighted and fused through its corresponding channel attention weight, and feature refinement is performed through 1×1Conv to obtain the initial feature fusion h t ;

[0022] The initial feature fusion h t is further subjected to frequency feature fusion to obtain the corresponding attention weight V i ; the output feature h p of the frequency domain feature fusion module is obtained,

[0023] Furthermore, the decoder inputs the output h p of the frequency domain feature fusion module into the third-scale decoder of the decoder to obtain the output h 3 of the third-scale decoder; the third-scale decoder is composed of N spatially domain feature fusion modules connected in series;

[0024] An upsampling operation is performed on h 3 , and it is input together with the output h 2 of the second-scale encoder into the SKFuion module for feature fusion to obtain the fused feature h 4 ;

[0025] h 4 is input into the second-scale decoder of the decoder to obtain the output h 5 of the second-scale decoder; the second-scale decoder is composed of N spatially domain feature fusion modules connected in series;

[0026] An upsampling operation is performed on h 5 , and it is input together with the output h 1 of the first-scale encoder into the SKFuion module for feature fusion to obtain the fused feature h 6 ;

[0027] h 6 ​The first-scale decoder of the input decoder obtains the output h of the first-scale decoder 7 ; the first-scale decoder is composed of N spatial-domain feature fusion modules connected in series;

[0028] Perform a block de-embedding operation on h 7 to obtain the final output h of the decoder part q .

[0029] Furthermore, the soft reconstruction module decomposes the output h q ∈R 4×H×W of the DFFNet decoder part into K∈R 1 ×H×W and B∈R 3×H×W ;

[0030] Then, the foggy input image h, the decomposed K, and B are used to reconstruct the fog-free image h D in the way of soft reconstruction: h D = Kh + B + h

[0031] Furthermore, the DFFNet network calculates the loss L D between the defogged image h D and the corresponding fog-free ground truth image g, and uses the loss to optimize the DFFNet network parameters through the backpropagation algorithm; during the iterative optimization process, based on the peak signal-to-noise ratio index on the test set, the network weight parameters with the best performance are saved as the source domain pre-training weights for transfer learning

[0032] Furthermore, the specific process of calculating the loss L D between the defogged image h D and the corresponding fog-free ground truth image g is as follows:

[0033] Calculate the spatial-domain loss D of the fog-free ground truth image g and the defogged image h where D(x, y) represents the L 1 distance between x and y, ω i is the weight, a is the hyperparameter, and R i is the i-th hidden feature extracted from ResNet-152;

[0034] Calculate the frequency-domain loss D of g and h where F represents the fast Fourier transform;

[0035] Combine the spatial-domain loss and the frequency-domain loss to obtain the total loss L D of the DFFNet network: β is a hyperparameter.

[0036] Further, the input foggy medical image is dehazed in a transfer learning manner, and the specific process of reconstructing the dehazed frame to obtain the dehazed medical image is as follows:

[0037] Extract a continuous frame sequence from the foggy medical image dataset X at a given sampling rate, and form a frame pair for every two adjacent frames pair = {frame n , frame n+1};

[0038] Calculate the fog index FI of frame n and frame n+1 in the frame pair pair: FI(frame) = -0.54 * dark_channel(frame) + 0.46 * contrast(frame); dark_channel() represents the dark channel value of the image, reflecting the overall brightness distribution and haze degree of the image. contrast() represents the contrast value of the image, reflecting the degree of light and dark change and detail clarity of the image. According to the calculated FI n and FI n+1 , perform the following operations: (1) When both FI n and FI n+1 are greater than the fog threshold 1: Judge that both frame n and frame n+1 are fog-free, and input the two frames into the pre-trained DFFNet network respectively, and perform dehazing in a transfer learning manner to obtain the dehazing results and (2) When FI n or FI n+1 is less than the fog threshold 1: Judge that there is a foggy frame, mark the frame with the higher FI as frame θ , and mark the frame with the lower FI as frame τ ;

[0039] Input frame θ and frame τ into the pre-trained Ef-RAFT network, and calculate the optical flow field flow θ between the two frames in a transfer learning manner;

[0040] Based on the displacement information provided by the optical flow field flow θ , deform the image of frame θ to obtain the aligned frame frame δ ;

[0041] Calculate frame δThe pixel feature mask mask δ , sum along the channel dimension, set the area with a value greater than 0.5 to 1, and set the area less than or equal to 0.5 to 0;

[0042] For frame δ and frame τ perform adaptive fusion to obtain the fused enhanced frame The fusion rule is as follows: (1) For the valid area where the value in the mask mask δ is 1, where represents the photometric consistency confidence between frame δ and frame θ , and λ is the fusion weight coefficient; (2) For the invalid area where the value in the mask mask δ is 0,

[0043] Use the pre-trained DFFNet network to perform defogging on frame θ and respectively through transfer learning to obtain the corresponding defogging results and

[0044] Batch process all frame pairs in the foggy medical image dataset X, and reconstruct all the defogged frame pairs into a clear surgical video X P .

[0045] An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the medical image defogging method based on transfer learning and optical flow estimation as described above.

[0046] A non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the medical image defogging method based on transfer learning and optical flow estimation as described above.

[0047] Compared with the prior art, the present invention and its preferred solutions have at least the following beneficial effects:

[0048] 1) By adopting the transfer learning method, first learn relevant knowledge from the source domain (RESIDE-Indoor and Sinel training sets), and then transfer the trained model to the target domain (foggy medical image dataset) for defogging, which alleviates the problems of scarce medical image datasets and high annotation costs.

[0049] 2) By making full use of the differences between clear / degraded image pairs in the spatial domain and the frequency domain, a spatial domain feature fusion module SFFM and a frequency domain feature fusion module FFFM are designed. The DFFNet constructed based on these two modules effectively solves the problem of insufficient dual-domain feature fusion in the prior art. This network not only has fewer parameters but also achieves state-of-the-art or competitive results on multiple hazy datasets.

[0050] 3) The designed PGVDNet integrates Ef-RAFT and DFFNet, effectively utilizes the temporal features of video sequences through optical flow estimation, and uses the relatively haze-free frames aligned in the frame pair as auxiliary information to better assist the hazy frames to be dehazed through the DFFNet network. Experiments show that this method exhibits excellent dehazing performance on the hazy medical image dataset. Brief Description of the Drawings

[0051] The present invention will be further described in detail below with reference to the drawings and specific embodiments:

[0052] Figure 1 This is the pre-training flowchart of the single-frame dehazing network DFFNet according to the embodiment of the present invention.

[0053] Figure 2 This is the module structure diagram of the single-frame dehazing network DFFNet according to the embodiment of the present invention.

[0054] Figure 3 This is the dehazing flowchart of the video dehazing network PGVDNet according to the embodiment of the present invention.

[0055] Figure 4 This is the dehazing effect diagram of the video dehazing network PGVDNet according to the embodiment of the present invention.

[0056] Figure 5 This is the comparison diagram of the dehazing effects of the single-frame dehazing network DFFNet and the video dehazing network PGVDNet according to the embodiment of the present invention. Detailed Description of the Embodiments

[0057] To make the features and advantages of this patent more obvious and understandable, specific embodiments are given below for detailed description as follows:

[0058] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further explanations for the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0059] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0060] As Figures 1 - 5 shown, an embodiment of the present invention provides a medical image dehazing method based on transfer learning and optical flow estimation, which is implemented according to the following steps:

[0061] Step S1: Construct a target domain medical image dataset in transfer learning. Taking the high-definition surgical video collected by the Da Vinci surgical robot platform as an example, professional physicians extract video segments with fog generated due to surgical operations to construct a foggy medical image dataset X;

[0062] Step S2: Obtain the source domain datasets required for transfer learning. Obtain the open-source optical flow estimation dataset Sintel, which contains a sequence frame dataset Frame = {frame 1 , frame 2 , …, frame n} and its corresponding optical flow field ground truth dataset Flow = {flow 1 , flow 2 , …, flow n-1}; Obtain the open-source foggy image dataset RESIDE-Indoor, which contains a foggy image training set H = {h 1 , h 2 , …, h m} and its corresponding fog-free ground truth image dataset G = {g 1 , g 2 , …, g m};

[0063] Step S3: Input adjacent frames frame n-1 , frame n in the sequence into the Ef-RAFT network for processing to obtain the optical flow field calculated by the network Calculate and the corresponding optical flow ground truth flow n-1 to obtain the loss L E , and use the loss L ETrain the Ef-RAFT network using backpropagation. Repeat the iterative step S3 for all sequence frames in the Sintel training set (including two rendering versions: Clean Pass and Final Pass) according to the batch size and the specified number of iterations, and save the network weight parameters with the lowest endpoint error (EPE) in the Sintel test set during the iteration process as the pre-trained weights of the source domain for transfer learning;

[0064] Step S4: Input the hazy image h into the encoder part of the DFFNet network for feature encoding to obtain the encoded feature h o ;

[0065] Step S5: Input the encoded feature h o into the frequency domain feature fusion module FFFM of the DFFNet network to amplify and fuse the frequency domain features, and obtain the output h of the FFFM module p ;

[0066] Step S6: Input the output feature h of the frequency domain feature fusion module FFFM p into the decoder part of the DFFNet network for feature decoding to obtain the decoded feature h q ;

[0067] Step S7: Input the decoded feature h q and the hazy image h into the soft reconstruction module of the DFFNet network for image reconstruction to obtain the output image h after haze removal by the DFFNet network D ;

[0068] Step S8: Calculate the loss L D between the haze-removed image h D and the corresponding haze-free ground truth image g, and optimize the DFFNet network parameters using this loss through the backpropagation algorithm. Repeat steps S4 - S8 for all images in the RESIDE-Indoor training set according to the preset batch size and number of iterations. During the iterative optimization process, based on the peak signal-to-noise ratio (PSNR) metric on the RESIDE-Indoor test set, save the network weight parameters with the optimal performance as the pre-trained weights of the source domain for transfer learning;

[0069] Step S9: Construct the PGVDNet network, which concatenates two pre-trained modules: the Ef-RAFT network trained based on the Sintel dataset and the DFFNet network trained based on the RESIDE-Indoor dataset. Input the hazy medical image dataset X into the PGVDNet network for haze removal through transfer learning to obtain the haze-removed surgical video X P .

[0070] In an embodiment of the present invention, in S4, the dehazing network DFFNet is one of the design points proposed by the present invention. The relevant design and working process of its encoder can refer to the following steps:

[0071] Step S41: Embed the input hazy image h ∈ R 3×H×W into blocks to obtain the features ψ ∈

[0072] R 24×H×W ;

[0073] Step S42: Input ψ into the first-scale encoder composed of N spatially domain feature fusion modules SFFM connected in series. The basic value of N in the DFFNet network is 1. When the value of N is larger, the dehazing performance is higher, but it will lead to a decrease in processing speed and an increase in memory occupancy. In the multi-head self-attention part of the spatially domain feature fusion module SFFM, first perform batch normalization BN, 1×1 convolution Conv, and Gaussian error linear unit GELU operations on ψ to obtain the intermediate feature ψ 1 ; respectively use large kernel attention LKA and pixel attention PA to selectively focus on the global features and local features of ψ 1 , splice the features enhanced by LKA and PA, and input them into the convolutional feed-forward neural network CFFN [1] ; connect the output features of CFFN [1] with the input feature ψ by residual connection to obtain the output ψ 2 of the multi-head self-attention part. In the feed-forward neural network part of SFFM, first perform BN on ψ 2 , then use CFFN [1] for feature mapping and connect the output features with ψ 2 by residual connection to obtain the final output ψ 3 of SFFM; after passing through N cascaded SFFMs, obtain the output h 1 of the first-scale encoder;

[0074] Step S43: Input h 1 after downsampling into the second-scale encoder composed of N SFFMs, and the processing process is the same as step S42; after passing through N cascaded SFFMs, obtain the output h 2 of the second-scale encoder;

[0075] Step S44: Input h 2 after downsampling into the third-scale encoder composed of N SFFMs, and the processing process is the same as step S42; after passing through N cascaded SFFMs, obtain the output, which is the final output h o of the encoder part of the DFFNet network;

[0076] In an embodiment of the present invention, in S5, the following steps are further included: The relevant design and working process of the Frequency Domain Feature Fusion Module (FFFM) can refer to the following steps:

[0077] Step S51: In the Frequency Domain Feature Fusion Module (FFFM), stack and use three convolutional blocks (each convolutional block consists of 3×3 Conv and GELU) to amplify the high-frequency features in the encoder output h o and enrich the diversity of its high-frequency features, obtaining three high-frequency features {u 1 , u 2 , u 3};

[0078] Step S52: Input the three amplified high-frequency features into the Multi-branch Channel Attention (MCA) to calculate the attention weights v i corresponding to each high-frequency feature. The operation process of MCA is as shown in the formula: i ∈ {1, 2, 3}. s i represents the intermediate attention weight corresponding to u i , and GAP represents the global average pooling operation. Connect the three intermediate attention weights and activate through Softmax to obtain the three channel attention weights v i corresponding to u i ;

[0079] Step S53: Weightedly fuse each high-frequency feature through its corresponding channel attention weight and refine the features through 1×1 Conv to obtain the initial feature fusion h t , and the process is as shown in the formula:

[0080] Step S54: Input h t into the subsequent frequency feature fusion part. This part adopts a design similar to the initial feature fusion part and uses MCA to fuse multiple high-frequency features to obtain the corresponding attention weight V i . The operation process refers to Steps S52 - S53. Finally, obtain the output feature h p of the FFFM, and the operation process is as shown in the formula:

[0081] In an embodiment of the present invention, in S6, the relevant design and working process of the decoder can refer to the following steps:

[0082] Step S61: Input the output h p of the FFFM into the third-scale decoder of the decoder. This decoder is composed of N SFFMs connected in series, and the processing process is the same as Step S42. After passing h p through the N SFFMs connected in series, obtain the output h3 ;

[0083] Step S62: Upsample h 3 and input it together with the output h 2 of the second-scale encoder into the SKFuion [2] module for feature fusion to obtain the fused feature h 4 ;

[0084] Step S63: Input h 4 into the second-scale decoder of the decoder. This decoder is composed of N SFFMs connected in series. The processing process is the same as that in Step S42. After passing through the N SFFMs connected in series, the output h 4 of the second-scale decoder is obtained 5 ;

[0085] Step S64: Upsample h 5 and input it together with the output h 1 of the first-scale encoder into the SKFuion [2] module for feature fusion to obtain the fused feature h 6 ;

[0086] Step S65: Input h 6 into the first-scale decoder of the decoder. This decoder is composed of N SFFMs connected in series. The processing process is the same as that in Step S42. After passing through the N SFFMs connected in series, the output h 6 of the first-scale decoder is obtained 7 ;

[0087] Step S66: Perform block de-embedding operation on h 7 to obtain the final output h q of the DFFNet decoder part

[0088] In an embodiment of the present invention, in S7, the relevant design and working process of the soft reconstruction module can refer to the following steps:

[0089] Step S71: Decompose the output h q ∈R 4×H×W of the DFFNet decoder part into K∈R 1×H×W and B∈R 3×H×W ;

[0090] Step S72: Reconstruct the haze-free image h D by soft reconstruction according to the input hazy image h and the decomposed K and B. The reconstruction process is as shown in the formula: h D =Kh + B + h;

[0091] In an example of the present invention, in S8, calculate the dehazed image hD The loss L between the corresponding haze-free ground truth image g D , and the specific process of using this loss to optimize the DFFNet network parameters through the backpropagation algorithm is as follows:

[0092] Step S81: Calculate the spatial domain loss of the haze-free ground truth image g and the dehazed image h D The process is as shown in the formula: The process is as shown in the formula: where D(x, y) represents the L distance between x and y 1 , ω i is the weight, the hyperparameter α is set to 0.1, and R i is the u-th hidden feature extracted from ResNet-152;

[0093] Step S82: Calculate the frequency domain loss of g and h D The process is as shown in the formula: The process is as shown in the formula: where F represents the fast Fourier transform;

[0094] Step S83: Combine the spatial domain loss and the frequency domain loss to obtain the total loss L of the DFFNet network D , the process is as shown in the formula: The hyperparameter β is set to 0.1;

[0095] In an example of the present invention, in S9, a PGVDNet network is constructed, and this network cascades two pre-trained modules: Ef-RAFT [3] network trained based on the Sintel dataset and the DFFNet network trained based on the RESIDE-Indoor dataset. The hazy medical image dataset X is input into the PGVDNet network, and dehazing is performed through transfer learning to obtain the dehazed surgical video X P :

[0096] The specific design of the constructed PGVDNet network and the specific process of transfer learning are as follows:

[0097] Step S91: Extract a continuous frame sequence from the hazy medical image dataset X at a sampling rate of 30 fps, and form a frame pair pair = {frame n , frame n+1};

[0098] Step S92: Input the frame pair pair into the PGVDNet network. First, calculate frame n and frame n+1The fogginess index FI. The fogginess index is composed of the dark channel value and the contrast value, and the calculation formula is FI(frame) = -0.54 * dark_channel(frame) + 0.46 * contrast(frame). According to the calculated FI n and FI n+1 , the following operations are performed: (1) When both FI n and FI n+1 are greater than the foggy threshold 1: It is determined that frame n and frame n+1 are both fog-free. The two frames are respectively input into the DFFNet network part in the PGVDNet network (trained based on the RESIDE-Indoor dataset), and defogging is performed through transfer learning to obtain the defogging results and (2) When FI n or FI n+1 is less than the foggy threshold 1: It is determined that there is a foggy frame. The frame with the higher FI is marked as frame θ , and the frame with the lower FI is marked as frame τ ;

[0099] Step S93: Input frame θ and frame τ into the Ef-RAFT network part in the PGVDNet network (trained based on the Sintel dataset), and calculate the optical flow field flow θ between the two frames through transfer learning;

[0100] Step S94: Based on the displacement information provided by the optical flow field flow θ , deform the image of frame θ to obtain the aligned frame frame δ ;

[0101] Step S95: Calculate the pixel feature mask mask δ of frame δ , sum along the channel dimension, set the area with a value greater than 0.5 to 1, and set the area less than or equal to 0.5 to 0;

[0102] Step S96: Perform adaptive fusion on frame δ and frame τ to obtain the fused enhanced frame The fusion rule is as follows: (1) For the valid area (the area with a value of 1) in the mask mask δ , where represents frameδ The photometric consistency confidence with the frame θ , where λ is the fusion weight coefficient with a value of 0.2; (2) For the invalid regions (regions with a value of 0) in the mask δ

[0103] Step S97: Use the DFFNet network part in the PGVDNet network (trained based on the RESIDE-Indoor dataset), and through transfer learning, dehaze frame θ and respectively to obtain the corresponding dehazed results and

[0104] Step S98: Batch process all frame pairs in the hazy medical image dataset X, repeatedly execute steps S92 - S97, and reconstruct all frame pairs dehazed by the PGVDNet network into a clear surgical video X P .

[0105] Figure 1 and Figure 2 are respectively the training flow chart and module structure diagram of the single-frame dehazing network DFFNet proposed in the embodiments of the present invention. Figure 3 is the inference flow chart of the video dehazing network PGVDNet proposed by integrating Ef-RAFT and DFFNet in the embodiments of the present invention. Figure 4 is the dehazing effect diagram of PGVDNet. As can be seen from Figure 4 , even in the case of severe hazy scenes, PGVDNet can effectively restore the clarity and color authenticity of the image. Figure 5 is the comparison diagram of the dehazing effects of DFFNet and PGVDNet. As can be seen from Figure 5 , PGVDNet effectively utilizes the temporal characteristics of the video sequence through optical flow estimation, and uses the relatively haze-free frames aligned in the frame pair as auxiliary information to better dehaze the hazy frames. It achieves a better dehazing effect compared to using DFFNet alone, with less haze residue and more accurate color restoration.

[0106] ​Based on the same inventive concept, the present invention further provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is configured to execute the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is used to implement one or more instructions, specifically for loading and executing one or more instructions in the computer storage medium to implement the above method.

[0107] It should be further noted that, based on the same inventive concept, the present invention further provides a computer storage medium, on which a computer program is stored, and the computer program, when run by a processor, executes the above method. The storage medium may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may, for example, but not be limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or combined with an instruction execution system, apparatus, or device.

[0108] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the field to which the present invention pertains. The "first", "second" and similar terms used in the present invention do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "comprising" or "including" mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0109] As described above, it is only the preferred embodiment of the present invention, and it is not a limitation to the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.

[0110] This patent is not limited to the above best implementation manner. Anyone inspired by this patent can obtain various other forms of medical image dehazing methods based on transfer learning and optical flow estimation. All equal changes and modifications made according to the scope of the patent application of the present invention shall fall within the coverage of this patent.

Claims

1. A medical image dehazing method based on transfer learning and optical flow estimation, characterized by: Based on the Ef-RAFT network obtained by pre-training with an optical flow estimation dataset and the DFFNet network obtained by pre-training with a foggy image dataset, foggy medical images are input for defogging in a transfer learning manner, and the defogged frames are reconstructed to obtain defogged medical images; the DFFNet network includes an encoder, a frequency domain feature fusion module, a decoder and a soft reconstruction module.

2. The medical image dehazing method based on transfer learning and optical flow estimation according to claim 1, characterized in that: The encoder specifically comprises: transforming the input foggy image h∈R 3×H×W Perform block embedding operation to obtain the feature ψ∈R after block embedding 24×H×W ; The feature ψ is input into the first scale encoder which is composed of N spatial domain feature fusion modules connected in series; In the multi-head self-attention part of the spatial domain feature fusion module, the feature ψ is first batch normalized, and the intermediate feature ψ is obtained after convolution 1×1Conv and Gaussian error linear unit operation. 1 ; Use large kernel attention and pixel attention to selectively focus on ψ 1 The global and local features of the multi-head self-attention part are concatenated, and the features enhanced by large kernel attention and pixel attention are input into the convolutional feedforward neural network CFFN; the output features are residually connected with the input features ψ to obtain the output ψ of the multi-head self-attention part 2 In the feedforward neural network part of the spatial domain feature fusion module, firstly, ψ 2 Perform batch normalization, then use the convolutional feedforward neural network CFFN for feature mapping and compare the output features with ψ 2 Perform residual connection to obtain the final output ψ of the spatial domain feature fusion module 3 ; After N serially connected spatial domain feature fusion modules, the output h of the first scale encoder is obtained 1 ; h 1 After downsampling, the second scale encoder composed of N spatial domain feature fusion modules is input to obtain the output h of the second scale encoder. 2 ; h 2 After downsampling, the third scale encoder consisting of N spatial domain feature fusion modules is input as the final output h of the encoder. o .

3. The medical image dehazing method based on transfer learning and optical flow estimation according to claim 2, characterized in that: In the frequency domain feature fusion module, three convolution blocks are stacked to output h o The high-frequency features in the image are amplified and the diversity of high-frequency features is enriched to obtain three high-frequency features {u 1 ,u 2 ,u 3 }; Each of the convolution blocks consists of 3×3Conv and GELU; The three high-frequency features are input into the multi-branch channel attention MCA to calculate each high-frequency feature to obtain u i The corresponding three channel attention weights v i , i=1,2,3; Each high-frequency feature is weightedly fused through its corresponding channel attention weight, and the feature is refined through 1×1Cony to obtain the initial feature fusion h t ; The initial features are fused into h t Further frequency feature fusion is performed to obtain the corresponding attention weight V i ; Get the output feature h of the frequency domain feature fusion module p , 4. The medical image dehazing method based on transfer learning and optical flow estimation according to claim 3, characterized in that: The decoder combines the output h of the frequency domain feature fusion module p The third scale decoder of the input decoder obtains the output h of the third scale decoder 3 ; The third scale decoder is composed of N spatial domain feature fusion modules connected in series; For h 3 The upsampling operation is performed and combined with the output h of the second scale encoder 2 Input them into the SKFusion module for feature fusion, and obtain the fused feature h 4 ; h 4 The second scale decoder of the input decoder obtains the output h of the second scale decoder 5 ; The second scale decoder is composed of N spatial domain feature fusion modules connected in series; For h 5 The upsampling operation is performed and combined with the output h of the first scale encoder 1 Input them into the SKFusion module for feature fusion, and obtain the fused feature h 6 ; h 6 The first scale decoder of the input decoder obtains the output h of the first scale decoder 7 ; The first scale decoder is composed of N spatial domain feature fusion modules connected in series; h 7 Perform block de-embedding operation to obtain the final output h of the decoder part q .

5. The medical image dehazing method based on transfer learning and optical flow estimation according to claim 4, characterized in that: The soft reconstruction module converts the output h of the DFFNet decoder part q ∈R 4×H×W Decompose into K∈R 1×H×W and B∈R 3×H×W : Then, the input foggy image h and the decomposed K and B are used to reconstruct the fog-free image h in a soft reconstruction manner. D :h D =Kh+B+h.

6. The medical image dehazing method based on transfer learning and optical flow estimation according to claim 1, characterized in that: The DFFNet network calculates the dehazed image h D The loss L between the corresponding haze-free ground-truth image g D , and use the loss to optimize the DFFNet network parameters through the back-propagation algorithm; During the iterative optimization process, based on the peak signal-to-noise ratio indicator on the test set, the network weight parameters with the best performance are saved as the source domain pre-training weights for transfer learning.

7. The medical image dehazing method based on transfer learning and optical flow estimation according to claim 6, characterized in that: Calculate the dehazed image h D The loss L between the corresponding haze-free ground-truth image g D The specific process is: Calculate the haze-free true image g and the dehazed image h D The spatial domain loss where D(x, y) represents the L1 distance between x and y, ω i is the weight, α is the hyperparameter, R i is the i-th hidden feature extracted from ResNet-152; Calculate g and h D The frequency domain loss Where F stands for Fast Fourier Transform; Combining spatial domain loss and frequency domain loss Get the total loss L of the DFFNet network D : β is a hyperparameter.

8. The medical image dehazing method based on transfer learning and optical flow estimation according to claim 1, characterized in that: The specific process of defogging the input foggy medical image by transfer learning and reconstructing the defogged frame to obtain the defogged medical image is as follows: Extract continuous frame sequences from the foggy medical image dataset X at a given sampling rate, and combine each two adjacent frames into a frame pair pair = {frame n ,frame n+1 }; Calculate the frame in the frame pair n and frame n+1 The haze index FI is: FI(frame) = -0.54*dark_channel(frame) + 0.46*contrast(frame); dark_channel() represents the dark channel value of the image, and contrast() represents the contrast value of the image; according to the calculated FI n and FI n+1 , perform the following operations: (1) When FI n and FI n+1 When both are greater than the fog threshold 1: judge the frame n and frame n+1 There is no fog in both frames. The two frames are input into the pre-trained DFFNet network respectively, and the defogging is performed through transfer learning to obtain the defogging results. and (2) When FI n or FI n+1 When the FI value is less than the fog threshold 1, it is determined that there is a foggy frame and the frame with a higher FI value is marked as a frame. θ , the frame with lower FI is marked as frame τ ; The frame θ and frame τ Input the pre-trained Ef-RAFT network and calculate the optical flow between two frames through transfer learning θ ; Based on optical flow θ The displacement information provided by the frame θ Perform image deformation to obtain aligned frames δ ; Calculating the frame δ The pixel feature mask mask δ , sum along the channel dimension, set the area with value greater than 0.5 to 1, and set the area with value less than or equal to 0.5 to 0; Frame δ and frame τ Perform adaptive fusion to obtain the fused enhanced frame The fusion rules are as follows: (1) For mask δ The value in the valid area is 1. in Represents frame δ With frame θ is the confidence of photometric consistency between them, and λ is the fusion weight coefficient; (2) for the mask δ Invalid area with value 0 in Using the pre-trained DFFNet network, frame θ and Perform dehazing and obtain the corresponding dehazing result and Batch process all frame pairs in the foggy medical image dataset X, and reconstruct all dehazed frame pairs into clear surgical videos X P .

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the medical image defogging method based on transfer learning and optical flow estimation as described in any one of claims 1 to 8 are implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the medical image defogging method based on transfer learning and optical flow estimation as described in any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Rain and fog removing method for image in water area scene

    CN120318117A