Self-supervised smoke removal method for surgical videos
Patent Information
- Application Number
- CN202410169725.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-06
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2044-02-06
AI Technical Summary
[0004]针对腹腔镜手术过程中产生的手术烟雾干扰医生视野的问题,本发明提供一种手术视频的自监督烟雾去除方法
[0044] The beneficial effects of this invention are as follows: The method of this invention fully utilizes the internal characteristics of surgical videos to achieve efficient self-supervised surgical smoke removal. Addressing the difficulty in obtaining clear target videos, the frame before the high-energy instruments begin operation (the pre-smoke frame) is treated as an misaligned supervision signal to complete parameter learning for the self-supervised video smoke removal model. To address the performance limitations in dense smoke scenes, the pre-smoke frame is further used as the model input, where a masking strategy and a regularized loss function are proposed to address overfitting.
Smart Images

Figure CN117876262B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a self-supervised smoke removal method for surgical videos, belonging to the field of image restoration technology. Background Technology
[0002] The frequent use of high-energy instruments during laparoscopic surgery generates surgical smoke, which interferes with the surgical field of vision and affects the surgeon's decision-making.
[0003] In real-world scenarios, clear videos paired with surgical smoke videos are unavailable. Existing methods for obtaining clear videos rely on simulated data for model training. Due to the differences between simulated and real smoke, the performance of existing methods is limited in real-world scenarios. Other methods employ unpaired learning for model training, but these still cannot effectively handle dense smoke scenes. Summary of the Invention
[0004] To address the problem of surgical smoke interfering with the surgeon's vision during laparoscopic surgery, this invention provides a self-supervised smoke removal method for surgical videos.
[0005] The present invention provides a self-supervised smoke removal method for surgical videos, comprising:
[0006] A video smoke removal model is built using a fog frame encoder, a reference frame encoder, an alignment module, a fusion module, and a reconstruction module, employing a smoke video sequence containing N smoke frames. Before the smoke S ref The video smoke removal model is trained to obtain the trained video smoke removal model. In practical applications, the trained video smoke removal model is based on the real smoke pre-frame and the real smoke video sequence to obtain the smoke-removed video sequence.
[0007] Using the smoke pre-frame S ref The method for training a video smoke removal model using supervised information is as follows:
[0008] The alignment module uses an optical flow estimation network based on the smoke-preceding frame S. ref and smoke frame S i Obtain the optical flow offset Ψ from the frame before the smoke to the frame before the smoke. i→ref ;
[0009] Based on optical flow offset Ψ i→ref S before the smoke ref To smoke frame S i Alignment yields the aligned smoke frame S ref→i ;
[0010] Smoke Pre-frame S ref The reference frame feature F is obtained by the reference frame encoder. ref Smoke Frame S iThe smoke frame feature F is obtained by the fog frame encoder. i Based on optical flow offset Ψ i→ref Reference frame features F ref To smoke frame features F i Alignment, obtaining the aligned feature F ref→i ;
[0011] Based on the aligned pre-smoke frame S ref→i and smoke frame S i Obtain the feature mask M i Using feature mask M i For aligned feature F ref→i Correction is performed to obtain the restored reference features.
[0012] The fusion module is based on smoke frame S i Restoration reference features and the adjacent previous smoke frame S i-1 Corresponding fusion feature H i-1 Obtain the fusion feature H i ; Fusion feature H i The smoke frame S is then obtained through the reconstruction module. i Reconstructed frames
[0013] During the training of the video smoke removal model, based on the reconstruction loss function... Regularized loss function and adversarial loss function The combined total loss function Adjust the network parameters of the video smoke removal model.
[0014] The self-supervised smoke removal method for surgical videos according to the present invention is based on the aligned pre-smoke frame S. ref→i and smoke frame S i Obtain the feature mask M i The method is as follows:
[0015] Aligned with smoke before frame S ref→i After dark channel prior processing and Gaussian blurring, the image is divided into P non-overlapping aligned blocks of the previous smoke frame image. Smoke Frame S i After dark channel prior processing and Gaussian blurring, the smoke frame image is segmented into P non-overlapping blocks.
[0016] For image patches and image blocks Perform structural similarity measurement to obtain image patch mask values.
[0017]
[0018] In the formula, SSIM represents the structural similarity measure, and ∈ is the similarity threshold hyperparameter;
[0019] Made from all image patch mask values Obtain the feature mask M i p = 1, 2, 3, ... P.
[0020] The self-supervised smoke removal method for surgical videos according to the present invention reconstructs the loss function. Used to calculate reconstructed frames With the smoke in the previous frame S ref Losses between:
[0021] The alignment module uses an optical flow estimation network to calculate the optical flow offset Ψ from the reconstructed frame to the frame before the smoke. ref→i :
[0022]
[0023] In the formula This represents the optical flow estimation network;
[0024] Reconstructed frames Smoke Before Frame S ref Alignment, then reconstruct the frame after alignment.
[0025]
[0026] In the formula Indicates alignment operation;
[0027] Reconstruction loss function for:
[0028]
[0029] In the formula V i A mask for the effective location of optical flow.
[0030] The self-supervised smoke removal method for surgical videos according to the present invention restores reference features. for:
[0031]
[0032] The self-supervised smoke removal method for surgical videos according to the present invention, with a regularized loss function. Used to suppress restored reference features Overfitting:
[0033]
[0034] The self-supervised smoke removal method for surgical videos according to the present invention, with adversarial loss function Used to train video smoke removal models:
[0035]
[0036] In the formula Let S represent the mathematical expectation, and S represent the smoke video sequence. DISC represents the probability distribution, and DISC represents the discriminator network. This represents a video smoke removal model.
[0037] The self-supervised smoke removal method for surgical videos according to the present invention has a total loss function. for:
[0038]
[0039] In the formula λ reg For the regularization loss weights, λ GAN To counteract the loss of weights.
[0040] The self-supervised smoke removal method for surgical videos according to the present invention, the training loss function of the discriminator network DISC. for:
[0041]
[0042] The self-supervised smoke removal method for surgical videos according to the present invention, in using the pre-smoke frame S ref Based on training the video smoke removal model using supervised information, the previous frame S... ref The smoke removal model was trained on the video using foggy frames as input, resulting in the reconstructed pre-smoke frames. Then reconstruct the frame before the smoke. The optimized video smoke removal model is fine-tuned using the supervised information after training, resulting in the optimal video smoke removal model.
[0043] According to the self-supervised smoke removal method for surgical videos of the present invention, the video smoke removal model is trained using the Adam optimization algorithm.
[0044] The beneficial effects of this invention are as follows: The method of this invention fully utilizes the internal characteristics of surgical videos to achieve efficient self-supervised surgical smoke removal. Addressing the difficulty in obtaining clear target videos, the frame before the high-energy instruments begin operation (the pre-smoke frame) is treated as an misaligned supervision signal to complete parameter learning for the self-supervised video smoke removal model. To address the performance limitations in dense smoke scenes, the pre-smoke frame is further used as the model input, where a masking strategy and a regularized loss function are proposed to address overfitting.
[0045] The method of this invention does not require the acquisition of clear videos paired with smoke videos; it can complete the training of the model solely based on smoke videos in real-world scenes.
[0046] This invention fully utilizes the internal characteristics of surgical videos, enabling efficient training of a video surgical smoke removal model for real-world scenarios. On one hand, compared to methods using simulated data for training, the model trained using this invention can effectively handle surgical smoke in real-world scenarios. On the other hand, compared to unpaired learning methods, the training method in this invention is more stable and can handle densely smoked scenes more efficiently.
[0047] Experiments show that, compared with the most advanced methods currently available, the method of this invention removes smoke more cleanly, restores details more clearly, and exhibits better temporal consistency in the restoration results. Attached Figure Description
[0048] Figure 1 This is a flowchart of the self-supervised smoke removal method for surgical videos described in this invention;
[0049] Figure 2 This is a training framework diagram for a video smoke removal model;
[0050] Figure 3 The feature mask M is obtained using a mask generator. i A schematic diagram;
[0051] Figure 4 It is the frame before the smoke was reconstructed. A schematic diagram of supervised model training;
[0052] Figure 5 This is a schematic diagram of the overfitting phenomenon in a specific embodiment;
[0053] Figure 6 It is a visual comparison of video reconstruction using different models on real data. Figure 1 Ours indicates that the smoke-preceding frame S is used. ref As supervisory information, Ours* indicates that the frame before the reconstructed smoke is used. As supervisory information;
[0054] Figure 7 It is a visual comparison of video reconstruction using different models on real data. Figure 2 Ours indicates that the smoke-preceding frame S is used. ref As supervisory information, Ours* indicates that the frame before the reconstructed smoke is used. As supervisory information;
[0055] Figure 8 It is a visual comparison of video reconstruction using different models on real data. Figure 3 Ours indicates that the smoke-preceding frame S is used. ref As supervisory information, Ours* indicates that the frame before the reconstructed smoke is used. As supervisory information;
[0056] Figure 9 It is a visual comparison of video reconstruction using different models on real data. Figure 4 Ours indicates that the smoke-preceding frame S is used. ref As supervisory information, Ours* indicates that the frame before the reconstructed smoke is used. As supervisory information. Detailed Implementation
[0057] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0058] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.
[0059] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but this is not intended to limit the scope of the invention.
[0060] Specific Implementation Method 1: Combination Figures 1 to 4 As shown, this invention provides a self-supervised smoke removal method for surgical videos, comprising:
[0061] A video smoke removal model is built using a fog frame encoder, a reference frame encoder, an alignment module, a fusion module, and a reconstruction module, employing a smoke video sequence containing N smoke frames. Before the smoke S ref The video smoke removal model is trained to obtain the trained video smoke removal model. In practical applications, the trained video smoke removal model is based on the real smoke pre-frame and the real smoke video sequence to obtain the smoke-removed video sequence.
[0062] Using the smoke pre-frame S ref The method for training a video smoke removal model using supervised information is as follows:
[0063] The alignment module uses an optical flow estimation network based on the smoke-preceding frame S. ref and smoke frame S i Obtain the optical flow offset Ψ from the frame before the smoke to the frame before the smoke. i→ref ;
[0064] Based on optical flow offset Ψi→ref S before the smoke ref To smoke frame S i Alignment yields the aligned smoke frame S ref→i ;
[0065] Smoke Pre-frame S ref The reference frame feature F is obtained by the reference frame encoder. ref Smoke Frame S i The smoke frame feature F is obtained by the fog frame encoder. i Based on optical flow offset Ψ i→ref Reference frame features F ref To smoke frame features F i Alignment, obtaining the aligned feature F ref→i ;F ref→i It is used as guiding information during the restoration process;
[0066] Based on the aligned pre-smoke frame S ref→i and smoke frame S i Obtain the feature mask M i Using feature mask M i For aligned feature F ref→i Correction is performed to obtain the restored reference features.
[0067] The fusion module is based on smoke frame S i Restoration reference features and the adjacent previous smoke frame S i-1 Corresponding fusion feature H i-1 Obtain the fusion feature H i ; Fusion feature H i The smoke frame S is then obtained through the reconstruction module. i Reconstructed frames
[0068] During the training of the video smoke removal model, based on the reconstruction loss function... Regularized loss function and adversarial loss function The combined total loss function Adjust the network parameters of the video smoke removal model.
[0069] This embodiment utilizes the internal characteristics of laparoscopic surgical videos to achieve highly efficient self-supervised surgical smoke removal. Because adjacent frames may contain information complementary to the current frame, the video smoke removal in this embodiment is more efficient than single-frame smoke removal.
[0070] The frame before the high-energy device starts working (the pre-smoke frame) is clearer than the subsequent smoke frames, and the content is largely consistent. Therefore, in this implementation, the pre-smoke frame is used as the misalignment supervision for the subsequent fog frames to complete the training of the self-supervised smoke removal model.
[0071] Furthermore, combined with Figure 3 As shown, based on the aligned pre-smoke frame S ref→i and smoke frame S i Obtain the feature mask M i The method is as follows:
[0072] In regions where optical flow estimation is inaccurate, the model is prone to overfitting to the reference frame, leading to deviations in the restored output from the input in terms of content. This implementation proposes a masking strategy and regularized loss to address this issue.
[0073] Aligned with smoke before frame S ref→i After dark channel prior processing and Gaussian blurring, the image is divided into P non-overlapping aligned blocks of the previous smoke frame image. Smoke Frame S i After dark channel prior processing and Gaussian blurring, the smoke frame image is segmented into P non-overlapping blocks. Dark channel prior and Gaussian blurring can avoid the interference of surgical smoke on the computational mask; for each set of image patches and The goal is to obtain the mask value. This indicates whether there is inaccurate optical flow at that location;
[0074] For image patches and image blocks Perform structural similarity measurement to obtain image patch mask values.
[0075]
[0076] In the formula, SSIM represents the structural similarity measure, ∈ is the similarity threshold hyperparameter; max represents calculating the maximum value, and sgn represents the sign function;
[0077] Made from all image patch mask values Obtain the feature mask M i p = 1, 2, 3, ... P.
[0078] In this embodiment, combined with Figure 2 As shown, the reconstruction loss function Used to calculate reconstructed frames With the smoke in the previous frame S ref Losses between:
[0079] Smoke Pre-frame S refConsidered misaligned supervision, optical flow networks are used to overcome the negative impact of spatial misalignment on model training. Estimated optical flow offset Ψ ref→i The output of the model Smoke Before Frame S ref After alignment, the reconstruction loss function is then applied. Calculation;
[0080] The alignment module uses an optical flow estimation network to calculate the optical flow offset Ψ from the reconstructed frame to the frame before the smoke. ref→i :
[0081]
[0082] In the formula This represents the optical flow estimation network;
[0083] Reconstructed frames Smoke Before Frame S ref Alignment, then reconstruct the frame after alignment.
[0084]
[0085] In the formula Indicates alignment operation;
[0086] Reconstruction loss function for:
[0087]
[0088] In the formula V i ⊙ represents the effective location of optical flow, and ⊙ is the dot product operation.
[0089] On the one hand, due to the inaccuracy of optical flow estimation, F ref→i Cannot be with F i Perfect alignment; on the other hand, F ref→i Compared to F i Containing more clear information, smoke removal models tend to utilize F ref→i This characteristic can cause the model to overfit to F-value. ref→i Inaccurate center alignment in certain areas causes the restored result to deviate from the input in terms of content. This implementation introduces a masking strategy and a regularization loss function to address this issue.
[0090] For areas with significant alignment inaccuracies, this implementation uses a masking strategy to remove F. ref→i Reference features corresponding to regions with inaccurate optical flow are obtained. Used as reference information for restoration;
[0091] Restoration reference features for:
[0092]
[0093] Combination Figure 2 As shown, regions with slight optical flow inaccuracies are not easily detected by the mask. To suppress overfitting in these slightly misaligned regions, this implementation introduces a regularized loss function. Suppress overfitting in these regions:
[0094] Regularized loss function Used to suppress restored reference features Overfitting:
[0095]
[0096] Adversarial loss function Used to train video smoke removal models:
[0097]
[0098] In the formula Let S represent the mathematical expectation, and S represent the smoke video sequence. DISC represents the probability distribution, and DISC represents the discriminator network. This represents a video smoke removal model.
[0099] The total loss function of this implementation method for:
[0100]
[0101] In the formula λ reg For the regularization loss weights, λ GAN To counteract the loss of weights.
[0102] As an example, λ reg =0.05, λ GAN =1.0.
[0103] The training loss function of the discriminator network DISC for:
[0104]
[0105] In surgical videos, the frame before the high-energy instruments begin operating (the pre-smoke frame) is clearer and has a largely consistent scene compared to the subsequent smoke frames. Using the pre-smoke frame as misalignment supervision for the subsequent smoke frames allows for the training of a self-supervised smoke removal model.
[0106] The optimized parameters of the video smoke removal model in this embodiment Represented as:
[0107]
[0108] In the formula This represents the parameters of the video smoke removal model.
[0109] Due to the smoke pre-frame S ref A small amount of smoke may still remain; this embodiment further enhances the smoke pre-frame S. ref To obtain better monitoring information.
[0110] Using the pre-smoke frame S ref Based on the pre-training of the smoke removal model using supervised information, the previous frame S is further used. ref Treating these as foggy frames and feeding them into a pre-trained smoke removal model yields cleaner results. As a better supervised information, the pre-trained smoke removal model is fine-tuned to eliminate S ref The impact of residual smoke.
[0111] Furthermore, combining Figure 4 As shown, in the smoke pre-frame S ref Based on training the video smoke removal model using supervised information, the previous frame S... ref The smoke removal model was trained on the video using foggy frames as input, resulting in the reconstructed pre-smoke frames. Then reconstruct the frame before the smoke. The optimized video smoke removal model is fine-tuned using the supervised information after training, resulting in the optimal video smoke removal model.
[0112] The smoke-before-the-flight frame can be obtained not only during model training but also during testing. This implementation further uses the smoke-before-the-flight frame S... ref As input to the model, it assists in smoke removal, thereby improving performance in dense smoke scenes:
[0113] Further optimized parameters Represented as:
[0114]
[0115] As an example, the video smoke removal model is trained using the Adam optimization algorithm. Specific Implementation Example 1:
[0117] The video smoke removal model is built on a unidirectional recurrent neural network. The fog frame encoder includes a first convolution operation, a first activation operation, a second convolution operation, a second activation operation, and E residual blocks. The reference frame encoder has the same structure as the fog frame encoder.
[0118] The alignment module uses an optical flow estimation network to calculate the optical flow from the previous frame to the current frame, and uses the estimated optical flow to align the temporal features E. i-1 Distort to the current frame to obtain E i-1→i .
[0119] The fusion module consists of a first convolution operation, a first activation operation, and F residual blocks:
[0120] 1) Transfer the current smoke frame features F i Aligned feature H i-1→i Features of the reference frame after masking Connect by channel and use the first convolution operation and the first activation operation to reduce the number of channels;
[0121] 2) Feed the features into F residual blocks in the order of processing;
[0122] 3) Compare the output features with H i-1→i Add them together to fuse the information of the current frame and obtain H. i .
[0123] The reconstruction module includes a first convolution operation, a first activation operation, R residual blocks, a first upsampling operation, a second upsampling operation, a third upsampling operation, a second convolution operation, a second activation operation, and a third convolution operation.
[0124] Except for the first and second convolutions in the fog frame encoder and the reference frame encoder, which have convolution kernels with a stride of 2, all other convolutions have a stride of 1. Specific Implementation Example 2:
[0126] Set E=5, F=60, R=5.
[0127] The fog frame encoder includes a first convolution operation, a first activation operation, a second convolution operation, a second activation operation, a first residual block, a second residual block, a third residual block, a fourth residual block, and a fifth residual block.
[0128] The reference frame encoder and the fog frame encoder have the same structure.
[0129] The alignment module uses an optical flow estimation network to calculate the optical flow from the previous frame to the current frame. The optical flow estimation network uses a pre-trained PWC-Net and the model parameters are fixed during training.
[0130] The fusion module includes the first convolution operation, the first activation operation, the first residual block, the second residual block, the third residual block, the fourth residual block, the fifth residual block, the sixth residual block, the seventh residual block, the eighth residual block, the ninth residual block, the tenth residual block, the eleventh residual block, the twelfth residual block, the thirteenth residual block, the fourteenth residual block, the fifteenth residual block, the sixteenth residual block, the seventeenth residual block, the eighteenth residual block, the nineteenth residual block, the twentieth residual block, the twenty-first residual block, the twenty-second residual block, the twenty-third residual block, the twenty-fourth residual block, the twenty-fifth residual block, the twenty-sixth residual block, the twenty-seventh residual block, the twenty-eighth residual block, the twenty-ninth residual block, the thirtieth residual block, and the thirty-first residual block. Residual block, 32nd residual block, 33rd residual block, 34th residual block, 35th residual block, 36th residual block, 37th residual block, 38th residual block, 39th residual block, 40th residual block, 41st residual block, 42nd residual block, 43rd residual block, 44th residual block, 45th residual block, 46th residual block, 47th residual block, 48th residual block, 49th residual block, 50th residual block, 51st residual block, 52nd residual block, 53rd residual block, 54th residual block, 55th residual block, 56th residual block, 57th residual block, 58th residual block, 59th residual block, 60th residual block.
[0131] The reconstruction module includes a first convolution operation, a first activation operation, a first residual block, a second residual block, a third residual block, a fourth residual block, a fifth residual block, a first upsampling operation, a second upsampling operation, a third upsampling operation, a second convolution operation, a second activation operation, and a third convolution operation.
[0132] The activation operations in the above modules are all LeakyReLU functions.
[0133] The upsampling operations in the above modules are all performed using the PixelShuffle function.
[0134] The residual blocks in the above modules all add the input features to the input features after passing through the first convolution operation, the first ReLU activation operation, and the second convolution operation, and use this as the output.
[0135] The first and second convolutions of the residual blocks in the above module are both convolutions with 64 3×3 kernels, a stride of 1, and padding of 1.
[0136] The first convolution of the fog frame encoder and the reference frame encoder is a convolution with three 3×3 kernels, a stride of 2, and padding of 1.
[0137] The second convolution of the fog frame encoder and the reference frame encoder is a convolution with 64 3×3 convolution kernels, a stride of 2, and padding of 1.
[0138] The first convolution of the fusion module consists of 192 3×3 convolution kernels with a stride of 1 and padding of 1.
[0139] The first convolution of the upsampling module consists of 128 3×3 convolution kernels with a stride of 1 and padding of 1; the second and third convolutions both consist of 64 3×3 convolution kernels with a stride of 1 and padding of 1.
[0140] The discriminator network in this embodiment includes a first convolution, a first activation function, a second convolution, a first batch normalization, a second activation function, a third convolution, a second batch normalization, a third activation function, a fourth convolution, a third batch normalization, a fourth activation function, a fifth convolution, a fourth batch normalization, and a fifth activation function.
[0141] The first convolution in the discriminator consists of 64 4×4 convolution kernels with a stride of 2 and padding of 1.
[0142] The first convolution in the discriminator consists of 128 4×4 convolution kernels with a stride of 2 and padding of 1.
[0143] The first convolution in the discriminator consists of 256 4×4 convolution kernels with a stride of 2 and padding of 1.
[0144] The first convolution in the discriminator consists of 512 4×4 convolution kernels with a stride of 2 and padding of 1.
[0145] The first convolution in the discriminator is a 4×4 convolution kernel with a stride of 1 and padding of 1.
[0146] The activation functions in the discriminator are all LeakyReLU functions.
[0147] Before use, the video smoke removal model is trained, and the training process can be implemented using the Adam optimization algorithm.
[0148] Combination Figures 5 to 9 As shown, compared with existing surgical smoke removal methods, the method of this invention fully utilizes the internal characteristics of surgical videos, enabling stable and efficient training of video surgical smoke removal models in real-world scenarios. Furthermore, this method can more effectively handle scenes with dense smoke, removing smoke more cleanly, restoring details more realistically, and resulting in better video stability. The video smoke removal model obtained by this invention can be deployed in imaging equipment used in laparoscopic surgery, helping surgeons to more clearly observe the surgical field of view.
[0149] While the invention has been described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the invention. Therefore, it should be understood that many modifications can be made to the exemplary embodiments, and other arrangements can be designed without departing from the spirit and scope of the invention as defined by the appended claims. It should be understood that different dependent claims and features described herein can be combined in ways different from those described in the original claims. It is also understood that features described in conjunction with individual embodiments can be used in other described embodiments.
Claims
1. A self-supervised smoke removal method for surgical videos, characterized in that... include, A video smoke removal model is built using a fog frame encoder, a reference frame encoder, an alignment module, a fusion module, and a reconstruction module, employing a smoke video sequence containing N smoke frames. Before the smoke S ref The video smoke removal model is trained to obtain the trained video smoke removal model; In practical applications, the trained video smoke removal model obtains the smoke-removed video sequence based on the real smoke-before frame and the real smoke video sequence. Using the smoke pre-frame S ref The method for training a video smoke removal model using supervised information is as follows: The alignment module uses an optical flow estimation network based on the smoke-preceding frame S. ref and smoke frame S i Obtain the optical flow offset Ψ from the frame before the smoke to the frame before the smoke. i→ref ; Based on optical flow offset Ψ i→ref S before the smoke ref To smoke frame S i Alignment yields the aligned smoke frame S ref→i ; Smoke Pre-frame S ref The reference frame feature F is obtained by the reference frame encoder. ref Smoke Frame S i The smoke frame feature F is obtained by the fog frame encoder. i Based on optical flow offset Ψ i→ref Reference frame feature F ref To smoke frame features F i Alignment, obtaining the aligned feature F ref→i ; Based on the aligned pre-smoke frame S ref→i and smoke frame S i Obtain the feature mask M i Using feature mask M i For aligned feature F ref→i Correction is performed to obtain the restored reference features. The fusion module is based on smoke frame S i Restoration reference features and the adjacent previous smoke frame S i-1 Corresponding fusion feature H i-1 Obtain the fusion feature H i ; Fusion feature H i The smoke frame S is then obtained through the reconstruction module. i Reconstructed frames During the training of the video smoke removal model, based on the reconstruction loss function... Regularized loss function and adversarial loss function The combined total loss function Adjust the network parameters of the video smoke removal model.
2. The self-supervised smoke removal method for surgical videos according to claim 1, characterized in that, Based on the aligned pre-smoke frame S ref→i and smoke frame S i Obtain the feature mask M i The method is as follows: Aligned with smoke before frame S ref→i After dark channel prior processing and Gaussian blurring, the image is divided into P non-overlapping aligned blocks of the previous smoke frame image. Smoke Frame S i After dark channel prior processing and Gaussian blurring, the smoke frame image is segmented into P non-overlapping blocks. For image patches and image blocks Perform structural similarity measurement to obtain image patch mask values. In the formula, SSIM represents the structural similarity measure, and ∈ is the similarity threshold hyperparameter; Made from all image patch mask values Obtain the feature mask M i p = 1, 2, 3, ... P.
3. The self-supervised smoke removal method for surgical videos according to claim 2, characterized in that, Reconstruction loss function Used to calculate reconstructed frames With the smoke in the previous frame S ref Losses between: The alignment module uses an optical flow estimation network to calculate the optical flow offset Ψ from the reconstructed frame to the frame before the smoke. ref→i : In the formula This represents the optical flow estimation network; Reconstructed frames Smoke Before Frame S ref Alignment, then reconstruct the frame after alignment. In the formula Indicates alignment operation; Reconstruction loss function for: In the formula V i A mask for the effective location of optical flow.
4. The self-supervised smoke removal method for surgical videos according to claim 3, characterized in that, Restoration reference features for:
5. The self-supervised smoke removal method for surgical videos according to claim 4, characterized in that, Regularized loss function Used to suppress restored reference features Overfitting:
6. The self-supervised smoke removal method for surgical videos according to claim 5, characterized in that, Adversarial loss function Used to train video smoke removal models: In the formula Let S represent the mathematical expectation, and S represent the smoke video sequence. DISC represents the probability distribution, and DISC represents the discriminator network. This represents a video smoke removal model.
7. The self-supervised smoke removal method for surgical videos according to claim 6, characterized in that, Total loss function for: In the formula λ reg For the regularization loss weights, λ GAN To counteract the loss of weights.
8. The self-supervised smoke removal method for surgical videos according to claim 7, characterized in that, Training loss function of the discriminator network DISC for:
9. The self-supervised smoke removal method for surgical videos according to claim 8, characterized in that, Using the pre-smoke frame S ref Based on training the video smoke removal model using supervised information, the previous frame S... ref The smoke removal model was trained on the video using foggy frames as input, resulting in the reconstructed pre-smoke frames. Then reconstruct the frame before the smoke. The optimized video smoke removal model is fine-tuned using the supervised information after training, resulting in the optimal video smoke removal model.
10. The self-supervised smoke removal method for surgical videos according to claim 1, characterized in that, The video smoke removal model was trained using the Adam optimization algorithm.