A depth information guided binocular video reflection removal method

The depth-guided binocular video reflection removal method, utilizing binocular cameras and feature fusion technology, solves the reflection artifact problem in scenes obscured by glass, achieving better image quality and performance in computer vision tasks.

CN119888424BActive Publication Date: 2026-04-07BEIJING JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

When processing images or videos of scenes obscured by glass, existing technologies suffer from reflection artifacts that cause clutter and blurring of image content, affecting the performance of computer vision tasks. Furthermore, existing methods lack robustness in dynamic scenes.

Method used

A depth-information-guided binocular video reflection removal method is adopted. The video stream is captured by a binocular camera, and a transmission map depth estimation module and a reflection removal module are used. Combined with cross-viewpoint and cross-frame feature fusion, a gating controller is designed to explore feature relationships and achieve effective reflection removal.

Benefits of technology

It improves the effect of reflection removal, showing better PSNR, SSIM and LMSE indicators, thus improving image quality and performance in computer vision tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119888424B_ABST
    Figure CN119888424B_ABST
Patent Text Reader

Abstract

This invention provides a depth-guided method for reflection removal in stereo video containing reflections, comprising a transmission map depth estimation module and a reflection removal module. The transmission map depth estimation module performs depth perception frame-by-frame on the stereo video stream, decoupling the mixed depth corresponding to the stereo video stream to obtain the true depth corresponding to the transmission layer and the pseudo depth corresponding to the reflection layer. The true depth corresponding to the transmission map is used to guide the reflection removal network, achieving effective reflection removal. The reflection removal module, by incorporating a feature fusion and enhancement module with a unified structure, achieves fusion and enhancement of cross-viewpoint features, depth-guided features, and cross-frame features. Gating controllers are designed for the CVEM, DGEM, and CFEM modules to control the feature relationship exploration range for the three different tasks. This achieves good reflection removal results for a given video stream containing reflections.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a depth-information-guided method for removing reflections in binocular video. Background Technology

[0002] When capturing images or recording videos through scenes obscured by glass, noticeable reflection artifacts are often observed. These artifacts result in cluttered, blurry images with lost details, leading to a decline in image quality and severely impacting the understanding of the image or video content. The presence of reflection artifacts in images or videos can significantly negatively affect downstream computer vision tasks such as image and video detection, tracking, segmentation, and classification. Research on reflection removal can promote the development and application of technologies in autonomous driving, video surveillance, intelligent photography, and special scene shooting.

[0003] With the development of computer vision algorithms, removing reflections from single images has become a common practice. However, reflection removal from single images is challenging, and without additional guidance, there are countless possible solutions. Therefore, introducing auxiliary information to guide reflection removal under these challenging conditions is an effective approach. Recent research has found that depth information is an effective guiding tool for reflection removal. For example, active near-infrared sensors can be used to acquire the depth of a transmitted scene to remove reflections. However, this method requires a specific sensor and a specific range of angles between the shooting direction and the glass, both of which limit its application in more general scenarios. For more widely used commercial monocular cameras, a common method for perceiving the depth of transmitted and reflected scenes is to adjust the focal length to keep the transmitted scene within the field of view, thus blurring the edges of the reflected scene, which is beneficial for reflection identification and removal. For example, using depth estimated from a single image to identify reflection edges provides a feasible solution for reflection elimination. However, when the mixed image contains complex textures, the robustness of the single-image depth estimation model decreases, providing limited information for reflection removal. To perceive depth more robustly, a strategy of capturing a pair of images with different focal settings can be used for depth perception. However, this strategy is limited to static scenes. In more general dynamic scenes, adjusting the focal length to keep up with the movement of objects within the scene is very challenging. Therefore, reflection removal depends on the robustness of depth estimation.

[0004] Compared to monocular cameras, binocular cameras can capture two-view images with known binocular camera parameters simultaneously. This property allows for pixel-level alignment of the two-view images and reflection-invariant flow estimation of the two views. However, these methods neglect the perception of scene depth. For human vision, rich prior knowledge enables us to intuitively distinguish between the true depth generated by a transmitted scene and the pseudo-depth generated by a reflected scene. Inspired by the human binocular vision system, we can design deep learning models through extensive training on binocular video stream data to obtain prior knowledge that can effectively distinguish between true and pseudo-depth. This learned knowledge can then be used to eliminate reflections. Therefore, this patent proposes an end-to-end depth-guided reflection removal network for binocular video reflection removal. This network includes a transmission map depth estimation module and a reflection removal module. Furthermore, binocular videos contain inherent cross-view and cross-frame information. Therefore, this patent utilizes depth guidance, supplemented by cross-view and cross-frame feature fusion and enhancement, to achieve effective reflection removal. To fully leverage this rich guidance, we designed a unified feature fusion and enhancement module to explore the correlation between local and global features, and equipped it with a specially designed gating controller based on the feature points of the three tasks to control the scope of feature relationship exploration for the three different tasks. Summary of the Invention

[0005] The embodiments of the present invention provide a depth information-guided binocular video reflection removal method to solve the technical problems existing in the prior art.

[0006] To achieve the above objectives, the present invention adopts the following technical solution.

[0007] A depth-information-guided binocular video reflection removal method includes:

[0008] Based on the left and right views of multiple reflected video frames, the transmission depth map of the multiple reflected video frames is obtained by processing through the transmission map depth estimation module of the network model; the multiple video frames include the current video frame, the previous video frame, and the subsequent video frame; the transmission depth map includes the left view of transmission depth, the right view of transmission depth, and the left and right views of transmission depth.

[0009] Based on the left and right views of multiple reflection video frames and the corresponding transmission depth maps, feature fusion and reflection removal processing are performed through the reflection removal module of the network model to obtain transmission map video frames.

[0010] The reflection removal module is based on an encoder-decoder structure, which includes a first encoder, a second encoder, a third encoder, and a fourth encoder that are set in parallel with each other, as well as a decoder. It also has a cross-view feature enhancement module, a depth-guided feature enhancement module, and a cross-frame feature enhancement module.

[0011] The first encoder and the second encoder are used to extract features step by step from the left and right views of the current video frame, respectively; the third encoder is used to extract features step by step from the left and right views of the transmission depth of the current video frame; and the fourth encoder is used to extract features step by step from the left and right views and the corresponding transmission depth maps of multiple reflection video frames.

[0012] The cross-view feature enhancement module performs a left-right view feature complementation operation on the outputs of the first and second encoders to obtain enhanced left-right feature maps of the reflected video frame, and aligns the features of the enhanced left-right feature maps of the current video frame with the transmission depth map at the pixel level; the depth-guided feature enhancement module performs depth information guidance and feature integration operations on the outputs of the cross-view feature enhancement module and the third encoder to obtain a feature map enhanced by depth-guided features; the inter-frame feature enhancement module is used to perform a fusion operation on the outputs of the depth-guided feature enhancement module and the fourth encoder to obtain a feature map enhanced by inter-frame features.

[0013] Preferably, the process of obtaining the transmission depth map of the multiple reflected video frames by processing them through a transmission map depth estimation module, based on the left and right views of the multiple reflected video frames, includes:

[0014] Video sequences captured by a binocular camera with multiple reflected video frames. The calculation formula of the transmission map depth estimation module

[0015]

[0016] Obtain the transmission depth map of the reflected video frame. In the formula, where Indicates the index of the video frame.

[0017] Preferably: the first encoder uses the left view of the current video frame. As input, output the left-look feature map of the current video frame. The second encoder uses the right view of the current video frame. As input, output the right-look feature map of the current video frame. The third encoder uses the transmission depth of the left and right views corresponding to the left and right views of the current video frame. As input, output the transmission depth feature map of the current video frame. The fourth encoder will process the preceding video frame, the current video frame, and the subsequent video frames ( ) and the corresponding left and right view transmission depth As input, the output is an inter-frame feature map. ;

[0018] The first encoder, second encoder, third encoder and fourth encoder each have three feature extraction blocks; the b-th feature extraction block of the three feature extraction blocks consists of b convolutional layers and b CBAMs (b = 1,2,3);

[0019] The processing steps of the cross-view feature enhancement module specifically include:

[0020] Set the left-look feature map of the current video frame as input. and right-view feature map Each position The feature vector is ∈ and ∈ Where c represents the feature dimension;

[0021] Left-view feature map of the current video frame and right-view feature map Perform global flat pooling to obtain global feature vectors. and ;

[0022] Through

[0023]

[0024] For eigenvectors and and global feature vectors and Combine them to obtain point feature groups In the formula, Representative feature combination operations;

[0025] Through

[0026]

[0027] Perform feature integration; where, This indicates that each feature group is processed separately using a multi-head attention mechanism. ;

[0028] Construct is used to build Similarity matrix of the relationships between the four components The data is then input into the first gating controller to control the feature extraction direction, thereby obtaining an updated feature map. and ;

[0029] Updated feature map and Perform fusion processing to obtain fused feature maps. ;

[0030] The processing steps of the deep guided feature enhancement module specifically include:

[0031] Set the fused feature map of the input. and the left and right feature maps of the current video frame Each position The feature vector is ∈ and ∈ Where c represents the feature dimension;

[0032] Feature maps of fusion and left and right feature maps Perform global flat pooling to obtain global feature vectors. and ;

[0033] Through

[0034]

[0035] For eigenvectors and and global feature vectors and Combine them to obtain point feature groups In the formula, Representative feature combination operations;

[0036] Through

[0037]

[0038] Perform feature integration; where, This indicates that each feature group is processed separately using a multi-head attention mechanism. ;

[0039] Construct is used to build Similarity matrix of the relationships between the four components The data is then input into the second gating controller to control the feature extraction direction, thereby obtaining an updated feature map. and ;

[0040] Updated feature map and Perform fusion processing to obtain fused feature maps. ;

[0041] The processing steps of the cross-frame feature enhancement module specifically include:

[0042] Set the fused feature map of the input. and inter-frame feature maps Each position The feature vector is ∈ and ∈ Where c represents the feature dimension;

[0043] Feature maps of fusion and inter-frame feature maps Perform global flat pooling to obtain global feature vectors. and ;

[0044] Through

[0045]

[0046] For eigenvectors and and global feature vectors and Combine them to obtain point feature groups In the formula, Representative feature combination operations;

[0047] Through

[0048]

[0049] Perform feature integration; where, This indicates that each feature group is processed separately using a multi-head attention mechanism. ;

[0050] Construct is used to build Similarity matrix of relationships between multiple components The data is then input into the third gating controller to control the feature extraction direction, thereby obtaining an updated feature map. and ;

[0051] Updated feature map and Perform fusion processing to obtain fused feature maps. .

[0052] Preferably, the processing procedure of the first gate controller includes:

[0053] Through

[0054]

[0055] Controlling the direction of feature extraction to obtain updated feature maps and In the formula, Indicates the first gate controller. for Matrix;

[0056] The processing steps of the second gate controller include:

[0057] Through

[0058]

[0059] Controlling the direction of feature extraction to obtain updated feature maps and In the formula, Indicates the second gate controller. It is an anti-diagonal matrix, and each anti-diagonal element of the anti-diagonal matrix has a value of -100;

[0060] The processing steps of the third gate controller include:

[0061] Through

[0062]

[0063] Controlling the direction of feature extraction to obtain updated feature maps and ; Indicates the third gate controller. for The matrix.

[0064] Preferably, the network model also has a loss function, which is expressed as follows:

[0065]

[0066] Calculate the overall loss to train the network model; where, They represent: Pixel loss, SSIM structure loss, reconstruction loss, mean square error loss, and perceptual loss; .

[0067] As can be seen from the technical solutions provided by the embodiments of the present invention above, the present invention provides a method for removing reflections from binocular videos containing reflections based on depth information guidance. This method includes a transmission map depth estimation module and a reflection removal module. The transmission map depth estimation module performs depth perception frame-by-frame on the binocular video stream, decoupling the mixed depth corresponding to the binocular video stream to obtain the true depth corresponding to the transmission layer and the pseudo depth corresponding to the reflection layer. The true depth corresponding to the transmission map is used to guide the reflection removal network, achieving effective reflection removal. The reflection removal module, by embedding a feature fusion and enhancement module with a unified structure, achieves guidance on cross-viewpoint features and depth features, as well as cross-frame feature fusion and enhancement. Gating controllers are designed for the CVEM, DGEM, and CFEM modules to control the feature relationship exploration range of the three different tasks. Ultimately, a good reflection removal effect is achieved for a given video stream containing reflections. The algorithm proposed in this invention shows better PSNR, SSIM, NCC, and LMSE indicators in comparison with existing single-map and multi-map reflection removal algorithms, and also outperforms the compared algorithms in terms of visualization effect after reflection removal.

[0068] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and will become apparent from the description or may be learned by practice of the invention. Attached Figure Description

[0069] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0070] Figure 1 A flowchart illustrating a depth-information-guided binocular video reflection removal method provided by this invention;

[0071] Figure 2 A flowchart of a preferred embodiment of a depth information-guided binocular video reflection removal method provided by the present invention;

[0072] Figure 3 A schematic diagram illustrating the interaction flow of three feature enhancement modules in a depth information-guided binocular video reflection removal method provided by the present invention.

[0073] Figure 4 A comparison diagram showing the effect of the depth information-guided binocular video reflection removal method provided by the present invention with existing algorithms. Detailed Implementation

[0074] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0075] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or couplings. The term “and / or” as used herein includes any and all combinations of one or more of the associated listed items.

[0076] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless defined as herein.

[0077] To facilitate understanding of the embodiments of the present invention, the following will provide further explanation and description with reference to the accompanying drawings and several specific embodiments. These embodiments do not constitute a limitation on the embodiments of the present invention.

[0078] This invention provides a depth-information-guided binocular video reflection removal method to address the following technical problems discovered by the applicant in practice:

[0079] (1) Previous reflection removal algorithms based on single images, multiple images, and monocular videos have neglected the perception principle of human binocular vision in scenes containing reflections. We propose to use a widely used binocular camera to capture dual-view videos to model the human eye's ability to perceive scene depth, thereby providing effective depth guidance information for reflection removal. (2) Previous binocular video reflection removal algorithms have neglected the perception of depth information in the scene. We propose to decouple the hybrid depth estimated from the binocular video to obtain the true depth corresponding to the transmission layer and the pseudo depth corresponding to the reflection layer, and use the true depth corresponding to the transmission map to guide the reflection removal network to achieve effective reflection removal. (3) Existing binocular video reflection removal algorithm models lack a fine design for feature fusion between viewpoint features, frame features, and guidance information features. Designing an effective feature fusion method is crucial for the integration of cross-viewpoint and cross-frame features, as well as the effectiveness of depth feature guidance.

[0080] In view of this, this invention employs a widely used binocular camera to capture left and right view videos for scene depth perception, thereby providing effective depth guidance information for reflection removal. The hybrid depth estimated from the binocular video yields the true depth corresponding to the transmission layer and the pseudo depth corresponding to the reflection layer. The true depth corresponding to the transmission map is used to guide the reflection removal network, achieving effective reflection removal. An effective feature enhancement method is designed in the reflection removal module to fully integrate and enhance cross-view, depth-guided features, and cross-frame features, achieving better binocular video reflection removal results.

[0081] See Figure 1 This invention provides a depth-information-guided binocular video reflection removal method, comprising:

[0082] Based on the left and right views of multiple reflected video frames, the transmission depth map of the multiple reflected video frames is obtained by processing through the transmission map depth estimation module of the network model; the multiple video frames include the current video frame, the previous video frame, and the subsequent video frame; the transmission depth map includes the left view of transmission depth, the right view of transmission depth, and the left and right views of transmission depth.

[0083] Based on the left and right views of multiple reflection video frames and the corresponding transmission depth maps, feature fusion and reflection removal processing are performed through the reflection removal module of the network model to obtain transmission map video frames.

[0084] The reflection removal module is based on an encoder-decoder structure, which includes a first encoder, a second encoder, a third encoder, and a fourth encoder that are set in parallel with each other, as well as a decoder. It also has a cross-view feature enhancement module, a depth-guided feature enhancement module, and a cross-frame feature enhancement module.

[0085] The first encoder and the second encoder are used to extract features step by step from the left and right views of the current video frame, respectively; the third encoder is used to extract features step by step from the left and right views of the transmission depth of the current video frame; and the fourth encoder is used to extract features step by step from the left and right views and the corresponding transmission depth maps of multiple reflection video frames.

[0086] The cross-view feature enhancement module performs a left-right view feature complementation operation on the outputs of the first and second encoders to obtain enhanced left-right feature maps of the reflected video frame, and aligns the features of the enhanced left-right feature maps of the current video frame with the transmission depth map at the pixel level; the depth-guided feature enhancement module performs depth information guidance and feature integration operations on the outputs of the cross-view feature enhancement module and the third encoder to obtain a feature map enhanced by depth-guided features; the inter-frame feature enhancement module is used to perform a fusion operation on the outputs of the depth-guided feature enhancement module and the fourth encoder to obtain a feature map enhanced by inter-frame features.

[0087] By repeatedly performing the above steps and fusing all the obtained transmission video frames, a video with reflections removed is obtained.

[0088] The reflection removal network designed in this invention mainly includes a transmission map depth estimation module and a reflection removal module. This network is based on a CNN and uses an iterative loop to remove reflections frame by frame. In a preferred embodiment, its specific execution process can be as follows:

[0089] Given a video sequence captured by a stereo camera ,in This represents the index of the video frame. We designed the depth estimation module. This module achieves depth perception and derives a transmission depth map from the mixed depth. It takes the left and right views of a captured video frame containing reflections as input and generates transmission depth maps for those left and right views. This module can be described as follows: The remaining part of the network (denoted as the reflection removal module) The purpose is to perform feature fusion and reflection removal. This involves combining the first N sets of left and right view video frames, the current left and right view video frame, and the subsequent N sets of left and right view video frames (…). ,common (Group) and its corresponding left and right view transmission depth As input, the final transmission map is generated. The process can be described as follows: In the reflection removal module In the feature extraction stage, this invention designs three key modules at different feature scales to interact with and enhance cross-view features, depth-guided features, and cross-frame features, respectively. First, a cross-view feature enhancement module (CVEM) is designed to fuse the left and right view features of the current video frame. Then, a depth-guided feature enhancement module (DGEM) is designed to fuse the features fused by the CVEM module with the depth features, leveraging the guiding role of the transmission map in reflection removal. Finally, a cross-frame feature enhancement module (CFEM) is designed to fuse the depth-guided feature map with… Feature maps from a series of video frame sequences are fused to fully utilize the inter-frame information of the stereo video. Through the aforementioned series of feature enhancement modules, the CVEM module enables the left and right view features to complement each other, while aligning the left and right view features with the depth map at the pixel level; the DGEM module guides depth information and integrates features; and the CFEM module further fuses the rich inter-frame information of the video, which also improves video smoothness. In this invention, the above three feature interaction modules are summarized into a unified feature interaction structure and paired with different gating controllers.

[0090] The transmission map depth estimation module of the reflection removal network provided by this invention It is designed based on the UNet architecture. In one feasible embodiment, the module takes left and right views containing reflected video frames as input and performs depth estimation from coarse to fine. First, the original image is obtained. A coarse-scale depth map. By continuously upsampling and integrating lower-layer features, the original depth map is gradually restored. scale, The module scales the data and ultimately obtains a full-scale transmission depth map. Furthermore, it employs two residual blocks, four dilated convolutional layers, and two convolutional block attention modules (CBAM) to expand the receptive field and enrich the diversity of intermediate convolutional layers.

[0091] The reflection removal module in this embodiment The design is based on an encoder-decoder architecture, consisting of four encoders and one decoder. The first encoder uses the left view of the current frame of the stereo video. As input, output the left-look feature map of the current video frame. The second encoder uses the right view of the current frame of the stereo video. As input, output the right-look feature map of the current video frame. The third encoder takes the left and right views of the current frame of the stereo video as input. Generated transmission depth left and right views As input, output the transmission depth feature map of the current video frame. The fourth encoder will process the first N sets of left and right view video frames, the current left and right view video frames, and the next N sets of left and right view video frames ( (and its corresponding left and right views of the transmission depth) As input, the output is an inter-frame feature map. These four encoders employ the same structure to extract features step-by-step. Each encoder uses three feature extraction blocks, where the b-th feature extraction block consists of b convolutional layers and b CBAMs (b = 1, 2, 3). Between the four encoders, a unified structure of CVEM, DGEM, and CFEM modules with different gating controllers, designed according to this invention, is embedded to perform cross-viewpoint, depth-guided, and cross-frame feature interaction and enhancement, respectively. The decoder fuses and enhances the feature maps. As input, four dilated convolutional blocks, two CBAM blocks, and two sets of deconvolution and convolutional blocks are used to generate the final transmission map video frame. .

[0092] In this embodiment, the CVEM, DGEM, and CFEM modules are cascaded and established among the four encoder branches of the reflection removal module. Aside from the difference in the gating controller, their structures are similar. The m-th CVEM maps the left and right view features at the current level. and As input, we get 3 outputs: Then, the enhanced features generated by CVEM are... with depth features Combining, we get: Then, CFEM will use the depth-enhanced feature map obtained from DGEM. Inter-frame feature maps Combined, we obtain The updated feature maps will be processed by convolutional blocks and CBAM, and then further feature enhancement will be performed in a higher dimension. These three modules are in the... The hierarchical feature map processing (update) process can be described as follows:

[0093]

[0094]

[0095]

[0096] in , , These represent the CVEM module, DGEM module, and CFEM module, respectively.

[0097] The following uses DEGM as an example to illustrate the specific operations within each module. Let's assume the m-th level feature map... and Each position The eigenvectors are denoted as ∈ , ∈ , where c represents the feature dimension. First, the feature map... and Perform global average pooling (GAP) to obtain the global feature vector. and For each position In this embodiment, the feature vector , and Combined, they form a point feature group:

[0098]

[0099] in This involves a feature combination operation. Then, a multi-head attention mechanism is employed. Process each point feature group separately Then they are combined, and linear layers are used to integrate the features. This process can be represented as:

[0100] .

[0101] In each In this embodiment, a similarity matrix is ​​constructed. To establish the relationship between the four components, the four components are the equation. The four elements in the cat(·) function Then, the similarity matrix The input gating controller controls the direction of various feature extraction processes. This embodiment will introduce the specific design of the gating controllers for CVEM, DGEM, and CFEM in the next section. Under the control of the gating controller, the unified feature fusion and enhancement module can effectively learn the correlation between the four components and dynamically adjust the feature vector group through matrix multiplication. Then, the adjusted feature vectors in the vector group are redistributed to their corresponding positions to obtain an updated feature map. To obtain fused features and provide global guidance for subsequent modules, this embodiment combines the two obtained feature maps and uses... Convolution integrates the features to obtain a fused feature map. .

[0102] Similarly, the CVEM processing procedure is as follows:

[0103] Set the left-look feature map of the current video frame as input. and right-view feature map Each position The feature vector is ∈ and ∈ Where c represents the feature dimension;

[0104] Left-view feature map of the current video frame and right-view feature map Perform global flat pooling to obtain global feature vectors. and ;

[0105] Through

[0106]

[0107] For eigenvectors and and global feature vectors and Combine them to obtain point feature groups In the formula, Representative feature combination operations;

[0108] Through

[0109]

[0110] Perform feature integration; where, This indicates that each feature group is processed separately using a multi-head attention mechanism. ;

[0111] The construction is used to establish four components ( Similarity matrix of the relationship between ) The data is then input into the first gating controller to control the feature extraction direction, thereby obtaining an updated feature map. and ;

[0112] Updated feature map and Perform fusion processing to obtain fused feature maps. .

[0113] Similarly, the CFEM processing procedure is as follows:

[0114] Set the fused feature map of the input. and inter-frame feature maps Each position The feature vector is ∈ and ∈ Where c represents the feature dimension;

[0115] Feature maps of fusion and inter-frame feature maps Perform global flat pooling to obtain global feature vectors. and ;

[0116] Through

[0117]

[0118] For eigenvectors and and global feature vectors and Combine them to obtain point feature groups In the formula, Representative feature combination operations;

[0119] Through

[0120]

[0121] Perform feature integration; where, This indicates that each feature group is processed separately using a multi-head attention mechanism. ;

[0122] The construction is used to establish four components ( Similarity matrix of the relationship between ) The data is then input into the third gating controller to control the feature extraction direction, thereby obtaining an updated feature map. and ;

[0123] Updated feature map and Perform fusion processing to obtain fused feature maps. .

[0124] In this embodiment, three unified feature fusion and enhancement modules—CVEM, DGEM, and CFEM—are cascaded to establish relationships among the four encoders. A gating controller is designed to control the feature relationship exploration range for the three different tasks. For the second gating controller of DGEM, the left and right views of the video frames used for feature extraction and their transmission depths are point-to-point aligned, so the second gating controller of this module uses the j-th similarity matrix... and an anti-angle matrix The sum is used as input, and a softmax layer is applied to constrain the inverse angle value to be close to 0. This is used in a gated controller within a DGEM. The interaction between global information of one modality feature map and local information of another modality feature map is suppressed. This gating controller can be described as follows:

[0125] ,in The anti-diagonal matrix has each anti-diagonal element with a value of -100. In contrast, the input paired feature maps of CVEM and CFEM are not aligned. Exploring the local and global correlations in the two feature maps is meaningful. Specifically, gating controllers are applied to the CVEM and CFEM modules. With the j-th similarity matrix Add a matrix The input is then activated through a softmax layer. This process can be represented as follows:

[0126] ; .in for .

[0127] In a preferred embodiment of the present invention, the designed model is further trained by introducing reconstruction loss, pixel loss, structural loss, perceptual loss, and adversarial loss to optimize network performance. The overall loss function used to train the network in this invention is defined as follows:

[0128]

[0129] in, They represent: Pixel loss, SSIM (Structural Similarity) structural loss, reconstruction loss, mean squared error loss, and perceptual loss; .

[0130] Appendix Figure 4 Two visualization examples of the present invention are provided. The first column is a frame of the left view of the binocular video; columns 2-6 are the visualization results of the comparison algorithm; columns 7 and 9 are the visualization results of the transmission map and transmission depth map of the present invention; and columns 8 and 10 are reference images of the actual results corresponding to the examples.

[0131] In summary, this invention provides a depth-information-guided method for reflection removal in stereo video containing reflections. This method includes a transmission map depth estimation module and a reflection removal module. The transmission map depth estimation module performs depth perception frame-by-frame on the stereo video stream, decoupling the mixed depth corresponding to the stereo video stream to obtain the true depth corresponding to the transmission layer and the pseudo depth corresponding to the reflection layer. The true depth corresponding to the transmission map is used to guide the reflection removal network, achieving effective reflection removal. The reflection removal module, by incorporating a unified structured feature fusion and enhancement module, achieves guidance on cross-viewpoint features, depth features, and cross-frame feature fusion and enhancement. Gating controllers are designed for the CVEM, DGEM, and CFEM modules to control the feature relationship exploration range of the three different tasks. Ultimately, this achieves good reflection removal results for a given video stream containing reflections. Compared with existing single-map and multi-map-based reflection removal algorithms, the proposed algorithm exhibits better PSNR, SSIM, NCC, and LMSE metrics, and also outperforms the compared algorithms in terms of post-reflection removal visualization.

[0132] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of one embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing the present invention.

[0133] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.

[0134] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for apparatus or system embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. The apparatus and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0135] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A depth-information-guided binocular video reflection removal method, characterized in that, include: Based on the left and right views of multiple reflected video frames, the transmission depth map of the multiple reflected video frames is obtained by processing through the transmission map depth estimation module of the network model; the multiple video frames include the current video frame, the previous video frame, and the subsequent video frame; the transmission depth map includes a left view of transmission depth, a right view of transmission depth, and left and right views of transmission depth. Based on the left and right views of the multiple reflection video frames and the corresponding transmission depth maps, feature fusion and reflection removal processing are performed through the reflection removal module of the network model to obtain transmission map video frames. The reflection removal module is based on an encoder-decoder structure, including a first encoder, a second encoder, a third encoder, and a fourth encoder arranged in parallel with each other, as well as a decoder. It also has a cross-view feature enhancement module, a depth-guided feature enhancement module, and a cross-frame feature enhancement module. The first encoder and the second encoder are used to perform stepwise feature extraction on the left and right views of the current video frame, respectively; the third encoder is used to perform stepwise feature extraction on the left and right views of the transmission depth of the current video frame; and the fourth encoder is used to perform stepwise feature extraction on the left and right views and the corresponding transmission depth maps of the multiple reflection video frames. The cross-view feature enhancement module performs a left-right view feature complementation operation on the outputs of the first encoder and the second encoder to obtain enhanced left-right feature maps of the reflected video frame, and aligns the features of the enhanced left-right feature maps of the current video frame with the transmission depth map at the pixel level; the depth-guided feature enhancement module performs depth information guidance and feature integration operations on the outputs of the cross-view feature enhancement module and the third encoder to obtain a feature map enhanced by depth guidance; the inter-frame feature enhancement module performs a fusion operation on the outputs of the depth-guided feature enhancement module and the fourth encoder to obtain a feature map enhanced by inter-frame features.

2. The method according to claim 1, characterized in that, The process of obtaining a transmission depth map of multiple reflected video frames by processing the left and right views based on multiple reflected video frames through a transmission map depth estimation module includes: Video sequences captured by a binocular camera with multiple reflected video frames. The calculation formula of the transmission map depth estimation module Obtain the transmission depth map of the reflected video frame. In the formula, where Indicates the index of the video frame.

3. The method according to claim 2, characterized in that: The first encoder uses the left view of the current video frame. As input, output the left-look feature map of the current video frame. The second encoder uses the right view of the current video frame. As input, output the right-look feature map of the current video frame. The third encoder uses the transmission depth left and right views corresponding to the left and right views of the current video frame. As input, output the transmission depth feature map of the current video frame. The fourth encoder will process the preceding video frame, the current video frame, and the subsequent video frames. and the corresponding left and right view transmission depth As input, the output is an inter-frame feature map. ; The first encoder, the second encoder, the third encoder and the fourth encoder each have three feature extraction blocks; the b-th feature extraction block of the three feature extraction blocks is composed of b convolutional layers and b CBAMs, where b = 1, 2, 3; The processing procedure of the cross-view feature enhancement module specifically includes: Set the left-look feature map of the current video frame as input. and right-view feature map Each position The feature vector is ∈ and ∈ Where c represents the feature dimension; Left-view feature map of the current video frame and right-view feature map Perform global flat pooling to obtain global feature vectors. and ; Through For eigenvectors and and global feature vectors and Combine them to obtain point feature groups In the formula, Representative feature combination operations; Through Perform feature integration; where, This indicates that each feature group is processed separately using a multi-head attention mechanism. ; Construct is used to build Similarity matrix of the relationships between the four components The data is then input into the first gating controller to control the feature extraction direction, thereby obtaining an updated feature map. and ; Updated feature map and Perform fusion processing to obtain fused feature maps. ; The processing procedure of the deep-guided feature enhancement module specifically includes: Set the fused feature map of the input. and the left and right feature maps of the current video frame Each position The feature vector is ∈ and ∈ Where c represents the feature dimension; Feature maps of fusion and left and right feature maps Perform global flat pooling to obtain global feature vectors. and ; Through For eigenvectors and and global feature vectors and Combine them to obtain point feature groups In the formula, Representative feature combination operations; Through Perform feature integration; where, This indicates that each feature group is processed separately using a multi-head attention mechanism. ; Construct is used to build Similarity matrix of the relationships between the four components The data is then input into the second gating controller to control the feature extraction direction, thereby obtaining an updated feature map. and ; Updated feature map and Perform fusion processing to obtain fused feature maps. ; The processing procedure of the cross-frame feature enhancement module specifically includes: Set the fused feature map of the input. and inter-frame feature maps Each position The feature vector is ∈ and ∈ Where c represents the feature dimension; Feature maps of fusion and inter-frame feature maps Perform global flat pooling to obtain global feature vectors. and ; Through For eigenvectors and and global feature vectors and Combine them to obtain point feature groups In the formula, Representative feature combination operations; Through Perform feature integration; where, This indicates that each feature group is processed separately using a multi-head attention mechanism. ; Construct is used to build Similarity matrix of the relationships between the four components The data is then input into the third gating controller to control the feature extraction direction, thereby obtaining an updated feature map. and ; Updated feature map and Perform fusion processing to obtain fused feature maps. .

4. The method according to claim 3, characterized in that, The processing procedure of the first gate controller includes: Through Controlling the direction of feature extraction to obtain updated feature maps and In the formula, Indicates the first gate controller. for Matrix; The processing procedure of the second gate controller includes: Through Controlling the direction of feature extraction to obtain updated feature maps and In the formula, Indicates the second gate controller. It is an anti-diagonal matrix, and each anti-diagonal element of the anti-diagonal matrix has a value of -100; The processing procedure of the third gate controller includes: Through Controlling the direction of feature extraction to obtain updated feature maps and ; Indicates the third gate controller. for The matrix.

5. The method according to claim 4, characterized in that, The network model also has a loss function, which is expressed as follows: Calculate the overall loss to train the network model; where, They represent: Pixel loss, SSIM structure loss, reconstruction loss, mean square error loss, and perceptual loss; .