Stereoscopic video compression method with double-branch attention
By introducing a codec block with a dual-branch attention mechanism and a high-frequency information fusion module in stereo video compression, the shortcomings of the existing technology in capturing non-repetitive textures and global information are solved, and higher quality image reconstruction and better compression performance are achieved.
Patent Information
- Application Number
- CN202510229638.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-28
Smart Images

Figure CN120050431A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of computer vision, deep learning, video compression, etc., and particularly relates to a stereoscopic video compression method with dual-branch attention. Background Art
[0002] Stereoscopic dual-view video is a video technology that simulates the parallax of human eyes by providing videos of two different perspectives (left perspective and right perspective), thereby generating a three-dimensional stereoscopic effect in the perception of viewers. With the popularization of autonomous vehicles (AVs) equipped with stereoscopic cameras and the development and application of virtual reality devices (VRs), the importance of stereoscopic video data (dual-view) has been significantly improved. During the application of AVs and VRs, the codec not only needs to maintain low latency but also needs to pay more attention to the quality of the reconstructed image after decoding. Therefore, how to efficiently compress stereoscopic video data has become a research hotspot.
[0003] Traditional stereoscopic video compression methods put each view into a single-view codec (such as H.264 / AVC or HEVC) separately, but they ignore the view redundancy caused by the similarity between the two views after compression. To solve this defect, people extended the single-view codec. For example, the extended H.264 / MVC of H.264 / AVC uses the method of disparity compensation to eliminate the disparity redundancy after compression, or the extended MV-HEVC of HEVC uses technologies such as coding tree units (CTUs) to eliminate the disparity redundancy. They all improve the compression performance by combining the prediction methods between frames and between views.
[0004] In recent years, due to the powerful learning ability of the Deep Neural Network (DNN), the stereo video compression and coding method based on DNN has been greatly developed and become a research hotspot (list all relevant literature). Luo et al. first proposed a Learning-based Stereo Video Compression (LSVC) framework based on DNN. Luo et al. reduced the temporal and binocular redundancies by performing motion compression and disparity compression in the feature space, thereby achieving efficient compression and coding of stereo videos. Hou et al. proposed a Low-Latency Neural Stereo Streaming (LLSS) method based on DNN. By introducing a bidirectional feature displacement module and sharing motion information and disparity information, parallel processing of the left and right views was achieved, significantly reducing the latency and improving the Rate Distortion (RD) performance. Although compared with traditional stereo video (two-view) compression and coding methods, the existing DNN-based stereo video compression methods can obtain better RD performance, almost all of them only use convolutional operations for feature extraction and fusion. Although convolutional operations can effectively extract local information, they have poor ability to capture non-repetitive and complex texture information, and convolutional operations cannot capture long-distance global information, resulting in defects such as the inability to effectively capture non-repetitive texture details within a local range and ignoring global features, seriously affecting the quality of image reconstruction during the decoding process. Summary of the Invention
[0005] To solve the problems existing in the background technology, the present invention provides a stereo video compression method with dual-branch attention, including:
[0006] S1: Divide the stereo video into a left video frame sequence and a right video frame sequence;
[0007] S2: Input the video frame at time t of the stereo video, its temporally adjacent reconstructed frame, and the view-adjacent reconstructed frame into the DAN stereo video compressor for compression to obtain the video reconstructed frame at time t; wherein, the video frame at time t of the stereo video includes: the left video frame at time t and the right video frame at time t; the temporally adjacent reconstructed frame of the left video frame at time t is the left video reconstructed frame at time t - 1, and its view-adjacent reconstructed frame is the right video reconstructed frame at time t - 1; the temporally adjacent reconstructed frame of the right video frame at time t is the right video reconstructed frame at time t - 1, and the view-adjacent reconstructed frame is the left video reconstructed frame at time t - 1;
[0008] The DAN stereoscopic video compressor includes: a feature extraction module, a motion estimation module, a disparity estimation module, an encoding and decoding module based on LGEDB, a motion compensation module, a disparity compensation module, a dual-branch high-frequency information fusion module, and an image reconstruction module;
[0009] S3: Store the left video reconstruction frame at time t and the right video reconstruction frame at time t to obtain the compressed stereoscopic video.
[0010] The present invention has at least the following beneficial effects
[0011] The LGEDB and DHFFM proposed by the present invention can achieve higher-quality image reconstruction with the same or lower Bits Per Pixel (BPP). A local and global dual-branch encoding and decoding block (LGEDB) based on Transformer and channel attention mechanism proposed by the present invention can accurately capture local non-repetitive texture details and global structure information by fusing the self-attention of pixel points in the local range and global channel attention. At the same time, the present invention also proposes a dual-branch high-frequency information fusion module (DHFFM) based on reversible neural network and gating mechanism, which can effectively extract the high-frequency information in the motion compensation feature and the disparity compensation feature, and realize the efficient fusion of the two types of features through the screening of pixel-by-pixel features. The LGEDB proposed by the present invention can achieve higher-quality image reconstruction while ensuring a lower bit rate. Description of the Drawings
[0012] Figure 1 It is a schematic diagram of the method flow of the present invention;
[0013] Figure 2 It is a schematic diagram of the compression process of each frame of the stereoscopic dual-view video in the embodiment of the present invention;
[0014] Figure 3 It is a schematic diagram of the overall framework of DAN in the embodiment of the present invention;
[0015] Figure 4 It is a schematic diagram of the LGEDB structure in the present invention;
[0016] Figure 5 It is a schematic diagram of the LWAB structure in the present invention;
[0017] Figure 6 It is a schematic diagram of the GCAB structure in the present invention;
[0018] Figure 7 It is a schematic diagram of the structure of the residual module in the present invention;
[0019] Figure 8 It is a schematic diagram of the structure of the dual-branch high-frequency information fusion module in the present invention;
[0020] Figure 9 Schematic diagram of the feature extraction module and the image reconstruction module in the present invention;
[0021] Figure 10 Schematic diagram of the motion estimation module and the disparity estimation module in the present invention;
[0022] Figure 11 Schematic diagram of the motion compensation module and the disparity compensation module in the present invention;
[0023] Figure 12 Schematic diagram of the NAF module in the present invention;
[0024] Figure 13 Schematic diagram of the INN module in the present invention;
[0025] Figure 14 Schematic diagram of the simulation comparison between the present invention and the prior art on different data sets. Detailed implementation manners
[0026] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention schematically. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0027] Please refer to Figure 1 , the present invention provides a stereoscopic video compression method with dual-branch attention, including:
[0028] S1: Divide the stereoscopic video into a left video frame sequence and a right video frame sequence;
[0029] S2: Input the video frame of the stereoscopic video at time t, its temporally adjacent reconstructed frame, and the view-adjacent reconstructed frame into the DAN stereoscopic video compressor for compression to obtain the video reconstructed frame at time t; wherein, the video frame of the stereoscopic video at time t includes: the left video frame at time t and the right video frame at time t; the temporally adjacent reconstructed frame of the left video frame at time t is the left video reconstructed frame at time t-1, and its view-adjacent reconstructed frame is the right video reconstructed frame at time t-1; the temporally adjacent reconstructed frame of the right video frame at time t is the right video reconstructed frame at time t-1, and the view-adjacent reconstructed frame is the left video reconstructed frame at time t-1;
[0030] The DAN stereoscopic video compressor includes: a feature extraction module, a motion estimation module, a disparity estimation module, an encoding and decoding module based on LGEDB, a motion compensation module, a disparity compensation module, a dual-branch high-frequency information fusion module, and an image reconstruction module;
[0031] S3: Store the left video reconstruction frame at time t and the right video reconstruction frame at time t to obtain the compressed stereoscopic video.
[0032] Preferably, for the first video frame in the left video frame sequence and the right video frame sequence, an image compression method is used to obtain the corresponding left video reconstruction frame and right video reconstruction frame.
[0033] In this embodiment, the first video frame can adopt the intra-frame prediction method of H.265. Since there is no previous frame for the first frame to be used as a reference, it can only rely on the pixel information inside the image for compression. In H.265, the image is divided into multiple small blocks. Each small block can use multiple prediction modes for compression to improve the compression ratio and reconstruction quality
[0034] In this embodiment, the compression of the first video frame can also use the JPEG image compression algorithm. The JPEG image compression algorithm converts the image from the time domain to the frequency domain, uses the discrete cosine transform (DCT) to remove redundant information, and achieves efficient lossy compression through quantization and entropy coding. The obtained left video reconstruction frame and right video reconstruction frame are used as reference frames for subsequent frames and input into the DAN network.
[0035] In this embodiment, in addition to the above two image compression algorithms that can compress the first video frame, other existing methods can also be used to compress the first video frame, which will not be elaborated too much in this embodiment.
[0036] Preferably, the reconstruction process of the DAN stereoscopic video compressor includes:
[0037] S21: Input the left video frame at time t the right video frame at time t the right video reconstruction frame at time t - 1 and the left video reconstruction frame at time t - 1 into the feature extraction module for feature extraction to obtain the corresponding feature F t L 、feature F t R 、feature and feature
[0038] S22: Input feature F t R and feature into the motion estimation module to calculate the motion information and input the feature F t L and the feature into the motion estimation module to calculate the motion information
[0039] S23: Input the motion information and into the encoding and decoding module based on LGEDB respectively to extract the local non-repetitive complex region texture features and global information, and obtain the motion reconstruction information and
[0040] S24: According to the motion reconstruction information use the motion compensation module to perform motion compensation on the feature F t R to obtain the feature F MR , and according to the motion reconstruction information use the motion compensation module to perform motion compensation on the feature F t L to obtain the feature F ML ;
[0041] S25: According to the feature F t R and the feature use the disparity estimation module to calculate the disparity information and according to the feature F t L and the feature use the disparity estimation module to calculate the disparity information
[0042] S26: Input the disparity information and into the encoding and decoding module based on LGEDB respectively to extract the local non-repetitive complex region texture features and global information, and obtain the disparity reconstruction information and
[0043] S27: According to the disparity reconstruction information use the disparity compensation module to perform disparity compensation on the feature F t R to obtain the feature F DR , according to the disparity reconstruction information use the disparity compensation module to perform disparity compensation on the feature F t L to obtain the feature F DL ;
[0044] S28: The feature F MR and the feature F DRThe input double-branch high-frequency information fusion module performs feature fusion to obtain prediction features The feature F ML and the feature F DL are input into the double-branch high-frequency information fusion module for feature fusion to obtain prediction features
[0045] S29: Subtract the feature t R from the feature F to obtain residual information Subtract the featurefrom the feature F t L to obtain residual information Input the residual information and the residual information into the LGEDB-based encoding and decoding module respectively to extract local non-repetitive complex region texture features and global information to obtain residual reconstruction information and Input the residual reconstruction information and the prediction features to perform feature addition to obtain reconstruction features Input the residual reconstruction information and the prediction features to perform feature addition to obtain reconstruction features Input the reconstruction features into the image reconstruction module to obtain the right video reconstruction frame at time t Input the reconstruction features into the image reconstruction module to obtain the left video reconstruction frame at time t
[0046] In this embodiment, the stereo binocular video is simultaneously captured by the left (L) and right (R) cameras for T frames, that is where the superscript represents the view and the subscript t represents the time step. The present invention uses the same process as LSVC to compress each frame, as Figure 2 shown, and will perform compression in turn in the order of first compressing the right view and then the left view
[0047] The left and right views of the present invention adopt the same compression coding network, as Figure 3 shown. We take the right view as an example to elaborate on the DAN proposed by the present invention in detail. First, input the current encoded frame the temporally adjacent reconstruction frame and the view-adjacent reconstruction frame into the feature extraction module to obtain the corresponding feature maps F t R and and Then, the feature F of the current encoded framet R And the features of the reconstructed frame adjacent to time Are fed into the motion estimation module to obtain motion information Then the motion information Is fed into the encoding and decoding module composed of local + global double-branch encoding and decoding blocks (LGEDB) to obtain motion reconstruction information Then for the motion reconstruction information And F t R Motion compensation is performed to obtain the motion compensation feature F ML . Similarly, the feature F of the current encoded frame t R And the features of the reconstructed frame adjacent to the view Are fed into the disparity estimation module to obtain disparity information The disparity information Is fed into the encoding and decoding module based on LGEDB to obtain disparity reconstruction information Then for the disparity reconstruction information And F t R Disparity compensation is performed to obtain the disparity compensation feature F DL . Next, F ML And F DL Are fed into the fusion module for feature fusion to obtain the predicted feature And And F t R Are subtracted to obtain the residual information That is Then the residual information Is fed into the encoding and decoding module based on LGEDB to obtain the residual reconstruction information Finally, the residual reconstruction information Is added to the predicted feature To obtain the reconstructed feature That is And the reconstructed feature is used for image reconstruction to obtain the reconstructed image And the reconstructed image is placed in the buffer. Among them, the feature extraction module, motion estimation module, motion compensation module, disparity estimation module, disparity compensation module, and image reconstruction module all adopt the same settings as LSVC, and the specific structure can be referred to Figures 9 to 11 As shown, and the encoding and decoding module based on LGEDB is the main contribution and innovation of the present invention
[0048] Preferably, the encoding and decoding module based on LGEDB includes: an encoder, a bottleneck layer, and a decoder; the encoder includes N convolutional modules and LGEDB modules connected in series between any two adjacent convolutional modules; the decoder includes N deconvolutional modules and LGEDB modules connected in series between any two adjacent deconvolutional modules;
[0049] The encoder is used to encode the input features of the encoding and decoding module based on LGEDB, the bottleneck layer is used to map the encoded features of the encoder, and the decoder is used to decode the mapped features of the bottleneck layer to obtain the output features of the encoding and decoding module based on LGEDB.
[0050] Preferably, the LGEDB module includes: an LWAB module, a GCAB module, a convolutional layer, and M cascaded residual modules; the feature extraction process of the LGEDB module includes:
[0051] S211: Input the input feature F of the LGEDB module into the LWAB module and the GCAB module respectively for feature extraction to obtain the feature f w and the feature f c ;
[0052] S212: Concatenate the feature f w and the feature f c , and then input them into the convolutional layer and M cascaded residual modules in sequence for feature processing to obtain the feature f;
[0053] S213: Add the feature f and the input feature F to obtain the output feature F of the LGEDB module LG-out .
[0054] Preferably, the feature extraction process of the LWAB module is as follows:
[0055] Given the input feature F ∈ R H×W×C , H represents the height of the input feature F, W represents the width of the input feature F, C represents the number of channels of the input feature F. First, divide the input feature into non-overlapping blocks of size p×p, that is k = 1, 2,..., L, represents the number of divided blocks; divide X k equally into multiple heads along the channel direction through a linear layer to generate queries keys and value V s k matrices, is the number of channels assigned to each head, s = 1, 2,..., S, S is the number of heads, and the generation formula is:
[0056]
[0057] Among them, respectively represent the linear projection matrix of the s-th head in the k-th block; calculate the feature similarity of the s-th head in the k-th block
[0058]
[0059] Among them, E is the relative position encoding; because there are S heads, S self-attention calculations are performed in parallel, and then the attention features calculated by each head are concatenated along the channel dimension to obtain feature A k ∈R p×p×C , and then feature A k is added to X k to obtain the self-attention feature of the k-th block
[0060]
[0061] Finally, the self-attention features of each block are concatenated back to their original positions to obtain the output feature f of LWAB w ∈R H ×W×C :
[0062]
[0063] Among them, concat(·) represents position concatenation.
[0064] Preferably, the feature extraction process of the GCAB module is as follows:
[0065] Given the input feature F ∈ R H×W×C , H represents the height of the input feature F, W represents the width of the input feature F, C represents the number of channels of the input feature F, and its calculation process is:
[0066] f' = Conv3(ReLU(Conv3(F)))
[0067] f” = ReLU(Conv1(AvgPool(f')))
[0068] f c = Sigmoid(Conv1(f”)) + f'
[0069] Among them, Conv3(·) represents the standard 3×3 convolution operation, Conv1(·) represents the standard 1×1 convolution operation, ReLU(·) represents the ReLU activation function, Sigmoid(·) represents the Sigmoid activation function; AvgPool(.) represents the average pooling operation.
[0070] In this embodiment, most existing stereo dual-view video compression and encoding models only use pure convolution to construct the codec. However, the operation of pure convolution not only fails to effectively extract features from complex regions with non-repetitive textures but also cannot capture long-range global information, resulting in poor compression performance. The LGEDB proposed by the present invention can balance the effective extraction of textures in non-repetitive complex regions and global information.
[0071] As Figure 5 shown, the LGEDB proposed by the present invention is composed of a local window attention branch (LWAB) and a global channel attention branch (GCAB), and effectively extracts local non-repetitive complex region texture features and global information through LWAB and GCAB respectively.
[0072] For LWAB, the present invention uses window-based multi-head attention to obtain local features, as Figure 3 shown. Given the input feature F ∈ R H×W×C , first divide the input feature into non-overlapping windows of size p×p, that is k = 1, 2,..., L( that is, the number of divided blocks). Then, through a linear layer, X k is equally divided into multiple heads in the channel direction to generate query, key, and value matrices: is the number of channels assigned to each head, s = 1, 2,..., S, S is the number of heads, and the generation formula is:
[0073]
[0074] Among them, respectively represent the linear projection matrices of the s-th head in the k-th block. At the same time, the parameters between different blocks are shared. Next, calculate the feature similarity
[0075]
[0076] where E is the relative position encoding. Since there are S heads, the self-attention calculation of formula (4) is performed S times in parallel, and then the attention features calculated by each head are concatenated in the channel dimension to obtain the feature A k ∈R p×p×C , and then the feature A k is added to X k to obtain the self-attention feature of the k-th block
[0077]
[0078] Finally, the self-attention features of each block are concatenated back to their original positions to obtain the LWAB output feature f of the entire image. w ∈R H×W×C :
[0079]
[0080] Among them, concat(·) represents position concatenation.
[0081] From formulas (1)-(5) and Figure 5 it can be seen that LWAB constructs the features of the local window by calculating the correlation between each pixel point in each window and other pixel points in the same window. Therefore, if there are repeated textures in the local window, their similarity is high, and the response after LWAB will be large; on the contrary, if there are non-repeated textures in the local window, the response after LWAB will be small. Therefore, LWAB can accurately express the non-repeated texture features in the local area.
[0082] For GCAB, the present invention adopts the idea of channel attention to effectively extract global information. As Figure 6 shown, given the input feature F ∈ R H×W×C , its calculation process is as follows:
[0083] f' = Conv3(ReLU(Conv3(F))) (7)
[0084] f” = ReLU(Conv1(AvgPool(f'))) (8)
[0085] f c = Sigmoid(Conv1(f”)) + f' (9)
[0086] Among them, Conv3(·) represents the standard 3×3 convolution operation, Conv1(·) represents the standard 1×1 convolution operation, ReLU(·) represents the ReLU activation function, and Sigmoid(·) represents the Sigmoid activation function.
[0087] After passing through LWAB and GCAB respectively, the present invention uses concatenation, convolution, and residual blocks to fuse the extracted local features and global information:
[0088] f = RB(RB(RB(Conv3(cat(f c , f w )))) (10)
[0089] Among them, cat(·) represents channel concatenation, and RB(·) represents a residual block, as Figure 7 shown. The final output of the LGEDB is:
[0090] F LG-out = f + F
[0091] Please refer to Figure 7 , in this embodiment, the residual module RB(·) consists of a cascaded first 3×3 convolution, a Relu activation function, and a second 3×3 convolution. The input feature of the residual module RB(·) and the output feature of the second 3×3 convolution are added to obtain the output feature of the residual module RB(·).
[0092] Preferably, the mapping of the encoded features of the encoder by the bottleneck layer includes:
[0093] (r, (μ, σ)) = CWAEM(y)
[0094]
[0095] Among them, y represents the encoded feature after encoding by the encoder, μ, σ, and r represent the mean, variance, and latent space residual respectively; Q(·) is the quantization process; CWAEM(·) represents the channel-wise autoregressive entropy model; AE(·) represents the arithmetic encoding process, and AD(·) represents the arithmetic decoding process.
[0096] After effectively extracting the texture features and global information of non-repetitive complex regions, the present invention uses the proposed LGEDB as a basic block and combines it with the channel-wise autoregressive entropy model (CWAEM) to construct an encoding and decoding module, as Figure 4 shown.
[0097] The constructed encoding and decoding module will act on the motion information disparity information and residual information respectively to generate corresponding reconstruction information. Taking the motion information as an example, the input is first used by the encoder composed of LGEDB to extract the latent representation y. Then y passes through quantization, CWAEM, arithmetic encoding (AE), arithmetic decoding (AD), and the decoder composed of LGEDB respectively to obtain the motion reconstruction information This process can be expressed as:
[0098]
[0099] (r, (μ, σ)) = CWAEM(y) (12)
[0100]
[0101]
[0102] Wherein, E(·) and D(·) respectively represent the encoder and decoder composed of LGEDB. Q(·) is the quantization process, CWAEM(·) represents the channel autoregressive entropy model, μ, σ and r respectively represent the mean, variance and latent space residual; AE(·) represents the arithmetic coding process, and AD(·) represents the arithmetic decoding process.
[0103]
[0104] Wherein, Q step is the quantization step size, and round represents the rounding function.
[0105] In the motion information and parallax information respectively pass through the encoding and decoding modules based on LGEDB, the obtained motion reconstruction information and parallax reconstruction information then pass through the motion compensation module and the parallax compensation module to obtain the motion compensation feature F M and the parallax compensation feature F D , and then use DHFFM to fuse F M and F D .
[0106] Preferably, the feature fusion process of the double-branch high-frequency information fusion module includes:
[0107]
[0108] θ = INN(INN(F M ) + INN(F D ))
[0109]
[0110] Wherein, F M ∈{F MR , F ML}, F D ∈{F DR , F DL}, NAF(·) represents the NAF module, INN(·) represents the invertible neural network; RB() represents the residual module.
[0111] In addition to effectively extracting local non-repetitive texture information and global information in the codec, the effective fusion of motion compensation information and disparity compensation information is another key factor affecting the quality of stereoscopic dual-view video compression coding. Based on this, the present invention proposes a novel dual-branch high-frequency information fusion module (DHFFM) to achieve effective extraction of high-frequency information and accurate screening of features in motion compensation features and disparity compensation features, thereby realizing the efficient fusion of motion compensation features and disparity compensation features.
[0112] Gomez et al. have demonstrated that the inverse transformation property of the invertible neural network (INN) can ensure that high-frequency information is not lost during forward transmission and inverse restoration. The INN block proposed by Zhou et al. can effectively extract high-frequency information. In addition, the gating mechanism in the NAF (Nonlinear Activation Free) block proposed by Chen et al. can achieve accurate feature screening. Therefore, as Figure 8 shown, the DHFFM proposed by the present invention first uses INN and NAF to respectively extract high-frequency information and screen features from the motion compensation information F M and the disparity compensation information F D . Then, the high-frequency information and screened features extracted from F M are element-wise added to the high-frequency information and screened features extracted from F D respectively; finally, the added high-frequency information and the added screened features are concatenated, convolved, and passed through a residual block to achieve the efficient fusion of motion compensation features and disparity compensation features, obtaining the predicted feature This process can be expressed as:
[0113]
[0114] θ = INN(INN(F M ) + INN(F D ))
[0115]
[0116] where F M ∈{F MR , F ML}, F D ∈{F DR , F DL}, NAF(·) represents the NAF module, INN(·) represents the invertible neural network; RB() represents the residual module.
[0117] The NAF block and the INN block are existing network structures, and their network structures can be referred to Figure 12 and Figure 13As shown, no specific description is made in this embodiment.
[0118] From the above analysis, it can be seen that due to the introduction of INN and NAF, the DHFFM proposed by the present invention can effectively extract the high-frequency information in the motion compensation feature and the disparity compensation feature, as well as the feature screening for each pixel point, so as to realize the efficient fusion of the motion compensation feature and the disparity compensation feature, and ensure the further improvement of the dual-view stereoscopic video compression coding performance.
[0119] After obtaining the predicted feature through DHFFM then is subtracted from the original feature F t R to obtain the residual information that is Then the residual information is sent to the codec module composed of LGEDB to obtain the residual reconstruction information Subsequently, the residual reconstruction information is added to the predicted feature to obtain the reconstructed feature that is and the reconstructed image is obtained through the image reconstruction module This reconstructed image will be stored in the buffer for the compression coding of the next frame.
[0120] In Figure 8 the residual module RB(·) consists of a cascaded first 3×3 convolution, a Relu activation function, and a second 3×3 convolution. The input feature of the residual module RB(·) and the output feature of the second 3×3 convolution are added to obtain the output feature of the residual module RB(·).
[0121] Preferably, the loss function used in the training of the DAN stereoscopic video compressor includes:
[0122]
[0123] where L represents the loss function, D(·) is the distortion loss, represents the video frame of view v at time t; represents the video reconstructed frame of view v at time t, and R(y) represents the bit rate of the encoded feature of the encoder during the compression process.
[0124] Following the idea of the end-to-end video compression method, all networks of the present invention use the same loss, and all parameters are updated simultaneously with this loss. The loss is set as:
[0125]
[0126] Among them, D(·) is the distortion loss, representing the distortion between the reconstructed frame and the original frame. In the present invention, the Mean Square Errors (MSE) loss and the Multi Scale Structural Similarity Index Measure (MS-SSIM) loss are used to calculate the distortion loss. R(·) represents the bitrate during the compression process. The superscript v indicates using the left view or the right view, and the subscript t represents the time step, where t ∈ {1, 2, 3, 4}. L represents the loss function, and D(·) is the distortion loss. represents the video frame of view v at time t; represents the video reconstructed frame of view v at time t, and R(y) represents the bitrate of the encoding features of the encoder during the compression process. λ is the rate-distortion control hyperparameter. In the present invention, four different λ values (λ = 512, 1024, 2048, 4096) are used.
[0127] Experimental settings
[0128] In the experiments of the present invention, first, the single-view video dataset Vimeo-90K is used for pre-training, and then the stereo dual-view video dataset Cityscape is used for training. The present invention uses the test datasets of the stereo dual-view videos Cityscape, KITTI2012, and KITTI 2015 for evaluation.
[0129] The training set and the test set of the stereo video dataset Cityscape respectively contain 2975 and 1525 stereo video sequence pairs. Each video sequence pair contains 30 frames, and the resolution of each picture is 2048×1024. KITTI 2012 and KITTI2015 respectively contain 195 and 200 stereo video sequence pairs. Each stereo video sequence pair contains 21 frames, and the resolution of each picture is 1241×376. The present invention follows the picture preprocessing process of LSVC and preprocesses the Cityscape test set and KITTI respectively. After preprocessing, the resolutions of all frames correspond to 1920×768 and 1152×256 respectively.
[0130] The present invention follows the training strategy of LSVC and trains the model in 3 stages. First, in the first stage, the feature extraction module, the motion estimation module, the encoding and decoding module based on LGEDB, the motion compensation module, and the image reconstruction module are pre-trained using the Vimeo-90K dataset. With a learning rate of 5×10 -5 for two million iterations. Then, initialize DAN using the pre-training weights of the first stage, and use the Cityscape training set and 1×10 -5The learning rate is used to perform 400,000 iterations on the LGEDB-based encoding and decoding module and the DHFFM, and the parameters of other modules are frozen. Finally, we use the Cityscape training set and 1×10 -5 The learning rate first performs 300,000 iterations on the entire DAN, and then adjusts the learning rate to 5×10 -6 for 200,000 iterations.
[0131] In addition, the present invention uses random rotations of 90 degrees, 180 degrees, 270 degrees, and horizontal flips to perform data augmentation on the training data set. The Adam optimizer is used. In the pre-training stage, the batch size is 8, and the size of the input Vimeo-90K is set to 256×256. In the subsequent stage, the batch size is 4, and the size of the input Cityscape training set is 512×256. During the testing process, we use the image compression method proposed by Cheng et al. to compress the first frame, and fine-tune it on the Cityscape training set according to the training method in the literature. All experiments of the present invention are trained and tested on the NVIDIA 4090 GPU and the Pytorch deep learning framework.
[0132] To prove the superiority of the method proposed by the present invention, we compare the network proposed by the present invention with the existing highly representative video compression methods. The comparison literature is shown in Table 1:
[0133]
[0134]
[0135] To verify the superiority of the method proposed by the present invention, we compare the proposed DAN with the six most advanced methods in recent years: LSVC, LLSS, FVC, DCVC, H.265, and MV-HEVC. Among them, LSVC, LLSS, FVC, and DCVC are DNN-based methods, while H.265 and MV-HEVC are traditional methods. For H.265, HM-16.20 is selected, and the configuration "lowdelay_P_main" is selected; for MV-HEVC, HTM-16.3 is selected, and the configuration is "baseCfg_2_view". For DNN-based methods, the settings in their respective literatures are used respectively. Figure 14 The rate-distortion curves of LSVC, LLSS, FVC, DCVC, H.265, MV-HEVC, and the DAN proposed by the present invention on the Cityscape, KITTI 2012, and KITTI 2015 test sets are shown. As Figure 14 shown, the rate-distortion curve of the DAN method proposed by the present invention is at the top, significantly superior to other methods.
[0136] Table 2 presents the BD-BR (with the baseline of BD-BR selected as MV-HEVC) of LSVC, LLSS, FVC, DCVC, H.265, and DAN proposed by the present invention on the Cityscape, KITTI 2012, and KITTI 2015 test sets respectively.
[0137] BD-BR results of all methods on the datasets Cityscape, KITTI 2012, and KITTI 2015 (with MV-HEVC as the baseline)
[0138]
[0139] As can be seen from Table 2, DAN proposed by the present invention outperforms the highly representative advanced methods in recent years on all test sets: compared with the second-best method LLSS, DAN proposed by the present invention has a BD-DR improvement of 3.6%, 7.3%, and 7.5% on the Cityscape, KITTI 2012, and KITTI 2015 test sets respectively.
[0140] In summary, the LGEDB and DHFFM proposed by the present invention can achieve higher-quality image reconstruction with the same or lower Bits Per Pixel (BPP). The locally and globally dual-branch encoder-decoder block (LGEDB) based on Transformer and channel attention mechanism proposed by the present invention realizes the precise capture of local non-repetitive texture details and global structural information by fusing the self-attention of pixel points in the local range and the global channel attention. At the same time, the present invention also proposes a dual-branch high-frequency information fusion module (DHFFM) based on reversible neural network and gating mechanism, which can effectively extract the high-frequency information in the motion compensation feature and the disparity compensation feature, and realize the efficient fusion of the two types of features through the screening of pixel-by-pixel features. The LGEDB proposed by the present invention can achieve higher-quality image reconstruction while ensuring a lower bit rate.
[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A dual-branch attention stereo video compression method, characterized in that: include: S1: Divide the stereoscopic video into a left video frame sequence and a right video frame sequence; S2: Inputting the video frame of the stereoscopic video at time t, its time-adjacent reconstructed frame and its view-adjacent reconstructed frame into the DAN stereoscopic video compressor for compression to obtain the video reconstructed frame at time t; wherein the video frame of the stereoscopic video at time t includes: a left video frame at time t and a right video frame at time t; the time-adjacent reconstructed frame of the left video frame at time t is the left video reconstructed frame at time t-1, and its view-adjacent reconstructed frame is the right video reconstructed frame at time t-1; the time-adjacent reconstructed frame of the right video frame at time t is the right video reconstructed frame at time t-1, and the view-adjacent reconstructed frame is the left video reconstructed frame at time t-1; The DAN stereo video compressor includes: a feature extraction module, a motion estimation module, a disparity estimation module, a coding and decoding module based on LGEDB, a motion compensation module, a disparity compensation module, a dual-branch high-frequency information fusion module and an image reconstruction module; S3: storing the left video reconstructed frame at time t and the right video reconstructed frame at time t to obtain a compressed stereoscopic video.
2. A dual-branch attention stereo video compression method according to claim 1, characterized in that: An image compression method is used for the first video frame in the left video frame sequence and the right video frame sequence to obtain the corresponding left video reconstructed frame and right video reconstructed frame.
3. A dual-branch attention stereo video compression method according to claim 1, characterized in that: The reconstruction process of the DAN stereo video compressor includes: S21: The left video frame at time t Right video frame at time t Reconstructed frame of the right video at time t-1 and the reconstructed frame of the left video at time t-1 Input feature extraction module to extract features and obtain corresponding features feature feature and Features S22: Features and Features Input motion estimation module to calculate motion information And the characteristics and Features Input motion estimation module to calculate motion information S23: The motion information and The LGEDB-based encoding and decoding module is input to extract the local non-repetitive complex area texture features and global information to obtain motion reconstruction information. and S24: Reconstructing information based on motion Using motion compensation module to Perform motion compensation to obtain feature F MR , and reconstruct information based on motion The motion compensation module is used to t L Perform motion compensation to obtain feature F ML ; S25: According to the characteristics and Features Calculate disparity information using the disparity estimation module And according to the characteristics and Features Calculate disparity information using the disparity estimation module S26: Parallax information and The LGEDB-based encoding and decoding module is input separately to extract the local non-repetitive complex area texture features and global information to obtain the disparity reconstruction information and S27: Reconstructing information based on parallax Using the parallax compensation module to correct the features Perform parallax compensation to obtain feature F DR , based on the disparity reconstruction information Using the parallax compensation module to correct the features Perform parallax compensation to obtain feature F DL ; S28: Feature F MR and feature F DR Input the dual-branch high-frequency information fusion module to perform feature fusion to obtain the prediction features The feature F ML and feature F DL Input the dual-branch high-frequency information fusion module to perform feature fusion to obtain the prediction features S29: Features Subtract Features Get residual information The characteristics Subtract Features Get residual information The residual information and residual information The LGEDB-based encoding and decoding module is input separately to extract the local non-repetitive complex area texture features and global information to obtain the residual reconstruction information and Reconstructing information from the residual and predictive features Add features to get reconstructed features Reconstructing information from the residual and predictive features Add features to get reconstructed features Reconstruct features Input the image reconstruction module to obtain the right video reconstructed frame at time t Reconstruct features Input the image reconstruction module to obtain the left video reconstructed frame at time t 4. A dual-branch attention stereo video compression method according to claim 1, characterized in that: The LGEDB-based encoding and decoding module includes: an encoder, a bottleneck layer and a decoder; the encoder includes N convolution modules and an LGEDB module connected in series between any two adjacent convolution modules; the decoder includes N deconvolution modules and an LGEDB module connected in series between any two adjacent deconvolution modules; The encoder is used to encode the input features of the LGEDB-based encoding and decoding module, the bottleneck layer is used to map the encoding features of the encoder, and the decoder is used to decode the mapping features of the bottleneck layer to obtain the output features of the LGEDB-based encoding and decoding module.
5. A dual-branch attention stereo video compression method according to claim 4, characterized in that: The LGEDB module includes: an LWAB module, a GCAB module, a convolutional layer, and M cascaded residual modules; the feature extraction process of the LGEDB module includes: S211: Input the input feature F of the LGEDB module into the LWAB module and the GCAB module for feature extraction to obtain feature f w and feature f c ; S212: The feature f w and feature f c After splicing, the convolutional layer and M cascaded residual modules are sequentially input for feature processing to obtain feature f; S213: Add feature f and input feature F to obtain output feature F of LGEDB module LG-out .
6. A dual-branch attention stereo video compression method according to claim 5, characterized in that: The feature extraction process of the LWAB module is as follows: Given input features F∈R H×W×C , H represents the height of the input feature F, W represents the width of the input feature F, and C represents the number of channels of the input feature F. First, the input feature is divided into non-overlapping blocks of size p×p, that is, Indicates the number of blocks divided; X is converted into k Divide into multiple heads along the channel direction and generate queries key Sum matrix, The number of channels assigned to each head, s = 1, 2, ..., S, S is the number of heads, and the generation formula is: in, Represent the linear projection matrix of the sth head in the kth block respectively; calculate the feature similarity of the sth head in the kth block Among them, E is the relative position encoding; because there are S heads, S self-attention calculations are performed in parallel, and then the attention features calculated by each head are Concatenate from the channel dimension to get feature A k ∈R p×p×C , and then feature A k With X k Add up to get the self-attention feature of the kth block Finally, the self-attention features of each block are spliced back to the original position to obtain the output feature f of LWAB w ∈R H×W×C : Among them, concat(·) means position concatenation.
7. A dual-branch attention stereo video compression method according to claim 5, characterized in that: The feature extraction process of the GCAB module is as follows: Given input features F∈R H×W×C , H represents the height of the input feature F, W represents the width of the input feature F, and C represents the number of channels of the input feature F. The calculation process is: f' = Conv3(ReLU(Conv3(F))) f”=ReLU(Conv1(AvgPool(f’))) f c =Sigmoid(Conv1(f”))+f' Among them, Conv3(·) represents a 3×3 standard convolution operation, Conv1(·) represents a 1×1 standard convolution operation, ReLU(·) represents a ReLU activation function, Sigmoid(·) represents a Sigmoid activation function; AvgPool(.) represents an average pooling operation.
8. A dual-branch attention stereo video compression method according to claim 4, characterized in that: The bottleneck layer maps the encoding features of the encoder including: (r,(μ,σ))=CWAEM(y) Among them, y represents the encoded features after encoding by the encoder, μ, σ and r represent the mean, variance and latent space residual respectively; Q(·) is the quantization process; CWAEM(·) represents the channel autoregressive entropy model; AE(·) represents the arithmetic encoding process, and AD(·) represents the arithmetic decoding process.
9. A dual-branch attention stereo video compression method according to claim 5, characterized in that: The feature fusion process of the dual-branch high-frequency information fusion module includes: θ=INN(INN(F M )+INN(F D )) Among them, F M ∈{F MR , F ML }, F D ∈{F DR , F DL }, NAF(·) represents NAF module, INN(·) represents reversible neural network; RB( ) represents residual module.
10. A dual-branch attention stereo video compression method according to claim 5, characterized in that: The loss function used in the training of the DAN stereo video compressor includes: Where L represents the loss function, D(·) is the distortion loss, represents the video frame of view v at time t; represents the reconstructed frame of the video of view v at time t, and R(y) represents the bit rate of the encoding feature of the encoder during the compression process.
Citation Information
Patent Citations
User abnormal comment detection method and system based on hierarchical multi-channel attention
CN112015862A
Video compression method and system based on attention mechanism
CN117061760A
Image deblurring method based on attention mechanism residual Fourier transform network
CN117455804A
Deep learning image compression method based on space-channel mixed attention
CN118612467A
Video defogging method and system based on multi-scale space-time fusion network
CN119107268A