A context-based binocular video compression method
By introducing optical flow motion and parallax estimation modules, combining pyramid optical flow and super-priori coding, the problem of large residual prediction of motion information and instability in binocular video compression is solved, and a more efficient and stable video compression effect is achieved.
Patent Information
- Application Number
- CN202510550540.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-29
AI Technical Summary
When the existing neural video compression method processes binocular video, the prediction residuals of motion information are large, the inter-frame encoding effect is poor, and the deformable convolution is unstable during the training process, affecting the compression efficiency and quality.
The optical flow motion estimation module and the optical flow parallax estimation module are used to calculate the motion and parallax information through the pyramid optical flow estimation network, combine the feature extraction and compensation module, and use the context information for video compression, and introduce the super prior coding model auxiliary entropy encoder and entropy decoder to improve model stability and compression efficiency.
The stability and efficiency of binocular video compression are improved, and the video reconstruction quality and compression ratio are improved through multi-scale optical flow estimation and context information fusion, which is suitable for video compression in real scenes.
Smart Images

Figure CN120075449B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video compression, and in particular to a context-based binocular video compression method. Background Art
[0002] With the rapid development of multimedia technology and the widespread popularity of Internet applications, the massive growth of high-definition video data has brought huge challenges to storage and transmission. Video compression technology effectively reduces the size of video files by removing redundant information in video data while trying to maintain the visual quality of the video as much as possible.
[0003] Most existing neural video compression methods adopt a predictive residual coding framework. However, due to video motion, new content often appears at object boundaries, making it difficult to find good references in previous frames, thus facing challenges such as large residuals and poor inter-frame coding effects. The context-based video compression framework uses the context in the feature domain as a condition to more flexibly learn spatio-temporal correlations and improve the quality of the reconstructed video. However, current research mostly focuses on monocular video compression and explores less on binocular video and multi-viewpoint stereo video compression. In binocular video, in addition to the redundancy within viewpoints, the correlation between viewpoints also provides room for optimization of the compression algorithm. Effectively utilizing this redundant information can further reduce the size of video files and improve compression performance.
[0004] The Chinese patent application with the publication number CN119232941A discloses a binocular video compression method based on deep learning, which uses deformable convolution to process motion information and inter-view information, achieving high-quality video compression by utilizing the redundancy between viewpoints and the redundancy within the same viewpoints. However, deformable convolution may have problems such as offset overflow during the training process and is relatively unstable. Summary of the Invention
[0005] Considering the relatively unstable training and other problems in the prior art, in contrast, optical flow, as a relatively mature motion modeling method in computer vision, already has many research results and high-quality pre-trained models. Based on this, this method introduces pyramid optical flow to replace the original deformable convolution scheme, enhancing the stability of the model while ensuring compression efficiency.
[0006] The technical solution adopted by the present invention to solve its technical problems is: to provide a context-based binocular video compression method, which uses the first channel and the second channel with the same structure to process the frame sequences of the right viewpoint and the left viewpoint frame by frame respectively. The specific process of channel processing is as follows:
[0007] The optical flow motion estimation module uses a spatial pyramid optical flow estimation network to calculate motion information. The feature extraction module extracts the shallow features of the motion reference frame. The motion information is input into the motion compression module for compression and then, together with the shallow features of the motion reference frame, is input into the motion compensation module to obtain motion context information. The optical flow disparity estimation module uses a spatial pyramid optical flow estimation network to calculate disparity information. The feature extraction module extracts the shallow features of the disparity reference frame. The disparity information is input into the disparity compression module for compression and then, together with the shallow features of the disparity reference frame, is input into the disparity compensation module to obtain disparity context information. The feature fusion module and the context refinement module are used to fuse the motion context information and the disparity context information to obtain the final context information, which is successively input into the context encoding module, the context compression module, and the context decoding module to obtain the reconstructed frame after compression of the current frame;
[0008] Among them, the first channel includes a first optical flow motion estimation module, a first motion compression module, a first motion compensation module, a first optical flow disparity estimation module, a first disparity compression module, a first disparity compensation module, a first feature extraction module, a first feature fusion module, a first context refinement module, a first context encoding module, a first context compression module, and a first context decoding module; the second channel includes a second optical flow motion estimation module, a second motion compression module, a second motion compensation module, a second optical flow disparity estimation module, a second disparity compression module, a second disparity compensation module, a second feature extraction module, a second feature fusion module, a second context refinement module, a second context encoding module, a second context compression module, and a second context decoding module.
[0009] Preferably, the context-based binocular video compression method includes the following steps:
[0010] Obtain the original frame at the current moment of the right view point as the first current frame, obtain the reconstructed frame at the previous moment of the right view point as the first motion reference frame, input it into the first optical flow motion estimation module to calculate the first motion information, and use the first motion compression module for compression to obtain the first compressed motion information; use the first feature extraction module to extract the shallow features of the first motion reference frame and input them into the first motion compensation module, and perform compensation based on the first compressed motion information to obtain the first motion context information; the first motion reference frame at the initial moment is the right view point I frame compressed from the first frame of the right view point;
[0011] Obtain the original frame of the left view at the current moment as the second current frame, obtain the reconstructed frame of the left view at the previous moment as the second motion reference frame, input it into the second optical flow motion estimation module to calculate the second motion information, and use the second motion compression module for compression to obtain the second compressed motion information; use the second feature extraction module to extract the shallow features of the second motion reference frame and input them into the second motion compensation module, and perform compensation based on the second compressed motion information to obtain the second motion context information; the second motion reference frame at the initial moment is the left view I frame compressed from the first frame of the left view;
[0012] The first optical flow disparity estimation module uses the second motion context information as the first disparity reference frame, combines it with the first current frame to calculate the first disparity information, and uses the first disparity compression module for compression to obtain the first compressed disparity information; uses the first feature extraction module to extract the shallow features of the first disparity reference frame and input them into the first disparity compensation module, and perform compensation based on the first compressed disparity information to obtain the first disparity context information; uses the first feature fusion module to fuse and encode the first motion context information and the first disparity context information to obtain the first rough context information; uses the first context refinement module to refine the first rough context to obtain the first final context information;
[0013] The first context encoding module maps the first current frame to the first latent representation with the first final context information as the condition; uses the first context compression module to compress the first latent representation to obtain the first compressed latent representation; uses the first context decoding module to upsample the first compressed latent representation into features with the original resolution, and then concatenate the upsampled features with the first final context information to obtain the first reconstructed frame at the current moment;
[0014] The second optical flow disparity estimation module uses the first reconstructed frame at the current moment as the second disparity reference frame, combines it with the second current frame to calculate the second disparity information, and uses the second disparity compression module for compression to obtain the second compressed disparity information; uses the second feature extraction module to extract the shallow features of the second disparity reference frame and input them into the second disparity compensation module, and perform compensation based on the second compressed disparity information to obtain the second disparity context information; uses the second feature fusion module to fuse and encode the second motion context information and the second disparity context information to obtain the second rough context information; uses the second context refinement module to refine the second rough context to obtain the second final context information;
[0015] The second context encoding module maps the second current frame to a second latent representation conditioned on the second final context information; compresses the second latent representation using the second context compression module to obtain a second compressed latent representation; uses the second context decoding module to upsample the second compressed latent representation into features with the original resolution, and then concatenates the upsampled features with the second final context information to obtain the second reconstructed frame at the current moment;
[0016] Taking the current moment as the previous moment and the next moment as the current moment, repeat the above steps until the reconstructed frames of all frames in the frame sequence are obtained, and combine the reconstructed frame sequences of the right view point and the left view point into a compressed stereoscopic video.
[0017] Preferably, the first motion information is expressed as:
[0018] ;
[0019] ;
[0020] ;
[0021] ;
[0022] ;
[0023] Wherein, represents the first motion information at time t; represents the spatial pyramid optical flow estimation network; represents the downsampling result of the original frame of the right view point at time t for inputting into the k-th pyramid level, represents the downsampling result of the reconstructed frame of the right view point at time t-1 for inputting into the k-th pyramid level; represents the motion optical flow field of the k-th pyramid level of the right view point, and the initial motion optical flow field of the right view point ; represents the motion residual optical flow of the k-th pyramid level of the right view point; represents the image deformation operation;
[0024] The second motion information is expressed as:
[0025] ;
[0026] ;
[0027] ;
[0028] ;
[0029] ;
[0030] wherein, represents the second motion information at time t; k represents the pyramid level, k = 0, 1, 2, 3, 4; represents the downsampling function; represents the downsampling result of the original left - view frame at time t for input to the k - th pyramid level, represents the downsampling result of the reconstructed left - view frame at time t - 1 for input to the k - th pyramid level; represents the upsampling function; represents the optical flow field of motion in the k - th pyramid level of the left - view point, the initial optical flow field of the left - view point ; represents the residual optical flow of motion in the k - th pyramid level of the left - view point; represents the trained convolutional neural network model.
[0031] Preferably, the first disparity information is expressed as:
[0032] ;
[0033] ;
[0034] ;
[0035] ;
[0036] ;
[0037] wherein, represents the first disparity information at time t; represents the downsampling result of the left - view motion context information at time t ; represents the view - point optical flow field of the k - th layer of the right - view pyramid; represents the view - point optical flow field of the (k - 1) - th layer of the right - view pyramid; represents the view - point residual optical flow field of the k - th layer of the right - view pyramid;
[0038] The second disparity information is expressed as:
[0039] ;
[0040] ;
[0041] ;
[0042] ;
[0043] ;
[0044] Among them, represents the disparity information of the left view point at time t; represents the downsampling result of the reconstructed frame of the right view point at time t ; represents the view point optical flow field of the k-th layer pyramid of the left view point; represents the view point optical flow field of the (k - 1)-th layer pyramid of the left view point; represents the view point residual optical flow field of the k-th layer pyramid of the left view point.
[0045] Preferably, the calculation process of the first motion compression module includes the following steps:
[0046] Use an encoding network to map the first motion information into a first motion latent representation, expressed as:
[0047] ;
[0048] Among them, represents the first motion latent representation at time t; represents a convolutional layer function with a convolutional kernel of ; represents a normalization layer;
[0049] Quantize and entropy code the first motion latent representation to obtain a first motion entropy coding result; the first entropy coding result is obtained through a decoding network to obtain first motion compression information, expressed as:
[0050] ;
[0051] Among them, represents the first motion compression information at time t; represents the first motion entropy coding result at time t; represents a transposed convolutional layer function with a convolutional kernel of ; represents an inverse normalization layer;
[0052] Use a hyperprior coding network to capture the hidden information of the first motion latent representation, generate the hyperprior edge information of the first motion latent representation; the hyperprior edge information of the first motion latent representation is quantized and entropy coded and then fed back to the entropy encoder and entropy decoder of the image compression autoencoder network to assist the entropy model to more accurately model the probability distribution of the latent representation;
[0053] The calculation process of the second motion compression module includes the following steps:
[0054] Use an encoding network to map the second motion information into a second motion latent representation, expressed as:
[0055] ;
[0056] Wherein, represents the second motion latent representation at time t;
[0057] Quantize and entropy-encode the second motion latent representation to obtain a second motion entropy encoding result. The second motion entropy encoding result is decoded through a decoding network to obtain second motion compression information, expressed as:
[0058] ;
[0059] Wherein, represents the second motion compression information at time t; represents the second motion entropy encoding result at time t;
[0060] Use a hyperprior encoding network to capture the hidden information of the second motion latent representation and generate the hyperprior marginal information of the second motion latent representation; the hyperprior marginal information of the second motion latent representation is quantized and entropy-encoded and then fed back to the entropy encoder and entropy decoder to assist the entropy model in more accurately modeling the probability distribution of the latent representation;
[0061] The calculation process of the first disparity compression module includes the following steps:
[0062] Use an encoding network to map the first disparity information into a first disparity latent representation, expressed as:
[0063] ;
[0064] Wherein, represents the first disparity latent representation at time t;
[0065] Quantize and entropy-encode the first disparity latent representation to obtain a first disparity entropy encoding result. The first disparity entropy encoding result is decoded through a decoding network to obtain first disparity compression information, expressed as:
[0066] ;
[0067] Wherein, represents the first disparity compression information at time t; represents the first disparity entropy encoding result at time t;
[0068] The hyperprior encoding network is used to capture the hidden information of the first disparity latent representation and generate the hyperprior edge information of the first disparity latent representation; the hyperprior edge information of the first disparity latent representation is fed back to the entropy encoder and entropy decoder of the image compression autoencoder network after quantization and entropy encoding, assisting the entropy model to more accurately model the probability distribution of the latent representation;
[0069] The calculation process of the second disparity compression module includes the following steps:
[0070] The encoding network is used to map the second disparity information into a second disparity latent representation, expressed as:
[0071] ;
[0072] where represents the second disparity latent representation at time t;
[0073] The second disparity latent representation is quantized and entropy encoded to obtain the second disparity entropy encoding result, and the second disparity entropy encoding result is decoded through the decoding network to obtain the second disparity compression information, expressed as:
[0074] ;
[0075] where represents the second disparity compression information at time t; represents the second disparity entropy encoding result at time t;
[0076] The hyperprior encoding network is used to capture the hidden information of the second disparity latent representation and generate the hyperprior edge information of the second disparity latent representation; the hyperprior edge information of the second disparity latent representation is fed back to the entropy encoder and entropy decoder after quantization and entropy encoding, assisting the entropy model to more accurately model the probability distribution of the latent representation.
[0077] Preferably, the first motion context information is expressed as:
[0078] ;
[0079] ;
[0080] ;
[0081] where represents the first motion context information at time t, represents the shallow features extracted by the first feature extraction module from the first motion reference frame at time t-1, represents the residual block with a convolution kernel of denotes the ReLU activation function, and x denotes the input feature of the residual block. denotes the output feature of the residual block;
[0082] The first parallax context information is expressed as:
[0083] ;
[0084] ;
[0085] where denotes the first parallax context information at time t; denotes the shallow feature extracted by the first feature extraction module from the first parallax reference frame at time t;
[0086] The second motion context information is expressed as:
[0087] ;
[0088] ;
[0089] where denotes the second motion context information at time t; denotes the shallow feature extracted by the second feature extraction module from the second motion reference frame at time t - 1;
[0090] The second parallax context information is expressed as:
[0091] ;
[0092] ;
[0093] where denotes the second parallax context information at time t; denotes the shallow feature extracted by the second feature extraction module from the second parallax reference frame at time t.
[0094] Preferably, the first final context information is expressed as:
[0095] ;
[0096] ;
[0097] where denotes the first final context information at time t; denotes the first rough context information at time t; denotes concatenating the features in the channel dimension;
[0098] The second final context information is expressed as:
[0099] ;
[0100] ;
[0101] Among them, represents the second final context information at time t, represents the second rough context information at time t.
[0102] Preferably, the first latent representation and the second latent representation are respectively represented as:
[0103] ;
[0104] ;
[0105] Among them, represents the first latent representation at time t, represents the second latent representation at time t; represents a normalization function.
[0106] Preferably, the process of the first context compression module compressing the first latent representation includes the following steps:
[0107] Using a hyperprior model to obtain first hyperprior edge information according to the first latent representation;
[0108] Using a time information extraction network to extract features from the first context information to obtain first context time prior information, represented as:
[0109] ;
[0110] Among them, represents the first context time prior information at time t;
[0111] Fusing the first hyperprior edge information and the first context time prior information to learn the probability distribution of the quantized first latent representation, and feeding it back to the entropy encoder and entropy decoder of the image compression autoencoder network to assist the entropy model in more accurately modeling the probability distribution of the first latent representation, obtaining a first compressed latent representation;
[0112] The process of the second context compression module compressing the second latent representation includes the following steps:
[0113] Using a hyperprior model to obtain second hyperprior edge information according to the second latent representation;
[0114] Using a time information extraction network Extract features from the second context information to obtain the second context time prior information, expressed as:
[0115] ;
[0116] where represents the second context time prior information at time t;
[0117] Fuse the second hyperprior edge information and the second context time prior information to learn the probability distribution of the quantized second latent representation, and feedback it to the entropy encoder and entropy decoder of the image compression autoencoder network, assisting the entropy model to more accurately model the probability distribution of the second latent representation, and obtaining the second compressed latent representation.
[0118] Preferably, the first reconstructed frame is expressed as:
[0119] ;
[0120] ;
[0121] where represents the first reconstructed frame at time t; represents the context decoder; represents the first compressed latent representation at time t;
[0122] The second reconstructed frame is expressed as:
[0123] ;
[0124] ;
[0125] where represents the second reconstructed frame at time t, represents the second compressed latent representation at time t.
[0126] The present invention has the following beneficial effects:
[0127] (1) The present invention designs an optical flow motion estimation module and an optical flow disparity estimation module to estimate motion information and disparity information, uses pyramid optical flow to model the motion information through a multi-scale structure, calculates the optical flow step by step on images with different resolutions, improves the motion modeling accuracy, and at the same time has a clearer structure and more stable training;
[0128] (2) The present invention introduces the context time prior information into the entropy coding model, uses the time prior encoder to mine the time correlation in the video sequence, thereby assisting the entropy model to more accurately estimate the probability distribution of the context latent representation and improving the video compression ratio;
[0129] (3) The present invention designs a context generation unit and a context encoding and decoding module, and uses high-dimensional context to transmit rich information to the context encoder and the context decoder, and assists video encoding and decoding with the context information in the feature domain, which helps to reconstruct high-frequency content to obtain higher video quality.
[0130] The following further describes the present invention in detail with reference to the drawings and embodiments, but the present invention is not limited to the embodiments. Brief Description of the Drawings
[0131] Figure 1 It is a flowchart of the steps of the context-based binocular video compression method according to an embodiment of the present invention;
[0132] Figure 2 It is a schematic diagram of a dual-channel model of the context-based binocular video compression method according to an embodiment of the present invention;
[0133] Figure 3 It is a schematic diagram of the context compression module of the context-based binocular video compression method according to an embodiment of the present invention;
[0134] Figure 4 It is a comparison chart of the compression effects between an embodiment of the present invention and LSVC. Specific Embodiments
[0135] The present invention provides a context-based binocular video compression method, which is implemented by using a dual-channel model. Specifically, the dual-channel model processes the frame sequences of the right view point and the left view point frame by frame respectively using a first channel and a second channel; wherein, the first channel includes a first context generation unit and a first video reconstruction unit, and the second channel includes a second context generation unit and a second video reconstruction unit.
[0136] Specifically, the first context generation unit includes a first optical flow motion estimation module, a first motion compression module, a first motion compensation module, a first optical flow disparity estimation module, a first disparity compression module, a first disparity compensation module, a first feature extraction module, a first feature fusion module and a first context refinement module; the first video reconstruction unit includes a first context encoding module, a first context compression module and a first context decoding module.
[0137] Specifically, the second context generation unit includes a second optical flow motion estimation module, a second motion compression module, a second motion compensation module, a second optical flow disparity estimation module, a second disparity compression module, a second disparity compensation module, a second feature extraction module, a second feature fusion module and a second context refinement module; the second video reconstruction unit includes a second context encoding module, a second context compression module and a second context decoding module.
[0138] See Figure 1 andFigure 2 As shown in the figure, the context-based binocular video compression method according to the embodiment of the present invention includes the following steps:
[0139] S101: Obtain the original frame of the right view point at the current moment as the first current frame, obtain the reconstructed frame of the right view point at the previous moment as the first motion reference frame, input the first optical flow motion estimation module to calculate the first motion information, and use the first motion compression module for compression to obtain the first compressed motion information; use the first feature extraction module to extract the shallow features of the first motion reference frame and input them into the first motion compensation module, and perform compensation based on the first compressed motion information to obtain the first motion context information;
[0140] S102: Obtain the original frame of the left view point at the current moment as the second current frame, obtain the reconstructed frame of the left view point at the previous moment as the second motion reference frame; input the second optical flow motion estimation module to calculate the second motion information, and use the second motion compression module for compression to obtain the second compressed motion information; use the second feature extraction module to extract the shallow features of the second motion reference frame and input them into the second motion compensation module, and perform compensation based on the second compressed motion information to obtain the second motion context information;
[0141] S103: The first optical flow disparity estimation module uses the second motion context information as the first disparity reference frame, combines it with the first current frame to calculate the first disparity information, and uses the first disparity compression module for compression to obtain the first compressed disparity information; uses the first feature extraction module to extract the shallow features of the first disparity reference frame and input them into the first disparity compensation module, and perform compensation based on the first compressed disparity information to obtain the first disparity context information; uses the first feature fusion module to perform fusion coding on the first motion context information and the first disparity context information to obtain the first rough context information; uses the first context refinement module to refine the first rough context to obtain the first final context information;
[0142] S104: The first context encoding module maps the first current frame to the first latent representation based on the first final context information; uses the first context compression module to compress the first latent representation to obtain the first compressed latent representation; uses the first context decoding module to upsample the first compressed latent representation into features with the original resolution, and then concatenates the upsampled features with the first final context information to obtain the first reconstructed frame at the current moment;
[0143] S105. The second optical flow parallax estimation module uses the first reconstructed frame at the current moment as the second parallax reference frame, combines it with the second current frame to calculate the second parallax information, and compresses it using the second parallax compression module to obtain the second compressed parallax information; the second feature extraction module extracts shallow features from the second parallax reference frame and inputs them into the second parallax compensation module, and compensates based on the second compressed parallax information to obtain the second parallax context information; the second feature fusion module fuses and encodes the second motion context information and the second parallax context information to obtain the second rough context information; the second context refinement module refines the second rough context to obtain the second final context information.
[0144] S106. The second context encoding module maps the second current frame to a second latent representation based on the second final context information; compresses the second latent representation using the second context compression module to obtain the second compressed latent representation; the second context decoding module upsamples the second compressed latent representation into features with the original resolution, and then cascades the upsampled features with the second final context information to obtain the second reconstructed frame at the current moment.
[0145] S107. Take the current moment as the previous moment and the next moment as the current moment, and repeat the above steps until the reconstructed frames of all frames in the frame sequence are obtained, and combine the reconstructed frame sequence of the right view point and the reconstructed frame sequence of the left view point into a compressed stereoscopic video.
[0146] Among them, the first motion reference frame and the second motion reference frame in the first iteration both use the I frame (i.e., the key frame) of the corresponding view point. Three types of frames are defined in H.264, namely I frame, B frame, and P frame. Among them, the I frame is an independent frame with all information, and only the data of this frame is required to complete decoding during decoding; in the embodiment of the present invention, the first frame of each video is compressed into an I frame.
[0147] Taking the left view point as an example below, the specific process of channel processing will be described in detail.
[0148] Specifically, the processing flow of the second context generation unit includes the following steps (the modules involved in the following content all refer to the modules in the second context generation unit).
[0149] S201. The optical flow motion estimation module is used to extract the motion information between the current frame and the motion reference frame. The motion information is expressed as follows:
[0150] ;
[0151] Among them, represents the input image of the left view point, represents the moment of the input frame, Representation The reconstructed frame of the left view point at the moment; among which represents the spatial pyramid optical flow estimation network. First, different-resolution image sequences are obtained through 2-fold downsampling layer by layer, and each time the downsampling halves the width and height of the image. Then, initial optical flow estimation is performed at the top layer of the pyramid. Assuming the initial motion optical flow field is 0, starting from the top layer, the residual flow at each pyramid level is learned through a coarse-to-fine spatial pyramid structure, and the corresponding motion optical flow field is obtained , where k (k = 0, 1, 2, 3, 4) represents the k-th pyramid level. The specific representation is as follows:
[0152] ;
[0153] ;
[0154] ;
[0155] ;
[0156] ;
[0157] Among them, represents the downsampling function; represents the upsampling function; represents the left view point image that has undergone (4 - k) times of 2-fold downsampling and is input to the k-th pyramid level, represents the reconstructed frame of the previous moment of the left view point that has undergone (4 - k) times of 2-fold downsampling and is input to the k-th pyramid level; represents the motion optical flow field of the k-th pyramid level; represents the motion residual optical flow of the k-th pyramid level; represents the trained convolutional neural network model, which uses the upsampled motion optical flow field of the previous pyramid level and the image frame received by the k-th layer to calculate the motion residual optical flow . represents the image warping operation. In optical flow motion estimation, each pixel in the input image is mapped to the target position according to the optical flow field, thereby generating the transformed output image.
[0158] S202. Use the spatial optical flow pyramid to extract the disparity features between the current frame and the disparity reference frame. The disparity information is represented as follows:
[0159] ;
[0160] ;
[0161] Among them, represents the disparity feature of the left view at time t; represents the reconstructed frame of the right view after (4 - k) times of 2x downsampling at the same time. represents the view optical flow field of the k-th layer pyramid; represents the view optical flow field of the (k - 1)-th layer pyramid; among which, the term represents the view residual optical flow field of the k-th layer pyramid, which can be expressed as .
[0162] S203, adopt a motion compression module, use a hyperprior model to compress the motion information, and obtain the compressed motion information , which is expressed as:
[0163] ;
[0164] represents the entropy coding function, which specifically includes the following two aspects.
[0165] On the one hand, compress the image through an autoregressive network composed of a motion encoding network and a motion decoding network. Map the motion information into a motion latent representation through the motion encoding network , and perform quantization and entropy coding on it to obtain , and obtain the compressed motion information through the motion decoding network. The specific representation is as follows:
[0166] ;
[0167] ;
[0168] Among them, represents the motion encoding network, which maps the motion information into a latent representation , represents the normalization layer; represents the motion decoding network, represents the transposed convolution layer function with a convolution kernel of , represents the inverse normalization layer.
[0169] On the other hand, introduce hyperprior motion edge information through the hyperprior model. Use the hyperprior encoding network to capture the hidden information of the motion latent representation and generate the hyperprior motion edge information of the latent representation , which produces after quantization and entropy coding, and is used as auxiliary information for encoding and transmission. Learn The probability distribution of is fed back to the entropy encoder (AE) and entropy decoder (AD) of the image compression autoencoder network, assisting the entropy model to represent the potential at the current moment The probability distribution of can achieve more accurate modeling, improve the compression ratio, and guide it to better complete image compression and restoration.
[0170] S204, using a disparity compression module to compress the disparity information using a super prior model to obtain compressed disparity information The specific expressions are as follows:
[0171] ;
[0172] The process of the disparity compression module is consistent with that of the motion coding network. On the one hand, the image is compressed through an autoregressive network consisting of a disparity coding network and a disparity decoding network, and the disparity information is mapped into a disparity potential representation through the disparity coding network. , and quantize and entropy encode it to get , compressed motion information is obtained through the motion decoding network On the other hand, the super-prior disparity edge information is introduced through the super-prior model. The super-prior encoding network is used to capture the hidden information of the disparity potential representation and generate the disparity potential representation The super prior parallax edge information , quantized and entropy encoded to produce , encoded and transmitted as auxiliary information, and learned through the hyper-prior model The probability distribution of is fed back to the entropy model of the image compression autoencoder network, namely the entropy encoder (AE) and the entropy decoder (AD), to assist the entropy model in the potential representation The probability distribution of can achieve more accurate modeling and guide it to better complete image compression and restoration.
[0173] S205, using a feature extraction module to extract features of the left viewpoint reconstructed frame at the previous moment and the right viewpoint reconstructed frame at the current moment, the features of the left viewpoint reconstructed frame at the previous moment It is expressed as follows:
[0174] ;
[0175] The feature extraction module It is expressed as follows:
[0176] ;
[0177] in, express The features of the left viewpoint reconstructed frame at time instant, Indicates the time of the input frame. Represents a convolutional layer function with a convolutional kernel of ; Represents a residual block with a convolutional kernel of . The residual block is represented as:
[0178] ;
[0179] where represents the RELU activation function, x represents the input feature, represents the output feature, represents the convolutional kernel, and when j = 3, it represents a convolutional kernel of .
[0180] The feature of the right - view reconstruction frame at time t is represented as follows:
[0181] ;
[0182] where is the right - view reconstruction frame at time t and serves as the disparity reference frame for the left view.
[0183] S206. Using the motion compensation module, the decoded motion feature will guide the network through a warping operation on where to extract context information to obtain high - dimensional motion context information, which is specifically represented as follows:
[0184] ;
[0185] where represents the motion context information of the left view at time t.
[0186] S207 uses the disparity compensation module, and the decoded disparity feature will guide the network through a warping operation on where to extract context information to obtain high - dimensional view - point context information, which is represented as follows:
[0187] ;
[0188] where represents the view - point context information of the left view at time t.
[0189] S208. Using the feature fusion module to fuse the left - view motion context information and the view - point context information together to obtain the rough context information of the left view at time t , which is represented as follows:
[0190] ;
[0191] where It means that the features are concatenated in the channel dimension.
[0192] S209. The adopted context refinement module further refines the rough context information of the left viewpoint obtained in the previous step through residual blocks and convolutional layers to obtain the final context information of the left viewpoint at time t, which is specifically represented as follows: The final context information of the left viewpoint at time t , which is specifically represented as follows:
[0193] ;
[0194] Among them, represents the context refinement module, which is specifically represented as follows:
[0195] .
[0196] Specifically, the processing flow of the video reconstruction module includes the following steps (the modules involved in the following content all refer to the modules in the second context generation unit).
[0197] S301. The context encoding module is adopted to encode the current frame of the left viewpoint at time t into a latent representation with the context information extracted in the previous step as a condition, which is specifically represented as follows: , which is specifically represented as follows:
[0198] ;
[0199] Among them, represents the context encoder, which is specifically represented as follows:
[0200] ;
[0201] Among them, represents the normalization function.
[0202] S302. The context compression module is adopted to quantize and compress the latent representation of the current frame of the left viewpoint. Specifically represented as follows:
[0203] ;
[0204] Among them, represents the fusion prior entropy model. is the hyperprior model, is the temporal prior model.
[0205] Specifically, as shown in Figure 3 input the latent representation of the current frame of the left viewpoint into the hyperprior model to obtain the edge information, and the final context information of the left viewpoint Input the temporal prior model to obtain the contextual temporal prior information. Learn the probability distribution of the quantized current-frame latent representation by fusing the edge information of the current-frame latent representation and the contextual temporal prior information, and feedback it to the entropy models of the image compression autoencoder network, namely the entropy encoder (the second AE) and the entropy decoder (the second AD), to assist the entropy models in more accurately modeling the probability distribution of the latent representation and guiding them to better complete image compression and restoration. The contextual temporal information feature is expressed as:
[0206] .
[0207] S303. Adopt the contextual decoding module. Upsample the reconstructed latent representation in the previous step to the features with the original resolution through the deconvolution layer, the inverse normalization layer, and the residual block, and then use the contextual information as a condition to concatenate the contextual information with the upsampled features to generate the final left-viewpoint reconstructed frame , which is specifically expressed as follows:
[0208] ;
[0209] Among them, represents the reconstruction module function, which is specifically expressed as follows:
[0210] ;
[0211] Among them, represents the contextual decoder, which is specifically expressed as follows:
[0212] ;
[0213] Among them, represents the deconvolution layer with a convolution kernel of , represents the inverse normalization layer.
[0214] When this embodiment works, the development environment adopts Pytorch, uses the Python language for programming, and uses the Nvidia RTX A6000 GPU for experiments. The datasets adopted in this embodiment are Vimeo-90K and Cityscapes. Vimeo-90K contains two subsets. The subset "Septuplet datase" of it is selected in this invention example to train the dual-channel model, with a total of 91,701 groups. Each group of data consists of a sequence of 7 consecutive frame images. Before training, the images are first randomly cropped, and the images with the original resolution of 448×256 are cropped into 256×256. Cityscapes consists of large-scale and diverse stereo video sequences collected on 50 different urban streets. Its training set contains 2,975 video pairs, the validation set contains 500 video pairs, and the test set contains 1,525 video pairs. Each video pair contains two 30-frame video sequences with a resolution of 2048×1034. To improve the adaptability of the model in real scenarios, the images need to be preprocessed before training: crop 64 pixels at the top and 128 pixels on the left side of each frame of the image to remove calibration artifacts; at the same time, crop 240 pixels at the bottom to remove the interference of the structure of the acquisition vehicle itself, so that the model can focus more on the learning of real scenario content. Subsequently, the random cropping technique is adopted to crop each training image into 320×256 for data augmentation.
[0215] The dual-purpose context video compression training of this invention example adopts a phased training strategy. First, the optical flow motion / disparity estimation, motion / disparity compression, and motion / disparity compensation parts are pre-trained on the Vimeo-90K dataset. The parameters of these modules are shared between the left-view and right-view modules. Then, the entire dual-channel compression model is trained end-to-end based on the Cityscapes dataset. During the training process, the initial learning rate is set to 5e -5 , and the Adam optimizer is used for optimization. The batchsize of the training network is set to 4. In terms of performance evaluation, Bpp is used to measure the bit rate, and PSNR and MS-SSIM are used to measure the distortion between the compressed video and the original video.
[0216] The relevant data and comparison results of the PSNR index are shown in Table 1 and Figure 4 As shown, it can be seen from Table 1 that the PSNR values of the embodiments of this invention are higher than 35 at different Bpps, and the compression effect meets the general video compression requirements; Figure 4This is a schematic diagram comparing the compression effects of the embodiments of the present invention with LSVC (learning-based stereo video compression framework). As can be seen from the figure, within a certain range, the compression effect of the embodiments of the present invention at the same bpp is higher than that of LSVC (learning-based stereo video compression framework, Z. Chen, G. Lu, Z. Hu, S. Liu, W. Jiang and D. Xu, "LSVC: A Learning-based Stereo Video Compression Framework," 2022 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 2022, pp. 6063-6072, doi: 10.1109 / CVPR52688.2022.00598.).
[0217] Table 1 - PSNR data at different Bpps:
[0218]
[0219] It can be seen that a context-based binocular video compression method proposed by the present invention utilizes an optical flow disparity estimation module and a disparity compression module to compress the redundancy between viewpoints and improve the video compression performance; the context learning module and the context encoding and decoding module transmit rich information to the context encoder and the context decoder using high-dimensional context, and assist video encoding and decoding conditional on the context information in the feature domain, which helps to reconstruct high-frequency content to obtain higher video quality. Optical flow explicitly models the inter-frame displacement through pixel-by-pixel motion vectors, and the estimated result is a physical meaningful motion field that can provide clear motion trajectories, is more suitable for precise motion analysis, and has higher interpretability and stability. Therefore, the present invention example adopts pyramid optical flow to estimate the motion optical flow step by step from low resolution to high resolution by constructing an image pyramid.
[0220] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A context-based binocular video compression method, characterized in that The first channel and the second channel with the same structure are respectively used to process the frame sequences of the right view point and the left view point frame by frame. The specific channel processing flow is as follows: The optical flow motion estimation module uses a spatial pyramid optical flow estimation network to calculate motion information. The feature extraction module extracts the shallow features of the motion reference frame. The motion information is input into the motion compression module for compression and then input into the motion compensation module together with the shallow features of the motion reference frame to obtain motion context information. The optical flow disparity estimation module uses a spatial pyramid optical flow estimation network to calculate disparity information. The feature extraction module extracts the shallow features of the disparity reference frame. The disparity information is input into the disparity compression module for compression and then input into the disparity compensation module together with the shallow features of the disparity reference frame to obtain disparity context information; The feature fusion module and the context refinement module are used to fuse the motion context information and the disparity context information to obtain the final context information, which is sequentially input into the context encoding module, the context compression module, and the context decoding module to obtain the reconstructed frame after compression of the current frame; Among them, the first channel includes a first optical flow motion estimation module, a first motion compression module, a first motion compensation module, a first optical flow disparity estimation module, a first disparity compression module, a first disparity compensation module, a first feature extraction module, a first feature fusion module, a first context refinement module, a first context encoding module, a first context compression module, and a first context decoding module; the second channel includes a second optical flow motion estimation module, a second motion compression module, a second motion compensation module, a second optical flow disparity estimation module, a second disparity compression module, a second disparity compensation module, a second feature extraction module, a second feature fusion module, a second context refinement module, a second context encoding module, a second context compression module, and a second context decoding module; The context-based binocular video compression method includes the following steps: Obtain the original frame of the right view point at the current moment as the first current frame, obtain the reconstructed frame of the right view point at the previous moment as the first motion reference frame, input it into the first optical flow motion estimation module to calculate the first motion information, and use the first motion compression module for compression to obtain the first compressed motion information; use the first feature extraction module to extract the shallow features of the first motion reference frame and input them into the first motion compensation module, and perform compensation based on the first compressed motion information to obtain the first motion context information; the first motion reference frame at the initial moment is the right view point I frame compressed from the first frame of the right view point; Obtain the original frame of the left view point at the current moment as the second current frame, obtain the reconstructed frame of the left view point at the previous moment as the second motion reference frame, input it into the second optical flow motion estimation module to calculate the second motion information, and use the second motion compression module for compression to obtain the second compressed motion information; use the second feature extraction module to extract the shallow features of the second motion reference frame and input them into the second motion compensation module, and perform compensation based on the second compressed motion information to obtain the second motion context information; the second motion reference frame at the initial moment is the left view point I frame compressed from the first frame of the left view point.
2. The context-based binocular video compression method according to claim 1, characterized in that, The context-based binocular video compression method further includes the following steps: The first optical flow disparity estimation module uses the second motion context information as the first disparity reference frame, combines it with the first current frame to calculate the first disparity information, and compresses it using the first disparity compression module to obtain the first compressed disparity information; uses the first feature extraction module to extract the shallow features of the first disparity reference frame and inputs them into the first disparity compensation module, and compensates based on the first compressed disparity information to obtain the first disparity context information; uses the first feature fusion module to fuse and encode the first motion context information and the first disparity context information to obtain the first rough context information; uses the first context refinement module to refine the first rough context to obtain the first final context information; The first context encoding module maps the first current frame into a first latent representation based on the first final context information; uses the first context compression module to compress the first latent representation to obtain the first compressed latent representation; uses the first context decoding module to upsample the first compressed latent representation into features with the original resolution, and then cascades the upsampled features with the first final context information to obtain the first reconstructed frame at the current moment; The second optical flow disparity estimation module uses the first reconstructed frame at the current moment as the second disparity reference frame, combines it with the second current frame to calculate the second disparity information, and compresses it using the second disparity compression module to obtain the second compressed disparity information; uses the second feature extraction module to extract the shallow features of the second disparity reference frame and inputs them into the second disparity compensation module, and compensates based on the second compressed disparity information to obtain the second disparity context information; uses the second feature fusion module to fuse and encode the second motion context information and the second disparity context information to obtain the second rough context information; uses the second context refinement module to refine the second rough context to obtain the second final context information; The second context encoding module maps the second current frame into a second latent representation based on the second final context information; uses the second context compression module to compress the second latent representation to obtain the second compressed latent representation; uses the second context decoding module to upsample the second compressed latent representation into features with the original resolution, and then cascades the upsampled features with the second final context information to obtain the second reconstructed frame at the current moment; Taking the current moment as the previous moment and the next moment as the current moment, repeat the above steps until the reconstructed frames of all frames in the frame sequence are obtained, and combine the reconstructed frame sequences of the right view point and the left view point into the compressed binocular video.
3. The context-based binocular video compression method according to claim 2, wherein The first motion information is expressed as: Among them, V t r represents the first motion information at time t; f opticFlow (·) represents the spatial pyramid optical flow estimation network; represents the downsampling result of the original frame of the right view point at time t for inputting into the k-th pyramid level, represents the downsampling result of the reconstructed frame of the right view point at time t-1 for inputting into the k-th pyramid level; represents the motion optical flow field of the k-th pyramid level of the right view point, and the initial motion optical flow field of the right view point represents the motion residual optical flow of the k-th pyramid level of the right view point; warp(·) represents the image warping operation; The second motion information is expressed as: Among them, V t l represents the second motion information at time t; k represents the pyramid level, k = 0, 1, 2, 3, 4; d(·) represents the downsampling function; represents the downsampling result of the original left-viewpoint frame at time t for input to the k-th pyramid level, represents the downsampling result of the reconstructed left-viewpoint frame at time t - 1 for input to the k-th pyramid level; u2(·) represents the upsampling function; represents the motion optical flow field of the k-th pyramid level of the left viewpoint, and the initial motion optical flow field of the left viewpoint represents the motion residual optical flow of the k-th pyramid level of the left viewpoint; G k (·) represents a trained convolutional neural network model.
4. The context-based binocular video compression method according to claim 3, characterized in that The first disparity information is expressed as: Among them, represents the first parallax information at time t; represents the downsampling result of the left view motion context information at time t; represents the view optical flow field of the k-th layer pyramid of the right view; represents the view optical flow field of the (k - 1)-th layer pyramid of the right view; represents the view residual optical flow field of the k-th layer pyramid of the right view; The second disparity information is expressed as: Among them, represents the disparity information of the left view at time t; represents the downsampling result of the reconstructed frame of the right view at time t ; represents the view optical flow field of the k-th layer pyramid of the left view; represents the view optical flow field of the (k - 1)-th layer pyramid of the left view; represents the view residual optical flow field of the k-th layer pyramid of the left view.
5. The context-based binocular video compression method according to claim 4, characterized in that The calculation process of the first motion compression module includes the following steps: Using an encoding network to map the first motion information into a first motion latent representation, expressed as: Among them, represents the first motion latent representation at time t; Conv 3×3 (·) represents the convolutional layer function with a convolution kernel of 3×3, and GDN(·) represents the normalization layer; Quantize and entropy-encode the first motion latent representation to obtain the first motion entropy-encoded result; the first entropy-encoded result is passed through a decoding network to obtain the first motion compression information, expressed as: Among them, represents the first motion compression information at time t; represents the first motion entropy coding result at time t; deconv 3×3 (·) represents a deconvolution layer function with a 3×3 convolution kernel, and IGND(·) represents an inverse normalization layer; Use a hyperprior encoding network to capture the hidden information of the first motion latent representation and generate the hyperprior marginal information of the first motion latent representation; the hyperprior marginal information of the first motion latent representation is quantized and entropy-encoded and then fed back to the entropy encoder and entropy decoder of the image compression autoencoder network to assist the entropy model in more accurately modeling the probability distribution of the latent representation; The calculation process of the second motion compression module includes the following steps: Use an encoding network to map the second motion information into a second motion latent representation, expressed as: Among them, represents the second motion latent representation at time t; Quantize and entropy-encode the second motion latent representation to obtain the second motion entropy-encoded result, and the second motion entropy-encoded result is passed through a decoding network to obtain the second motion compression information, expressed as: Among them, represents the second motion compression information at time t; represents the second motion entropy coding result at time t; Use a hyperprior encoding network to capture the hidden information of the second motion latent representation and generate the hyperprior marginal information of the second motion latent representation; the hyperprior marginal information of the second motion latent representation is quantized and entropy-encoded and then fed back to the entropy encoder and entropy decoder to assist the entropy model in more accurately modeling the probability distribution of the latent representation; The calculation process of the first disparity compression module includes the following steps: Use an encoding network to map the first disparity information into a first disparity latent representation, expressed as: Among them, represents the first parallax latent representation at time t; Quantize and entropy-encode the first disparity latent representation to obtain the first disparity entropy-encoded result, and the first disparity entropy-encoded result is passed through a decoding network to obtain the first disparity compression information, expressed as: Among them, represents the first parallax compression information at time t; represents the first parallax entropy coding result at time t; Use a hyperprior encoding network to capture the hidden information of the first disparity latent representation and generate the hyperprior marginal information of the first disparity latent representation; the hyperprior marginal information of the first disparity latent representation is quantized and entropy-encoded and then fed back to the entropy encoder and entropy decoder of the image compression autoencoder network to assist the entropy model in more accurately modeling the probability distribution of the latent representation; The calculation process of the second disparity compression module includes the following steps: Use an encoding network to map the second disparity information into a second disparity latent representation, expressed as: Among them, represents the second parallax latent representation at time t; Quantize and entropy-encode the second disparity latent representation to obtain the second disparity entropy-encoded result, and the second disparity entropy-encoded result is passed through a decoding network to obtain the second disparity compression information, expressed as: Among them, represents the second parallax compression information at time t; represents the second parallax entropy coding result at time t; Use a hyperprior encoding network to capture the hidden information of the second disparity latent representation and generate the hyperprior marginal information of the second disparity latent representation; the hyperprior marginal information of the second disparity latent representation is quantized and entropy-encoded and then fed back to the entropy encoder and entropy decoder to assist the entropy model in more accurately modeling the probability distribution of the latent representation.
6. The context-based binocular video compression method according to claim 5, wherein The first motion context information is expressed as: Among them, represents the first motion context information at time t, represents the shallow features extracted by the first feature extraction module from the first motion reference frame at time t-1, res 3×3 (·) represents a residual block with a 3×3 convolution kernel, represents the RELU activation function, x represents the input features of the residual block, res 3×3 (x) represents the output features of the residual block; The first disparity context information is expressed as: Among them, represents the first parallax context information at time t; represents the shallow features extracted by the first feature extraction module from the first parallax reference frame at time t; The second motion context information is expressed as: Among them, represents the second motion context information at time t; represents the shallow features extracted by the second feature extraction module from the second motion reference frame at time t-1; The second disparity context information is expressed as: Among them, represents the second parallax context information at time t; represents the shallow features extracted by the second feature extraction module from the second parallax reference frame at time t.
7. The context-based binocular video compression method according to claim 6, characterized in that The first final context information is expressed as: Among them, represents the first final context information at time t; represents the first rough context information at time t; concat(·) means concatenating features in the channel dimension; The second final context information is expressed as: Among them, represents the second final context information at time t, represents the second rough context information at time t.
8. The context-based binocular video compression method according to claim 7, wherein The first latent representation and the second latent representation are respectively expressed as: Among them, represents the first latent representation at time t, represents the second latent representation at time t; GDN(·) represents a normalization function.
9. The context-based binocular video compression method according to claim 8, wherein The process of the first context compression module compressing the first latent representation includes the following steps: Obtain the first hyperprior edge information according to the first latent representation by using a hyperprior model; Adopt a time information extraction network f tr (·) Extract features from the first context information to obtain the first context time prior information, expressed as: Among them, represents the first context time prior information at time t; Fuse the first hyperprior edge information and the first context temporal prior information to learn the probability distribution of the quantized first latent representation, and feedback it to the entropy encoder and entropy decoder of the image compression autoencoder network to assist the entropy model to more accurately model the probability distribution of the first latent representation, and obtain the first compressed latent representation; The process of the second context compression module compressing the second latent representation includes the following steps: Obtain the second hyperprior edge information according to the second latent representation by using a hyperprior model; Adopt the time information extraction network f tr (·) Extract features from the second context information to obtain the second context time prior information, expressed as: Among them, represents the second context time prior information at time t; Fuse the second hyperprior edge information and the second context temporal prior information to learn the probability distribution of the quantized second latent representation, and feedback it to the entropy encoder and entropy decoder of the image compression autoencoder network to assist the entropy model to more accurately model the probability distribution of the second latent representation, and obtain the second compressed latent representation.
10. The context-based binocular video compression method according to claim 9, wherein The first reconstructed frame is expressed as: Among them, represents the first reconstructed frame at time t; f dec (·) represents the context decoder; represents the first compressed latent representation at time t; The second reconstructed frame is expressed as: Among them, represents the second reconstructed frame at time t, represents the second compressed latent representation at time t.
Citation Information
Patent Citations
Binocular video compression method and device based on deep learning, and readable medium
CN119232941A
End-to-end stereo image compression method and device based on bidirectional conditional coding
CN114697632A
Unsupervised multi-scale disparity / optical flow fusion
US20220147776A1