Binocular video compression method and device, equipment and storage medium
By using quality factors to generate quality maps in the binocular video compression model, the problem of inability to adapt to video compression in different bit rates in the prior art is solved, and more efficient compression efficiency and better rate distortion performance are achieved.
Patent Information
- Application Number
- CN202510120617.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-25
- Publication Date
- 2025-05-06
AI Technical Summary
The existing binocular video compression scheme cannot adapt to different bitrate video compression, which increases training and storage costs and has low compatibility.
By obtaining the current frame and reconstructed frame in the binocular video, the target quality factor is determined according to the correspondence between the bit rate and the quality factor, a quality map is generated, and input it into the pre-trained binocular video compression model for reconstruction, covering a larger bit rate compression range.
It significantly reduces training time and resource consumption, improves compression efficiency and rate distortion performance, and supports dynamic adjustment of compression quality to cover different bit rate requirements.
Smart Images

Figure CN119946282A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a binocular video compression method, device, equipment and storage medium. Background Art
[0002] Stereo cameras are widely used in autonomous driving scenarios because they can obtain more accurate depth information than monocular cameras and are less expensive. However, with the large amount of binocular video data generated, how to efficiently store and transmit binocular videos has become an important research topic. In the existing technology, binocular video compression requires training multiple models for different bit rates, which increases training and storage costs, making it less compatible. Summary of the invention
[0003] The object of the present invention is to provide a binocular video compression method, device, equipment and storage medium to solve the problem that the existing binocular video compression scheme cannot adapt to different bit rate video compression.
[0004] The first aspect of the present application provides a binocular video compression method, comprising:
[0005] Acquire a current frame of a first viewpoint in a binocular video, a first reconstructed frame of a previous frame of the first viewpoint, and a second reconstructed frame of a current frame of a second viewpoint, wherein the first viewpoint includes a left viewpoint and a right viewpoint;
[0006] Determine a target quality factor of the binocular video according to a corresponding relationship between a bit rate and a quality factor, and generate a quality map based on the target quality factor;
[0007] The quality map, the current frame, the first reconstructed frame and the second reconstructed frame are input into a pre-trained binocular video compression model to reconstruct the current frame, obtain a reconstructed frame of the current frame, and store it.
[0008] Optionally, the binocular video compression model includes: a feature extraction module, an estimation module, a joint compression module, a compensation fusion module, a residual compression module and an image reconstruction module;
[0009] The step of inputting the quality map, the current frame, the first reconstructed frame and the second reconstructed frame into a pre-trained binocular video compression model to reconstruct the current frame to obtain the reconstructed frame of the current frame includes:
[0010] Inputting the current frame, the first reconstructed frame and the second reconstructed frame into the feature extraction module to generate corresponding latent representations, wherein the latent representations include the current frame latent representation, the first reconstructed frame latent representation and the second reconstructed frame latent representation;
[0011] Inputting the current frame potential representation, the first reconstructed frame potential representation and the second reconstructed frame potential representation into the estimation model to perform information estimation to obtain motion information and disparity information of the current frame;
[0012] Inputting the motion information and the disparity information into the joint compression module for joint compression encoding and decoding to obtain reconstructed motion information and reconstructed disparity information of the current frame;
[0013] Inputting the reconstructed motion information, the reconstructed disparity information, the first reconstructed frame potential representation and the second reconstructed frame potential representation into the compensation fusion module for compensation fusion to obtain a comprehensive prediction feature of the current frame;
[0014] The compression code rate of the residual compression module is controlled by the quality map to perform residual compression reconstruction on the potential representation of the current frame and the comprehensive prediction features, and a reconstructed frame in the pixel domain of the current frame is output through the image reconstruction module.
[0015] Optionally, the step of inputting the motion information and the disparity information into the joint compression module for joint compression encoding and decoding to obtain the reconstructed motion information and the reconstructed disparity information of the current frame includes:
[0016] Inputting the motion information and the disparity information into an encoder of the joint compression module to generate a joint latent representation;
[0017] quantizing the joint potential representation, and inputting the quantized reconstructed joint potential representation into the super-a priori encoder of the joint compression module for processing to obtain super-a priori information;
[0018] Inputting the super-prior information and context information into the entropy modeling module of the joint compression module for encoding to obtain distribution parameters of the reconstructed joint potential representation, wherein the context information is a reconstructed joint potential representation sequence of each previous frame;
[0019] Decoding is performed based on the distribution parameters, the reconstructed joint potential representation and the context information to obtain reconstructed motion information and reconstructed disparity information of the current frame.
[0020] Optionally, the step of inputting the super-prior information and the context information into the entropy modeling module of the joint compression module for encoding to obtain distribution parameters for reconstructing the joint potential representation includes:
[0021] generating a content-aware tag based on the hyper-prior information and the contextual information;
[0022] Generate fusion information based on the reconstructed joint latent representation and the content-aware label fusion;
[0023] The content-aware label and the reconstructed joint latent representation are subjected to bidirectional interaction and cross-attention fusion to obtain distribution parameters of the reconstructed joint latent representation.
[0024] Optionally, the generating fusion information based on the reconstructed joint potential representation and the content-aware label fusion includes:
[0025] The sum of the dot product operations of the reconstructed joint potential representation and the content-aware label and the occlusion label is calculated to obtain fusion information.
[0026] Optionally, after inputting the quality map, the current frame, the first reconstructed frame and the second reconstructed frame into a pre-trained binocular video compression model to reconstruct the current frame and obtain the reconstructed frame of the current frame, the method further includes:
[0027] The temporal consistency loss is introduced to optimize the loss of motion information and visual information of each frame in the binocular video, and to optimize the smoothness loss between frames.
[0028] Optionally, the introducing of the temporal consistency loss to optimize the loss of motion information and visual information of each frame in the binocular video, and the optimization of the smoothness loss between frames, includes:
[0029] Calculate the norm of the motion information of the current frame of the first viewpoint and the motion information of the previous frame and the sum of the norm of the motion information of the current frame of the second viewpoint and the motion information of the previous frame, and optimize the motion information loss based on the sum of the norms;
[0030] Calculate the norm of the visual information of the current frame of the first viewpoint and the visual information of the previous frame after the warping operation and the sum of the norms of the visual information of the current frame of the second viewpoint and the visual information of the previous frame after the warping operation, and optimize the visual information loss based on the sum of the norms;
[0031] Calculate the norm of the reconstructed potential representation of the current frame of the first viewpoint and the reconstructed potential representation of the previous frame after the warping operation and the sum of the norms of the reconstructed potential representation of the current frame of the second viewpoint and the reconstructed potential representation of the previous frame after the warping operation, and optimize the inter-frame smoothness loss based on the sum of the norms.
[0032] A second aspect of the present application provides a binocular video compression device, comprising:
[0033] An acquisition module, used to acquire a current frame of a first viewpoint, a first reconstructed frame of a previous frame of the first viewpoint, and a second reconstructed frame of a current frame of a second viewpoint in a binocular video, wherein the first viewpoint includes a left viewpoint and a right viewpoint;
[0034] A determination module, configured to determine a target quality factor of the binocular video according to a correspondence between a bit rate and a quality factor, and generate a quality map based on the target quality factor;
[0035] A compression module is used to input the quality map, the current frame, the first reconstructed frame and the second reconstructed frame into a pre-trained binocular video compression model to reconstruct the current frame, obtain a reconstructed frame of the current frame, and store it.
[0036] The third aspect of the present application provides an electronic device, comprising: a memory, at least one processor and a display, wherein instructions are stored in the memory; the at least one processor calls the instructions in the memory so that the electronic device executes the above-mentioned binocular video compression method.
[0037] A fourth aspect of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the above-mentioned binocular video compression method.
[0038] In the technical solution provided by the present application, the current frame of the first viewpoint in the binocular video, the first reconstructed frame of the previous frame of the first viewpoint, and the second reconstructed frame of the current frame of the second viewpoint are obtained, wherein the first viewpoint includes the left viewpoint and the right viewpoint; the target quality factor of the binocular video is determined according to the correspondence between the bit rate and the quality factor, and a quality map is generated based on the target quality factor; the quality map, the current frame, the first reconstructed frame, and the second reconstructed frame are input into a pre-trained binocular video compression model to reconstruct the current frame, obtain the reconstructed frame of the current frame, and store it. The present application generates a quality map through a quality factor, thereby guiding the transmission information to be compressed at different bit rates, covering a larger bit rate compression range in a single model, thereby significantly reducing training time and resource consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 A schematic diagram of an embodiment of a binocular video compression method in an embodiment of the present application;
[0040] Figure 2 This is a structural diagram of a binocular video compression model in an embodiment of the present application;
[0041] Figure 3 It is a motion parallax joint compression network framework based on conditional entropy coding in an embodiment of the present application;
[0042] Figure 4 It is the Transformer entropy model network framework of the joint context in the embodiment of the present application;
[0043] Figure 5 A schematic diagram of an embodiment of a binocular video compression device in an embodiment of the present application;
[0044] Figure 6 This is a schematic diagram of an embodiment of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION
[0045] The present application provides a binocular video compression method, apparatus, device and storage medium, which covers a larger code rate compression range in a single model through quality factor guidance, thereby significantly saving training time. In addition, the method reduces intra-frame redundancy by jointly compressing disparity information and motion information, and uses context information as a priori condition to reduce inter-frame redundancy. At the same time, the Transformer entropy model architecture is used to fuse context information and super prior information to achieve more efficient conditional entropy coding. In order to further improve the stability of inter-frame rate distortion performance, temporal consistency loss is introduced, which greatly improves the consistency and performance of video compression in the temporal dimension.
[0046] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0047] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 , an embodiment of the binocular video compression method in the embodiment of the present application includes:
[0048] 101. Acquire a current frame of a first viewpoint in a binocular video, a first reconstructed frame of a previous frame of the first viewpoint, and a second reconstructed frame of a current frame of a second viewpoint, wherein the first viewpoint includes a left viewpoint and a right viewpoint.
[0049] It is understandable that the execution subject of the present application may be a binocular video compression device, or an electronic device, an electronic device or a server, which is not specifically limited here. The present application embodiment is described by taking a server as the execution subject as an example.
[0050] In this embodiment, the method includes right viewpoint compression and left viewpoint compression, wherein each frame of the binocular video is first compressed from the right viewpoint and then compressed from the left viewpoint, where compression is frame reconstruction. The left viewpoint and the right viewpoint are distinguished according to the position of the camera on the vehicle.
[0051] If the first viewpoint is a right viewpoint, the second reconstructed frame of the current frame of the second viewpoint is obtained by prediction, specifically, by first completing the motion estimation of the left viewpoint to obtain motion information Then, a motion-compensated warping operation is used to transform the latent representation of the previous frame of the left viewpoint According to sports information Compensate to get the potential representation of the left viewpoint of the current frame as a reference frame; if the first viewpoint is a left viewpoint, a second reconstructed frame of the current frame of the second viewpoint is obtained from the buffer.
[0052] 102. According to the corresponding relationship between the bit rate and the quality factor, determine the target quality factor of the binocular video, and generate a quality map based on the target quality factor.
[0053] It should be noted that the quality factor is one of the improvements in this solution. By setting the quality factor, the dual-target video compression model is matched with video data of different bit rates, thereby achieving compression processing.
[0054] The following takes the compression of the left view as an example to explain that a quality map is generated for each frame according to the quality factor q, thereby controlling the compression bit rate. Figure 2 As shown, the generator regulates the current frame of the left view in the pixel domain through the quality factor q The quality of the image is used to generate a quality map m. The quality map m is used to adjust the compression quality in a weighted manner in the compression network to control the compression bit rate of the binocular video.
[0055] At the same time, each frame compresses the right view first, and then the left view. The right view compression framework is similar to the left view compression framework. Compared with the left view network framework, the difference of the right view network is that the left view reference frame used to calculate the disparity information is not directly obtained from the buffer, but the motion information is obtained by completing the motion estimation of the left view first. Then, a motion-compensated warping operation is used to transform the latent representation of the previous frame of the left viewpoint According to sports information Compensate to get the potential representation of the left viewpoint of the current frame as a reference frame.
[0056] 103. Input the quality map, the current frame, the first reconstructed frame and the second reconstructed frame into a pre-trained binocular video compression model to reconstruct the current frame, obtain a reconstructed frame of the current frame, and store it.
[0057] It should be noted that the binocular video compression model includes: a feature extraction module, an estimation module, a joint compression module, a compensation fusion module, a residual compression module and an image reconstruction module; wherein the estimation module includes a motion estimation module and a disparity estimation module, the compensation fusion module includes a motion compensation module, a disparity compensation module and a fusion module, and the residual compression module is a variable rate motion disparity joint compression network (Variable-rate JointCompression module) based on conditional entropy coding. Figure 2 shown.
[0058] It can be understood that the step of inputting the quality map, the current frame, the first reconstructed frame and the second reconstructed frame into a pre-trained binocular video compression model to reconstruct the current frame to obtain the reconstructed frame of the current frame includes:
[0059] Inputting the current frame, the first reconstructed frame and the second reconstructed frame into the feature extraction module to generate corresponding latent representations, wherein the latent representations include the current frame latent representation, the first reconstructed frame latent representation and the second reconstructed frame latent representation;
[0060] Inputting the current frame potential representation, the first reconstructed frame potential representation and the second reconstructed frame potential representation into the estimation model to perform information estimation to obtain motion information and disparity information of the current frame;
[0061] Inputting the motion information and the disparity information into the joint compression module for joint compression encoding and decoding to obtain reconstructed motion information and reconstructed disparity information of the current frame;
[0062] Inputting the reconstructed motion information, the reconstructed disparity information, the first reconstructed frame potential representation and the second reconstructed frame potential representation into the compensation fusion module for compensation fusion to obtain a comprehensive prediction feature of the current frame;
[0063] The compression code rate of the residual compression module is controlled by the quality map to perform residual compression reconstruction on the potential representation of the current frame and the comprehensive prediction features, and a reconstructed frame in the pixel domain of the current frame is output through the image reconstruction module.
[0064] Here we take the processing of the left view as an example to illustrate the current frame of the left view in the pixel domain. First reconstructed frame of the previous left view frame and the second reconstructed frame of the current frame of the right view Enter the feature extraction module ( Figure 2 The Feature Extraction module in and Among them, the current frame potential representation and the first reconstructed frame latent representation As the input of the Motion Estimation module, it is used to estimate the motion information of the current frame of the left viewpoint The current frame potentially represents and the second reconstructed frame latent representation As the input of the disparity estimation module, it is used to estimate the disparity information from the right viewpoint to the left viewpoint of the current frame.
[0065] Estimated motion information and parallax information The image is then fed into the variable-rate joint compression network (Variable-rate Joint Compression module) based on conditional entropy coding for joint compression. After the joint compression coding transmission, the decoder obtains the reconstructed motion information. and reconstruct disparity information Reconstruction of sports information The potential representation of the first reconstructed frame with the previous frame of the left view As an input to the Motion Compensation module, the prediction features of the current frame of the left view are generated. Similarly, the disparity information is reconstructed The second reconstructed frame potential representation of the right view The two prediction features are input into the disparity compensation module to generate the prediction features of the current frame of the left viewpoint. The two prediction features are fused through the fusion network to generate the comprehensive prediction features of the current frame of the left viewpoint. The current frame potential representation of the left view With prediction features Perform residual calculation and generate residual information Residual information Then enter the variable-rate residual compression module based on conditional entropy coding (Variable-rate ResidualCompression module), and obtain the reconstructed residual information after compression and reconstruction Residual information after reconstruction The predicted features of the current frame with the left view Perform residual compensation and finally reconstruct the reconstruction features of the current frame of the left viewpoint Rebuild Features Then input the image reconstruction module (Image Reconstruction) to generate a reconstructed frame in the pixel domain of the current frame of the left viewpoint and will reconstruct the frame Store into the frame buffer.
[0066] In another feasible implementation manner, the step of inputting the motion information and the disparity information into the joint compression module for joint compression encoding and decoding to obtain the reconstructed motion information and the reconstructed disparity information of the current frame includes:
[0067] Inputting the motion information and the disparity information into an encoder of the joint compression module to generate a joint latent representation;
[0068] quantizing the joint potential representation, and inputting the quantized reconstructed joint potential representation into the super-a priori encoder of the joint compression module for processing to obtain super-a priori information;
[0069] Inputting the super-prior information and context information into the entropy modeling module of the joint compression module for encoding to obtain distribution parameters of the reconstructed joint potential representation, wherein the context information is a sequence of reconstructed joint potential representations of each previous frame;
[0070] Decoding is performed based on the distribution parameters, the reconstructed joint potential representation and the context information to obtain reconstructed motion information and reconstructed disparity information of the current frame.
[0071] It can be understood that the inputting of the super-prior information and the context information into the entropy modeling module of the joint compression module for encoding to obtain the distribution parameters for reconstructing the joint potential representation includes:
[0072] generating a content-aware tag based on the hyper-prior information and the contextual information;
[0073] Generate fusion information based on the reconstructed joint latent representation and the content-aware label fusion;
[0074] The content-aware label and the reconstructed joint latent representation are subjected to bidirectional interaction and cross-attention fusion to obtain distribution parameters of the reconstructed joint latent representation.
[0075] The generating fusion information based on the reconstructed joint potential representation and the content-aware label fusion includes:
[0076] The sum of the dot product operations of the reconstructed joint potential representation and the content-aware label and the occlusion label is calculated to obtain fusion information.
[0077] like Figure 3-4 As shown, the motion information V t and disparity information D t Input the encoder module together to extract features and obtain joint potential representation Joint Latent Representation The reconstructed latent representation obtained after quantization (Q) Input the hyperprior encoder to get the hyperprior information The other side buffer will extract the previously reconstructed joint latent representation Generate context information s through ConvLSTM-based ContextGenerator t . Super Prior Information and context information t The joint context Transformer Entropy Model module (see the detailed structure of the Joint-context Transformer Entropy Model module) is input together. Figure 4 , to estimate the joint latent representation The distribution parameter (μ t ,σ t ) thus represents the joint potential Perform encoding and decoding.
[0078] At the decoding end, the motion information decoder and the disparity information decoder will jointly represent the latent and context information t As input, the reconstructed motion information is recovered and the reconstructed disparity information The method will be used to calculate the motion information V t and reconstructed motion information between, and the disparity information D t and the reconstructed disparity information A constraint is made between them. For example, the method in Formula 5 uses context information for conditional entropy coding and effectively reduces inter-frame and intra-frame redundancy by jointly compressing motion information and disparity information, thereby effectively improving rate-distortion performance.
[0079] Furthermore, the super prior information After the hyper-prior decoder (HPD) and context information s t Input the prior fusion module together to generate content-aware label c v , as shown in formula (1), the motion disparity is jointly represented as the potential and content-aware label c v Fusion generates fusion information u t , and in the subsequent self-attention operation of the Swin Transformer module, the content-aware label c v Joint latent representation of motion disparity Achieve two-way interaction to dynamically update the prior information. The potential representation after cross-attention fusion is finally estimated through the pre-trained EPM module to estimate the joint potential representation of motion parallax. The probability distribution parameter (μ t ,σ t ). Joint latent representation of motion-disparity and content-aware label c v The fusion process can be expressed as:
[0080]
[0081] Among them, u t To finally fuse the information, is the joint latent representation of motion disparity, ⊙ is the dot product operation, and m1 is the occlusion label.
[0082] In another feasible implementation, in order to improve the consistency of binocular video compression in the time dimension, the present invention introduces a time consistency loss function, which includes motion consistency loss, parallax consistency loss and inter-frame smoothness loss. That is, after inputting the quality map, the current frame, the first reconstructed frame and the second reconstructed frame into a pre-trained binocular video compression model to reconstruct the current frame and obtain the reconstructed frame of the current frame, it also includes:
[0083] The temporal consistency loss is introduced to optimize the loss of motion information and visual information of each frame in the binocular video, and to optimize the smoothness loss between frames.
[0084] The introduction of the temporal consistency loss optimizes the loss of motion information and visual information of each frame in the binocular video, and optimizes the smoothness loss between frames, including:
[0085] Calculate the norm of the motion information of the current frame of the first viewpoint and the motion information of the previous frame and the sum of the norm of the motion information of the current frame of the second viewpoint and the motion information of the previous frame, and optimize the motion information loss based on the sum of the norms;
[0086] It should be noted that the optimization of motion information loss based on the sum of norms is used to constrain the optical flow changes between adjacent frames in the time dimension and ensure the smooth transition of motion information in the time dimension. Specifically, as shown in formula (2), is the current frame motion information of the left viewpoint, is the motion information of the left viewpoint t-1 frame, is the current frame motion information of the right viewpoint, is the motion information of the right viewpoint t-1 frame, and ||||1 is the L1 norm calculation. By calculating the L1 norm of the optical flow difference between the left view and the right view between adjacent frames, the discontinuity of motion information is reduced.
[0087] Calculate the norm of the visual information of the current frame of the first viewpoint and the visual information of the previous frame after the warping operation and the sum of the norms of the visual information of the current frame of the second viewpoint and the visual information of the previous frame after the warping operation, and optimize the visual information loss based on the sum of the norms;
[0088] It should be noted that the optimization of visual information loss based on the sum of norms is used to constrain the consistency of disparity information between adjacent frames in the time dimension. Specifically, as shown in formula (3), is the disparity information of the current frame of the left viewpoint, is the disparity information of the left viewpoint t-1 frame, is the current frame motion information of the left viewpoint, is the disparity information of the current frame of the right viewpoint, is the disparity information of the right viewpoint t-1 frame, is the current frame motion information of the right viewpoint, is the warping operation, and ||||1 is the L1 norm calculation. The difference between the current frame disparity and the predicted disparity obtained by warping the previous frame disparity and the current frame motion information is taken as the optimization target, and the L1 norm is calculated for the left view and the right view respectively to ensure the smoothness of the disparity information in the time series.
[0089] Calculate the norm of the reconstructed potential representation of the current frame of the first viewpoint and the reconstructed potential representation of the previous frame after the warping operation and the sum of the norms of the reconstructed potential representation of the current frame of the second viewpoint and the reconstructed potential representation of the previous frame after the warping operation, and optimize the inter-frame smoothness loss based on the sum of the norms.
[0090] It should be noted that the optimization of inter-frame smoothness loss based on the sum of norms is used to ensure the smoothness of the reconstructed image between temporally adjacent frames. Specifically, as shown in formula (4), is the potential representation of the reconstructed left viewpoint current frame, is the potential representation of the reconstructed left viewpoint t-1 frame, is the current frame motion information of the left viewpoint, is the potential representation of the current frame of the reconstructed right viewpoint, is the potential representation of the reconstructed right viewpoint t-1 frame, is the current frame motion information of the right viewpoint, is the warping operation, and ||||1 is the L1 norm calculation. The L1 norm is calculated by calculating the difference between the reconstructed features of the current frame and the predicted features obtained after compensation based on the reconstructed features of the previous frame and the motion information of the current frame, constraining the visual smoothness between video frames.
[0091]
[0092] Furthermore, formula (5) is the complete loss function introduced in the entire binocular video compression model. For the weighted square error term where x i and Represent the original pixel value and the reconstructed pixel value respectively, λ i is a weighting factor controlled by the quality map m. Specifically, the quality map m is generated by a quality factor and the information to be transmitted and encoded. i With the corresponding weight λ i Through a monotonically increasing function association, i.e. λ i =f(m i ). In this framework, the higher the quality factor, the larger the distortion weight λ of the corresponding pixel. i The larger the quality factor, the higher the reconstruction effect of high-quality areas. The introduction of this constraint makes it possible to control the bitrate allocation by adjusting the quality factor to achieve changes in compression quality in different areas, thereby supporting variable bitrate video compression. Specifically, by adjusting the quality factor, different bitrate ranges can be covered in a single model, which greatly saves training time and achieves flexible spatial bit allocation. At the same time, combined with end-to-end optimization, this method has significantly improved rate-distortion performance, especially in application scenarios where compression quality needs to be adjusted dynamically.
[0093]
[0094] In addition, V t Indicates the motion information of the current frame. Represents the motion information reconstruction result of the current frame, D t Indicates the disparity information of the current frame, represents the disparity information reconstruction result of the current frame, d() represents the mean square error calculation, represents the motion consistency loss (Formula (2)), represents the parallax consistency loss (Formula (3)), represents the inter-frame smoothness loss (Formula (4), λ α and λ β represents the weighting factor for the control loss constraint weight, and They respectively represent the bit rates of the prior information and super prior information that need to be encoded and transmitted.
[0095] By implementing the method provided above, using quality factors to guide a larger bit rate compression range in a single model, jointly encoding motion information and disparity information, and using super-prior information and context information for efficient entropy modeling, inter-frame and intra-frame redundancy is significantly reduced, achieving higher compression efficiency and excellent rate-distortion performance, while supporting dynamic adjustment of compression quality to cover different bit rate requirements.
[0096] Furthermore, after introducing temporal consistency loss, the inter-frame quality smoothness of the video sequence reconstruction results is significantly improved. Compared with the test results without adding temporal consistency loss, the fluctuation of PSNR between frames is significantly reduced, avoiding large fluctuations in visual quality, thereby improving the overall visual effect.
[0097] The binocular video compression method in the embodiment of the present application is described above. The binocular video compression device in the embodiment of the present application is described below. Figure 5 In the embodiment of the present application, one embodiment of the binocular video compression device includes:
[0098] An acquisition module 510 is used to acquire a current frame of a first viewpoint, a first reconstructed frame of a previous frame of the first viewpoint, and a second reconstructed frame of a current frame of a second viewpoint in a binocular video, wherein the first viewpoint includes a left viewpoint and a right viewpoint;
[0099] A determination module 520, configured to determine a target quality factor of the binocular video according to a correspondence between a bit rate and a quality factor, and generate a quality map based on the target quality factor;
[0100] The compression module 530 is used to input the quality map, the current frame, the first reconstructed frame and the second reconstructed frame into a pre-trained binocular video compression model to reconstruct the current frame, obtain the reconstructed frame of the current frame, and store it.
[0101] Optionally, the binocular video compression model includes: a feature extraction module, an estimation module, a joint compression module, a compensation fusion module, a residual compression module and an image reconstruction module;
[0102] The compression module 530 is specifically used for:
[0103] Inputting the current frame, the first reconstructed frame and the second reconstructed frame into the feature extraction module to generate corresponding latent representations, wherein the latent representations include the current frame latent representation, the first reconstructed frame latent representation and the second reconstructed frame latent representation;
[0104] Inputting the current frame potential representation, the first reconstructed frame potential representation and the second reconstructed frame potential representation into the estimation model to perform information estimation to obtain motion information and disparity information of the current frame;
[0105] Inputting the motion information and the disparity information into the joint compression module for joint compression encoding and decoding to obtain reconstructed motion information and reconstructed disparity information of the current frame;
[0106] Inputting the reconstructed motion information, the reconstructed disparity information, the first reconstructed frame potential representation and the second reconstructed frame potential representation into the compensation fusion module for compensation fusion to obtain a comprehensive prediction feature of the current frame;
[0107] The compression code rate of the residual compression module is controlled by the quality map to perform residual compression reconstruction on the potential representation of the current frame and the comprehensive prediction features, and a reconstructed frame in the pixel domain of the current frame is output through the image reconstruction module.
[0108] Optionally, the compression module 530 is specifically used for:
[0109] Inputting the motion information and the disparity information into an encoder of the joint compression module to generate a joint latent representation;
[0110] quantizing the joint potential representation, and inputting the quantized reconstructed joint potential representation into the super-a priori encoder of the joint compression module for processing to obtain super-a priori information;
[0111] Inputting the super-prior information and context information into the entropy modeling module of the joint compression module for encoding to obtain distribution parameters of the reconstructed joint potential representation, wherein the context information is a reconstructed joint potential representation sequence of each previous frame;
[0112] Decoding is performed based on the distribution parameters, the reconstructed joint potential representation and the context information to obtain reconstructed motion information and reconstructed disparity information of the current frame.
[0113] Optionally, the compression module 530 is specifically used for:
[0114] generating a content-aware tag based on the hyper-prior information and the contextual information;
[0115] Generate fusion information based on the reconstructed joint latent representation and the content-aware label fusion;
[0116] The content-aware label and the reconstructed joint latent representation are subjected to bidirectional interaction and cross-attention fusion to obtain distribution parameters of the reconstructed joint latent representation.
[0117] Optionally, the compression module 530 is specifically used for:
[0118] The sum of the dot product operations of the reconstructed joint potential representation and the content-aware label and the occlusion label is calculated to obtain fusion information.
[0119] Optionally, the device further includes a loss optimization module 540, which is used to:
[0120] The temporal consistency loss is introduced to optimize the loss of motion information and visual information of each frame in the binocular video, and to optimize the smoothness loss between frames.
[0121] Optionally, the loss optimization module 540 is specifically used to:
[0122] Calculate the norm of the motion information of the current frame of the first viewpoint and the motion information of the previous frame and the sum of the norm of the motion information of the current frame of the second viewpoint and the motion information of the previous frame, and optimize the motion information loss based on the sum of the norms;
[0123] Calculate the norm of the visual information of the current frame of the first viewpoint and the visual information of the previous frame after the warping operation and the sum of the norms of the visual information of the current frame of the second viewpoint and the visual information of the previous frame after the warping operation, and optimize the visual information loss based on the sum of the norms;
[0124] Calculate the norm of the reconstructed potential representation of the current frame of the first viewpoint and the reconstructed potential representation of the previous frame after the warping operation and the sum of the norms of the reconstructed potential representation of the current frame of the second viewpoint and the reconstructed potential representation of the previous frame after the warping operation, and optimize the inter-frame smoothness loss based on the sum of the norms.
[0125] In an embodiment of the present application, by obtaining the current frame of the first viewpoint in the binocular video, the first reconstructed frame of the previous frame of the first viewpoint, and the second reconstructed frame of the current frame of the second viewpoint, wherein the first viewpoint includes the left viewpoint and the right viewpoint; according to the correspondence between the bit rate and the quality factor, the target quality factor of the binocular video is determined, and a quality map is generated based on the target quality factor; the quality map, the current frame, the first reconstructed frame, and the second reconstructed frame are input into a pre-trained binocular video compression model to reconstruct the current frame, obtain the reconstructed frame of the current frame, and store it. The present application generates a quality map through the quality factor, thereby guiding the transmission information to be compressed at different bit rates, covering a larger bit rate compression range in a single model, thereby significantly reducing training time and resource consumption.
[0126] above Figure 5 The binocular video compression device in the embodiment of the present application is described in detail from the perspective of modular functional entities, and the electronic device in the embodiment of the present application is described in detail from the perspective of hardware processing.
[0127] See also Figure 6As shown, the electronic device includes a processor 600, a memory 601 and a display 602, wherein the display 602 is used to provide a human-computer interaction interface; the memory 601 stores machine executable instructions that can be executed by the processor 600, and the processor 600 executes the machine executable instructions to implement the above-mentioned binocular video compression method.
[0128] Further, Figure 6 The electronic device shown further includes a bus 603 and a communication interface 606 , and the processor 600 , the communication interface 606 and the memory 601 are connected via the bus 603 .
[0129] The memory 601 may include a high-speed random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), for example, at least one disk storage. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 606 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used. The bus 603 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0130] The processor 600 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 600. The above processor 600 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present disclosure can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present disclosure can be directly embodied as a hardware decoding processor to be executed, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 601 , and the processor 600 reads the information in the memory 601 and completes the method steps of the above-mentioned embodiment in combination with its hardware.
[0131] The present application also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer executes the steps of the binocular video compression method.
[0132] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0133] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., and other media that can store program codes.
[0134] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A binocular video compression method, characterized in that: The method comprises: Acquire a current frame of a first viewpoint in a binocular video, a first reconstructed frame of a previous frame of the first viewpoint, and a second reconstructed frame of a current frame of a second viewpoint, wherein the first viewpoint includes a left viewpoint and a right viewpoint; Determine a target quality factor of the binocular video according to a corresponding relationship between a bit rate and a quality factor, and generate a quality map based on the target quality factor; The quality map, the current frame, the first reconstructed frame and the second reconstructed frame are input into a pre-trained binocular video compression model to reconstruct the current frame, obtain a reconstructed frame of the current frame, and store it.
2. The binocular video compression method according to claim 1, characterized in that: The binocular video compression model includes: a feature extraction module, an estimation module, a joint compression module, a compensation fusion module, a residual compression module and an image reconstruction module; The step of inputting the quality map, the current frame, the first reconstructed frame and the second reconstructed frame into a pre-trained binocular video compression model to reconstruct the current frame to obtain the reconstructed frame of the current frame includes: Inputting the current frame, the first reconstructed frame and the second reconstructed frame into the feature extraction module to generate corresponding latent representations, wherein the latent representations include the current frame latent representation, the first reconstructed frame latent representation and the second reconstructed frame latent representation; Inputting the current frame potential representation, the first reconstructed frame potential representation and the second reconstructed frame potential representation into the estimation model to perform information estimation to obtain motion information and disparity information of the current frame; Inputting the motion information and the disparity information into the joint compression module for joint compression encoding and decoding to obtain reconstructed motion information and reconstructed disparity information of the current frame; Inputting the reconstructed motion information, the reconstructed disparity information, the first reconstructed frame potential representation and the second reconstructed frame potential representation into the compensation fusion module for compensation fusion to obtain a comprehensive prediction feature of the current frame; The compression code rate of the residual compression module is controlled by the quality map to perform residual compression reconstruction on the potential representation of the current frame and the comprehensive prediction features, and a reconstructed frame in the pixel domain of the current frame is output through the image reconstruction module.
3. The binocular video compression method according to claim 2, characterized in that: The step of inputting the motion information and the disparity information into the joint compression module for joint compression encoding and decoding to obtain the reconstructed motion information and the reconstructed disparity information of the current frame includes: Inputting the motion information and the disparity information into an encoder of the joint compression module to generate a joint latent representation; quantizing the joint potential representation, and inputting the quantized reconstructed joint potential representation into the super-a priori encoder of the joint compression module for processing to obtain super-a priori information; Inputting the super-prior information and context information into the entropy modeling module of the joint compression module for encoding to obtain distribution parameters of the reconstructed joint potential representation, wherein the context information is a reconstructed joint potential representation sequence of each previous frame; Decoding is performed based on the distribution parameters, the reconstructed joint potential representation and the context information to obtain reconstructed motion information and reconstructed disparity information of the current frame.
4. The binocular video compression method according to claim 3, characterized in that: The step of inputting the super-prior information and the context information into the entropy modeling module of the joint compression module for encoding to obtain distribution parameters for reconstructing the joint potential representation includes: generating a content-aware tag based on the hyper-prior information and the contextual information; Generate fusion information based on the reconstructed joint latent representation and the content-aware label fusion; The content-aware label and the reconstructed joint latent representation are subjected to bidirectional interaction and cross-attention fusion to obtain distribution parameters of the reconstructed joint latent representation.
5. The binocular video compression method according to claim 4, characterized in that: The generating fusion information based on the reconstructed joint potential representation and the content-aware label fusion includes: The sum of the dot product operations of the reconstructed joint potential representation and the content-aware label and the occlusion label is calculated to obtain fusion information.
6. The binocular video compression method according to any one of claims 1 to 5, characterized in that: After inputting the quality map, the current frame, the first reconstructed frame and the second reconstructed frame into a pre-trained binocular video compression model to reconstruct the current frame and obtain the reconstructed frame of the current frame, the method further includes: The temporal consistency loss is introduced to optimize the loss of motion information and visual information of each frame in the binocular video, and to optimize the smoothness loss between frames.
7. The binocular video compression method according to claim 6, characterized in that: The introduction of the temporal consistency loss optimizes the loss of motion information and visual information of each frame in the binocular video, and optimizes the smoothness loss between frames, including: Calculate the norm of the motion information of the current frame of the first viewpoint and the motion information of the previous frame and the sum of the norm of the motion information of the current frame of the second viewpoint and the motion information of the previous frame, and optimize the motion information loss based on the sum of the norms; Calculate the norm of the visual information of the current frame of the first viewpoint and the visual information of the previous frame after the warping operation and the sum of the norms of the visual information of the current frame of the second viewpoint and the visual information of the previous frame after the warping operation, and optimize the visual information loss based on the sum of the norms; Calculate the norm of the reconstructed potential representation of the current frame of the first viewpoint and the reconstructed potential representation of the previous frame after the warping operation and the sum of the norms of the reconstructed potential representation of the current frame of the second viewpoint and the reconstructed potential representation of the previous frame after the warping operation, and optimize the inter-frame smoothness loss based on the sum of the norms.
8. A binocular video compression device, characterized in that: The device comprises: An acquisition module, used to acquire a current frame of a first viewpoint, a first reconstructed frame of a previous frame of the first viewpoint, and a second reconstructed frame of a current frame of a second viewpoint in a binocular video, wherein the first viewpoint includes a left viewpoint and a right viewpoint; A determination module, configured to determine a target quality factor of the binocular video according to a correspondence between a bit rate and a quality factor, and generate a quality map based on the target quality factor; A compression module is used to input the quality map, the current frame, the first reconstructed frame and the second reconstructed frame into a pre-trained binocular video compression model to reconstruct the current frame, obtain a reconstructed frame of the current frame, and store it.
9. An electronic device, characterized in that: The electronic device comprises: a memory, at least one processor and a display, wherein instructions are stored in the memory; The at least one processor calls the instructions in the memory so that the electronic device executes the binocular video compression method according to any one of claims 1 to 7.
10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the binocular video compression method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Context-guided converter entropy coding video compression device and method
CN120343275A