A video compression method and system considering long-distance time series information
The proposed video compression framework addresses the limitation of short-distance time sequence focus by integrating long-distance temporal context, enhancing motion and context prediction for improved video compression efficiency.
Patent Information
- Application Number
- CN202211706230.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-29
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-12-29
AI Technical Summary
The existing end-to-end video compression methods fail to effectively utilize long-distance timing information, limiting the performance optimization of the codec.
A video compression method and system that considers long-distance timing information is constructed. By iteratively recycling the timing prior information, it assists in the compression and decompression of motion information and context information, including motion estimation, motion information encoding and decoding, motion compensation, context encoding and decoding, etc., the timing prior initialization and update modules are used for efficient use.
It realizes efficient use of long-distance timing information, improves the accuracy and efficiency of video compression, enhances the prediction ability of motion information and context information, and improves the overall performance of the codec.
Smart Images

Figure CN116033169B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of end-to-end optimized video compression, and relates to a video compression method and system, in particular to a video compression method and system that takes long-distance temporal information into account. Background Art
[0002] Video compression is a fundamental technology in the fields of signal processing and computer vision. Nowadays, the public's increasing demand for video content with higher resolutions, frame rates, and dynamic ranges has made the need for efficient video compression methods even more urgent. The widely adopted traditional video codecs have achieved quite efficient rate-distortion performance with continuous improvements by global experts in recent decades. Nevertheless, they still encounter limitations in terms of optimization schemes and objectives. Traditional video codecs are designed in the form of a hybrid coding framework and rely heavily on various manually designed techniques to improve compression efficiency. The sub-modules are designed and optimized separately but cannot be optimized simultaneously, thus limiting the compression potential. In addition, it is difficult for the hybrid coding framework to be optimized for complex objective metrics such as multi-scale structural similarity metrics and visual perception metrics. In summary, considering the uniqueness of video content and the rapid development of deep learning technology, exploring new video coding frameworks and paradigms has become an important trend in the development of future video codecs.
[0003] With the rapid development of deep learning, the potential of artificial neural networks has been further explored, and the concept of an image compression framework based on deep learning has also been formed. Since the end-to-end optimized compression method can jointly train the parameters of the entire framework, the improvement of the performance of each module will naturally promote the achievement of the final goal. The existing end-to-end optimized video compression process consists of the following basic modules and steps: motion estimation, motion compensation, context / residual compression, transformation, quantization, and entropy coding. First, the motion estimation module is used to predict the motion information between the reference frame and the current frame, and then the motion information is converted into a latent representation through a compression transformation operation, and then the quantized latent representation is compressed using entropy coding. The decoded motion information is used to align with the reference frame to generate the context. The generated context information is used to assist in the compression, decompression, and entropy estimation of the current frame. However, the current end-to-end video compression methods basically only consider short-distance temporal information and do not consider long-distance temporal information, thus limiting the performance of the entire codec.
[0004] How to construct an effective and efficient video compression scheme that utilizes long-distance temporal information is a challenge. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides a video compression method and system that takes long-distance temporal information into account.
[0006] The technical solution adopted by the method of the present invention is as follows: A video compression method considering long-distance temporal information, which inputs a video sequence or video features into a video compression model. During encoding, the video compression model generates a corresponding bitstream, and during decoding, the video compression model reconstructs the corresponding video sequence or features from the bitstream;
[0007] The video compression model includes a motion estimation module, a motion information encoding and decoding module, a prior encoding network, a motion information entropy estimation module, a motion compensation module, a context encoder, a context decoder, a context information entropy estimation module, and a temporal prior initialization, supplementation, and update module;
[0008] The compression process is carried out in an iterative loop. The previously decoded video frame is taken as the reference frame, and the currently encoded video frame is the current frame. First, the temporal prior initialization, supplementation, and update module initializes or supplements the current temporal prior using the reference frame, and then the motion estimation module is used to generate the motion information between the reference frame and the current frame. Then, the temporal prior is used to assist the motion information encoding and decoding module to compress and decompress the motion information, generate the corresponding bitstream, and recover the motion information from the bitstream. Here, the motion information entropy estimation module is used to generate Gaussian parameters, estimate the probabilities of each element of the motion information, and assist in compression and decompression. Then, the decoded motion information and the reference frame are used to complete motion compensation through the motion compensation module to obtain the context information of the current frame. At the same time, the temporal prior is also updated through the motion information. The current frame information and the context information are fused, and the temporal prior is used to assist the context encoder to compress the fused information. Finally, the context decoder is used to reconstruct the current video frame from the bitstream. At the same time, the context information entropy estimation module is used to generate Gaussian parameters, estimate the probabilities of each element, and assist in compression and decompression. After the current reference frame completes the encoding and decoding process, it will be used as the reference frame for the next compression cycle to assist in the compression of the next video frame.
[0009] The technical solution adopted by the system of the present invention is as follows: A video compression system considering long-distance temporal information, including:
[0010] One or more processors;
[0011] A storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the video compression method considering long-distance temporal information described above.
[0012] The present invention has the following advantages:
[0013] 1) The present invention constructs an end-to-end optimized video compression framework, generates the context of the current compression object through motion compensation, and then assists in the compression of the current frame according to the context information.
[0014] 2) The present invention proposes to use temporal prior to achieve efficient utilization of long - distance temporal information. In a group of pictures, the temporal prior is initialized by the first reference frame and iteratively updated during the compression process to always maintain spatial correlation with the current compressed frame. The temporal prior is used to assist in predicting motion information and predicting the Gaussian distribution parameters of the current frame. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a rough flowchart of data acquisition and compression in an embodiment of the present invention.
[0016] Figure 2 It is a flowchart of inter - frame compression in an embodiment of the present invention.
[0017] Figure 3 It is a flowchart of temporal prior initialization, supplementation, and update operations in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0018] To facilitate understanding and implementation of the present invention by those of ordinary skill in the art, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments described herein are only for the purpose of illustrating and explaining the present invention and are not intended to limit the present invention.
[0019] Currently, most deep - learning - based video compression methods focus on the utilization of short - distance temporal information, such as motion compensation and multi - frame enhanced recovery. Short - distance temporal information refers to the temporal information contained in a fixed number of decoded video frames (usually one to four frames). Such temporal information can directly assist in the reconstruction and recovery of the current frame through motion compensation, and the information modeling method is relatively simple and direct. However, in video compression tasks, long - distance temporal information also has important practical value. Long - distance temporal information can be used to assist in predicting motion information and entropy estimation of context information, making the parameter prediction process more accurate. Moreover, accurate and effective long - distance temporal information can also assist in the reconstruction of the current compressed frame. The present invention addresses this design shortcoming of current end - to - end optimized video compression methods and provides a video compression method and system that consider long - distance temporal information.
[0020] This method proposes to construct a temporal prior to preserve and utilize the global temporal information generated during the compression process. After the intra - frame (I - frame) in a group of pictures (GOP) is decoded, the temporal prior initialization operation is used to generate the temporal prior information of the current group of pictures using the I - frame. During the subsequent inter - frame (Predictive frame, P - frame) compression process, the temporal prior is continuously updated according to the inter - frame motion information and assists in predicting the Gaussian distribution parameters and recovering the transformation of the video frame during the compression of motion information and context information.
[0021] See Figure 1 Figure 1 , a video compression method considering long - distance temporal information provided by the present invention inputs a video sequence or video features into a video compression model. During encoding, the video compression model of this embodiment generates a corresponding bitstream, and during decoding, the video compression model of this embodiment reconstructs a corresponding video sequence or features from the bitstream;
[0022] The video compression model of this embodiment includes a motion estimation module, a motion information encoding and decoding module, a prior encoding network, a motion information entropy estimation module, a motion compensation module, a context encoder, a context decoder, a context information entropy estimation module, and a temporal prior initialization, supplementation and update module;
[0023] The compression process is carried out in an iterative loop. The previously decoded video frame is taken as a reference frame, and the currently encoded video frame is the current frame. First, the temporal prior initialization, supplementation and update module of this embodiment is used to initialize or supplement the current temporal prior using the reference frame, and then the motion estimation module of this embodiment is used to generate the motion information between the reference frame and the current frame. Then, the temporal prior is used to assist the motion information encoding and decoding module to compress and decompress the motion information, generate a corresponding bitstream, and recover the motion information from the bitstream. During this process, the motion information entropy estimation module is used to generate Gaussian parameters, estimate the probabilities of each element of the motion information, and assist in compression and decompression. Then, the motion compensation is completed using the decoded motion information and the reference frame through the motion compensation module of this embodiment to obtain the context information of the current frame, and at the same time, the temporal prior is also updated through the motion information. The current frame information and the context information are fused, and the temporal prior is used to assist the context encoder of this embodiment to compress the fused information. Finally, the context decoder of this embodiment is used to reconstruct the current video frame from the bitstream. During this process, the context information entropy estimation module is used to generate Gaussian parameters to estimate probabilities and assist in compression and decompression. After the current reference frame completes the encoding and decoding process, it will be used as the reference frame for the next compression cycle to assist in the compression of the next video frame.
[0024] The motion estimation module of this embodiment is composed of six cascaded basic feature extraction modules, and the basic feature extraction module is composed of five consecutive convolutional layers; the convolutional kernel size of the convolutional layer is 7, the stride is 1, and an activation layer is added after each convolution for feature non - linear transformation; the output of the final motion estimation module is an optical flow field with 2 channels and the same resolution as the original Figure 1 one.
[0025] The motion information encoding and decoding module of this embodiment is composed of a motion information encoder and a motion information decoder;
[0026] The motion information encoder of this embodiment consists of four cascaded feature encoding modules; the feature encoding module includes a convolutional layer, a residual connection layer, and a downsampling module; among them, the convolutional kernel size of the convolutional layer is 5 and the stride is 1; the convolutional kernel sizes in the residual connection layer are 1, 3, and 1 in sequence, and the stride is 1; the downsampling module is a convolutional layer with a convolutional kernel size of 5 and a stride of 2.
[0027] The motion information decoder of this embodiment consists of four cascaded feature decoding modules; the feature encoding module includes a convolutional layer, a residual connection layer, and an upsampling module; among them, the convolutional kernel size of the convolutional layer is 5 and the stride is 1; the convolutional kernel sizes in the residual connection layer are 1, 3, and 1 in sequence, and the stride is 1; the upsampling module is a convolutional layer plus a spatial mixing operation, and the convolutional kernel size of its convolutional layer is 3 and the stride is 1.
[0028] The motion information entropy estimation module of this embodiment is a masked checkerboard convolutional layer with a convolutional kernel size of 5 and a stride of 1, which is used to predict the Gaussian parameters of the motion information elements to be compressed.
[0029] The motion compensation module of this embodiment is a spatial position warping operation based on the optical flow field, and pixel resampling is performed based on the pixel offset values in the optical flow field.
[0030] The context encoder and context decoder of this embodiment; their network hierarchical structures are basically the same as those of the motion information encoder and motion information decoder, except that the input of the encoder is modified from the optical flow field to context feature information, and the output of the decoder is modified to the reconstructed current frame information.
[0031] The context information entropy estimation module of this embodiment is a masked checkerboard convolutional layer with a convolutional kernel size of 5 and a stride of 1, which is used to predict the Gaussian parameters of each element to be compressed.
[0032] For the network structure of the temporal prior initialization, supplementation, and update module of this embodiment, please refer to Figure 3 . It includes a temporal prior initialization sub-module, a temporal prior supplementation sub-module, and a temporal prior update sub-module;
[0033] In the temporal prior initialization sub-module of this embodiment, the reference frame passes through a convolutional layer, an activation function layer, a convolutional layer, an activation function layer, and a convolutional layer in sequence, and then outputs the temporal prior.
[0034] In the temporal prior supplementation sub-module of this embodiment, the reference frame features pass through a convolutional layer and an activation function layer in sequence, and after being fused with the current temporal prior, they are input into a convolutional layer, an activation function layer, and a convolutional layer in sequence, and then output the temporal prior.
[0035] In the temporal prior update sub-module of this embodiment, after the current temporal prior is fused with the motion information, it is sequentially input into the motion compensation layer, convolutional layer, activation function layer, and convolutional layer, and then the temporal prior is output.
[0036] This embodiment uses temporal prior storage and utilization of the global temporal information in the encoding and decoding process; specifically including:
[0037] During the encoding process, the first frame of the video cannot refer to the information of other frames, so it is called an intra-frame; all subsequent video frames can be obtained by prediction with reference to the previously decoded video frames, so they are called predicted frames; when encoding the first predicted frame, the intra-frame data is used for the initialization of the temporal prior.
[0038] When encoding the motion information, a prior encoding network including two cascaded convolutional layers is used to process the temporal prior to obtain the corresponding encoded prior, where the convolutional kernel size of the convolutional layer is 5 and the stride is 2, and an activation layer is added after each convolution; then the encoded prior is sent into the motion information entropy estimation module to assist in the prediction of Gaussian parameters; the Gaussian parameters are used for the probability prediction of each element to be encoded, and the predicted probability is used to calculate the information entropy.
[0039] After obtaining the decoded motion information, an explicit feature resampling strategy based on the optical flow field is used to process the current temporal prior to increase its spatial correlation with the current video frame.
[0040] When encoding subsequent predicted frames, the reference frame data is used to update the temporal prior in the current loop; the updated temporal prior is used for context compression, including context encoding, context decoding, and context information entropy estimation module.
[0041] In the compression process of this embodiment, the flowchart of inter-frame compression is as Figure 2 shown, and the specific implementation steps are as follows:
[0042] Step 1: Initialize the temporal prior. Use the reference frame as the input to extract features for the initialization of the temporal prior. As the encoding and decoding progress, the temporal information contained in the temporal prior will become richer and richer. See the prior initialization step in Figure 3 .
[0043] Step 2: Construct the motion estimation module. In order to enable the motion estimation module of the present invention to be trained end-to-end synchronously with the overall model, the motion information is modeled as an optical flow field, and an optical flow field prediction network is used to estimate the motion information between the reference frame and the current frame. The optical flow field information is represented as the pixel offsets in the horizontal and vertical directions.
[0044] Step 3: Construction of the motion information encoding and decoding module. The present invention uses multiple cascaded convolutional operation layers and generalized split normalization to construct the motion information encoding and decoding module. This non-linear transformation encoding method completes the mutual conversion between data and latent expressions through an encoder-decoder. The goal of compression is to reduce the information entropy of the latent expression. Finally, the motion information entropy estimation module assists in generating a compressed bitstream by generating probability estimation results for corresponding elements.
[0045] The input of the motion information compression model of the present invention is an optical flow field, and the output of the encoder is a latent expression variable that conforms to a Gaussian distribution. This output result needs to be input into the motion information entropy estimation module for further feature extraction and data distribution prediction. The latent expression variable is quantized through the assistance of Gaussian distribution data and written into the bitstream for storage. During decoding, the latent expression variable is decoded from the bitstream, and the optical flow field information is restored through the motion information decoder.
[0046] Step 4: Construction of the motion compensation module and update of the temporal prior. The motion information obtained by decoding is used to process the pixel information of the reference frame to align the reference frame information with the content of the current frame. At the same time, the current temporal prior is updated using the motion information to keep the temporal prior highly spatially correlated with the current frame. The steps for updating the temporal prior are shown in Figure 3 .
[0047] Step 5: Compression of the current frame based on context information. The model construction is similar to the motion information compression model. The context information encoding and decoding module is used to compress the video frame information, so that the output of the encoder is a latent expression variable that conforms to a Gaussian distribution. This output result needs to be input into the context information entropy estimation module for further feature extraction and data distribution prediction. The latent expression variable is quantized through the assistance of Gaussian parameters predicted by the entropy estimation module and written into the bitstream for storage. During decoding, the latent expression variable is decoded from the bitstream, and the current frame information is restored through the context information decoder.
[0048] Step 6: Construction of the rate-distortion loss function. The training objective of the lossy compression model is to minimize the expected length of the bitstream and the distortion of the reconstructed video frame relative to the original image, which can be summarized as a rate-distortion optimization problem:
[0049] Loss = R + λ·D;
[0050] where λ is the Lagrange multiplier, which determines the desired rate-distortion trade-off result of the model; R represents the bit rate predicted by the entropy estimation module, and D represents the distortion between the reconstructed video frame and the real video frame.
[0051] The first fraction in the loss function, i.e., the rate term, corresponds to the cross-entropy between the marginal distribution of the latent representation and the estimation result of the learned entropy model. Minimizing the cross-entropy makes these two distributions as identical as possible. The second fraction in the loss function, the distortion term, corresponds to the likelihood of the approximate form of the original image and the reconstruction result.
[0052] Step 7: End-to-end training of the model. The present invention applies the gradient descent method to the loss function in Step 6 to optimize the overall model, thereby completing the end-to-end training of the overall compression framework. During training, the Lagrange multiplier λ is selected to construct the loss function, and further an end-to-end optimized video compression model is constructed.
[0053] It should be understood that the above description of the preferred embodiment is relatively detailed, and it should not be considered as a limitation to the protection scope of the invention patent. Under the inspiration of the present invention, those of ordinary skill in the art can also make substitutions or deformations without departing from the protection scope defined by the claims of the present invention, and all fall within the protection scope of the present invention. The scope of the present invention claimed shall be subject to the appended claims.
Claims
1. A video compression method considering long-distance temporal information, characterized in that: Input a video sequence or video features into a video compression model. During encoding, the video compression model generates a corresponding bitstream, and during decoding, the video compression model reconstructs the corresponding video sequence or features from the bitstream. The video compression model includes a motion estimation module, a motion information encoding / decoding module, a prior encoding network, a motion information entropy estimation module, a motion compensation module, a context encoder, a context decoder, a context information entropy estimation module, and a temporal prior initialization, supplementation, and update module. The compression process is carried out in an iterative loop. The previously decoded video frame is taken as the reference frame, and the currently encoded video frame is the current frame. First, use the reference frame to initialize or supplement the current temporal prior through the temporal prior initialization, supplementation, and update module. Then use the motion estimation module to generate the motion information between the reference frame and the current frame. Then use the temporal prior to assist the motion information encoding / decoding module to compress and decompress the motion information through the prior encoding network, generate the corresponding bitstream, and recover the motion information from the bitstream. The motion information entropy estimation module is used to generate Gaussian parameters, estimate the probabilities of each element of the motion information, and assist in compression and decompression. Then use the decoded motion information and the reference frame to complete motion compensation through the motion compensation module to obtain the context information of the current frame. At the same time, the temporal prior is also updated through the motion information. Integrate the current frame information and the context information, and use the temporal prior to assist the context encoder to compress the integrated information through the prior encoding network. Finally, use the context decoder to reconstruct the current video frame from the bitstream. During this process, the context information entropy estimation module is used to generate the corresponding Gaussian distribution parameters, estimate the probabilities of each pixel, and assist in compression and decompression. After the current reference frame completes the encoding / decoding process, it will be used as the reference frame for the next compression cycle to assist in the compression of the next video frame.
2. The video compression method considering long-distance timing information according to claim 1, characterized in that: The motion estimation module consists of six cascaded basic feature extraction modules, and the basic feature extraction module consists of five consecutive convolutional layers. The convolutional kernel size of the convolutional layer is 7, the stride is 1, and an activation layer is added after each convolution for feature non-linear transformation. The output of the final motion estimation module is an optical flow field with 2 channels and the same resolution as the original image.
3. The video compression method considering long-distance timing information according to claim 1, characterized in that: The motion information encoding / decoding module consists of a motion information encoder and a motion information decoder. The motion information encoder consists of four cascaded feature encoding modules. The feature encoding module includes a convolutional layer, a residual connection layer, and a downsampling module. The convolutional kernel size of the convolutional layer is 5, and the stride is 1. The convolutional kernel sizes in the residual connection layer are 1, 3, 1 in sequence, and the stride is 1. The downsampling module is a convolutional layer with a convolutional kernel size of 5 and a stride of 2. The motion information decoder consists of four cascaded feature decoding modules; the feature encoding module includes a convolutional layer, a residual connection layer, and an upsampling module; the convolutional kernel size of the convolutional layer is 5, and the stride is 1; the convolutional kernel sizes in the residual connection layer are 1, 3, 1 in sequence, and the stride is 1; the upsampling module is a convolutional layer plus a spatial mixing operation, and the convolutional kernel size of its convolutional layer is 3, and the stride is 1.
4. The video compression method considering long-distance timing information according to claim 1, characterized in that: The motion information entropy estimation module and the context information entropy estimation module are both masked checkerboard convolutional layers, with a convolutional kernel size of 5 and a stride of 1.
5. The video compression method considering long-distance timing information according to claim 1, wherein: The motion compensation module is a spatial position distortion operation based on the optical flow field, and pixel resampling is performed based on the pixel offset values in the optical flow field.
6. The video compression method considering long-distance timing information according to claim 1, characterized in that: The context encoder and the context decoder; Its network hierarchical structure is basically the same as that of the motion information encoder and the motion information decoder, except that the input of the encoder is modified from the optical flow field to the context feature information, and the output of the decoder is modified to the reconstructed current frame information.
7. The video compression method considering long-distance timing information according to claim 1, characterized in that: The temporal prior initialization, supplementation, and update module includes a temporal prior initialization sub-module, a temporal prior supplementation sub-module, and a temporal prior update sub-module; In the temporal prior initialization sub-module, the reference frame passes through a convolutional layer, an activation function layer, a convolutional layer, an activation function layer, and a convolutional layer in sequence, and then outputs the temporal prior; In the temporal prior supplementation sub-module, the reference frame features pass through a convolutional layer and an activation function layer in sequence, and after being fused with the current temporal prior, they are input into a convolutional layer, an activation function layer, and a convolutional layer in sequence, and then output the temporal prior; In the temporal prior update sub-module, the current temporal prior is fused with the motion information and then input into a motion compensation layer, a convolutional layer, an activation function layer, and a convolutional layer in sequence, and then outputs the temporal prior.
8. The video compression method considering long-distance timing information according to claim 1, characterized in that: The global temporal information in the encoding and decoding process is stored and utilized by using the temporal prior; specifically including: During the encoding process, the first frame of the video cannot refer to the information of other frames, so it is called an intra-frame; all subsequent video frames can be predicted by referring to the previously decoded video frames, so they are called predicted frames; when encoding the first predicted frame, the intra-frame data is used for the initialization of the temporal prior; When encoding the motion information, a prior encoding network containing two cascaded convolutional layers is used to process the temporal prior to obtain the corresponding encoded prior, where the convolutional kernel size of the convolutional layer is 5 and the stride is 2, and an activation layer is added after each convolution; then the encoded prior is sent into the motion information entropy estimation module to assist in the prediction of Gaussian parameters; the Gaussian parameters are used for the probability prediction of each element to be encoded, and the predicted probability is used to calculate the information entropy; After obtaining the decoded motion information, an explicit feature resampling strategy based on the optical flow field is used to process the current temporal prior to increase the spatial correlation between it and the current video frame; When encoding subsequent predicted frames, the reference frame data is used to update the temporal prior in the current loop; the updated temporal prior is used for context compression, including context encoding, context decoding, and the context information entropy estimation module.
9. The video compression method considering long-distance timing information according to any one of claims 1-8, characterized in that: The video compression model is a video compression model trained by using the gradient descent method; The training is carried out in a multi-stage manner; after the training of the motion-related modules, including the motion estimation module, the motion information encoding / decoding module, and the motion compensation module, converges, the training of all modules of the overall framework is then carried out.
10. A video compression system considering long-distance timing information, characterized in that, Including: One or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the video compression method considering long-distance temporal information according to any one of claims 1 to 9.
Citation Information
Patent Citations
Variable code rate video compression method, system and device and storage medium
CN114501013A
End-to-end intelligent video coding method and device
CN115278262A