An end-to-end intelligent video encoding method and device

The end-to-end intelligent video coding method addresses the lack of temporal context exploration in existing methods by utilizing long-term temporal relationships and encoded context information, achieving better coding efficiency and reduced bit rate.

CN115278262BActive Publication Date: 2025-07-15TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210915058.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-01
Publication Date
2025-07-15
Estimated Expiration
2042-08-01

AI Technical Summary

Technical Problem

The existing end-to-end intelligent video encoding methods lack effective exploration of the video time domain context, resulting in the need to improve the encoding performance.

Method used

The end-to-end intelligent video encoding network framework is built, and by exploring the long-term timing relationship of video sequences, the global time-domain reference feature generation module and time-domain prior encoder are used to improve the accuracy of motion compensation prediction, and supplement the high-frequency information loss in the motion vector and residual coding process.

Benefits of technology

It improves video encoding performance, effectively saves coding rate, improves coding efficiency, and performs better especially in complex sports scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115278262B_ABST
    Figure CN115278262B_ABST
Patent Text Reader

Abstract

The present invention discloses an end-to-end intelligent video coding method and apparatus. The method includes: constructing an end-to-end intelligent video coding network framework composed of a feature extraction module, a global temporal reference feature generation module, a motion estimation module, a temporal prior encoder, a motion compensation module, and a reconstruction module; the feature extraction module maps the current coding frame to a feature space to obtain the current coding frame feature; obtaining a global reference feature through the global temporal reference feature generation module; estimating a motion vector between the global reference feature and the current coding frame feature through the motion estimation module; compressing the motion vector through the temporal prior encoder; based on the compressed motion vector and the global reference feature, obtaining the predicted feature of the current coding frame using the motion compensation module; compressing the residual between the predicted feature and the current coding frame feature using the temporal prior encoder; generating a reconstructed feature based on the compressed residual and the predicted feature, and obtaining the final reconstructed frame through the reconstruction module. The apparatus includes: a processor and a memory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of deep learning and video coding, and particularly to an end-to-end intelligent video coding method and apparatus. Background Art

[0002] With the introduction of high-quality videos such as ultra-high definition, high dynamic range, and high frame rate, the data volume of videos has become increasingly large, posing a huge challenge to video transmission and storage systems. To reduce the data volume of videos, the international standard organization has developed a series of video coding standards. These video coding standards usually use a variety of manually designed coding tools to remove video signal redundancy. Specifically, intra-frame prediction and inter-frame prediction techniques are used to remove spatial and temporal redundancy information respectively. Subsequently, the prediction residuals are transformed, quantized, and entropy encoded to further remove visual and statistical redundancy information in the frequency domain space. However, since these manually designed coding tools cannot be jointly optimized using the rate-distortion optimization function, it is difficult to further improve the coding performance.

[0003] Benefiting from the powerful feature extraction and non-linear expression capabilities of deep learning, end-to-end image coding methods have been widely studied in recent years and have achieved coding performance comparable to traditional image coding methods. Inspired by this, some scholars have begun to focus on researching end-to-end intelligent video coding methods. The end-to-end intelligent video coding method aims to use a neural network to implement a complete video coding framework. Lu et al. established the first end-to-end intelligent video coding framework based on a traditional hybrid video coding framework, obtained a predicted frame through a motion estimation network and a motion compensation network, and further compressed the estimated motion vectors and residuals using an end-to-end image coding method. Lin et al. designed an end-to-end intelligent video coding model based on multi-reference frame prediction, which effectively improved the accuracy of motion compensation prediction using multiple reference frames, thereby improving the efficiency of end-to-end intelligent video coding. Hu et al. proposed an adaptive resolution motion vector coding method, which adaptively selects the optimal resolution optical flow map for the motion vectors of coding blocks, improving the efficiency of motion vector coding.

[0004] However, existing methods mainly improve coding performance by modeling the short-term temporal correlation of videos. Due to the lack of effective exploration of the video temporal context, the coding performance needs to be further improved. Summary of the Invention

[0005] The present invention provides an end-to-end intelligent video coding method and apparatus. The present invention improves the prediction accuracy of motion compensation by exploring the long-term temporal relationship of video sequences, and at the same time uses the encoded temporal context information to supplement the high-frequency information loss in the motion vector and residual coding processes, as described in detail below:

[0006] An end-to-end intelligent video coding method, the method comprising:

[0007] Construct an end-to-end intelligent video coding network framework consisting of a feature extraction module, a global temporal reference feature generation module, a motion estimation module, a temporal prior encoder, a motion compensation module, and a reconstruction module;

[0008] The feature extraction module maps the current encoded frame to the feature space to obtain the current encoded frame feature, obtains the global reference feature through the global temporal reference feature generation module, and estimates the motion vector between the global reference feature and the current encoded frame feature through the motion estimation module;

[0009] Compress the motion vector through the temporal prior encoder. Based on the compressed motion vector and the global reference feature, use the motion compensation module to obtain the predicted feature of the current encoded frame. Use the temporal prior encoder to compress the residual between the predicted feature and the current encoded frame feature. Generate the reconstructed feature based on the compressed residual and the predicted feature, and obtain the final reconstructed frame through the reconstruction module.

[0010] Among them, the global temporal reference feature generation module generates the global reference feature f for the motion estimation module and the motion compensation module r and dynamically updates the temporal context state information once where the subscript t represents the current time t, and the superscript l i represents the i-th level of context state information;

[0011] The feature extraction module extracts the short-term temporal context from the reference frame which is expressed as:

[0012]

[0013] After cascading the updated temporal context state information through two convolutional layers and then fusing through two convolutional layers, the long-term temporal context is obtained The calculation is as follows:

[0014]

[0015] where Fusion(·), h0(·), h1(·), h2(·) all represent two convolutional layers, represents the concatenation of the channel dimension.

[0016] Aggregate the short-term temporal context and the long-term temporal context to generate the final global temporal reference feature f t r :

[0017]

[0018] where g(·) represents one convolutional layer.​

[0019] Furthermore, the temporal prior encoder takes the motion vector v t and the temporal reference information as inputs, and outputs the compressed motion vector including: a temporal context generator, an encoder, a decoder, and a conditional entropy codec.

[0020] Among them, the temporal context generator is used to extract the multi-level temporal context information {m1, m2, m3} in the temporal reference information , and aggregates the frequency subbands from the same level through concatenation and convolution operations to generate the multi-level temporal context information {m1, m2, m3}:

[0021]

[0022] Among them, g(·) represents a convolutional layer, represents the concatenation of the channel dimensions.

[0023] Among them, the encoder compresses the input motion vector v t into a compact latent representation y3. The encoder consists of three stacked encoding units, and each encoding unit includes: two residual units and a downsampling convolutional layer. Add the output of each encoding unit to the corresponding multi-level temporal context information generated by the temporal context generator. The calculation formula is expressed as:

[0024]

[0025] Among them, enc(·) represents an encoding unit.

[0026] Furthermore, the conditional entropy codec rounds and quantizes the latent representation y3 to generate the quantized latent variable Extracts the temporal prior from the temporal reference information using two stacked encoding units and a convolutional layer , and uses a convolutional layer to fuse the hyperprior and the autoregressive prior to generate the spatial prior. Fuse the temporal prior and the spatial prior through 3 layers of 1×1 convolution to obtain the mean μ and variance σ of the Gaussian distribution of the latent variable ;

[0027] Decodes the quantized latent variable through an inverse transform to obtain the compressed motion vector The decoder consists of three stacked decoding units, and each decoding unit includes: two residual units and an upsampling convolutional layer;

[0028] Add the output of each decoding unit to the corresponding multi-level temporal context information generated by the temporal context generator, and finally obtain the compressed motion vector.

[0029] Wherein, the method further includes: training the end-to-end intelligent video coding network framework on single videos and multi-videos respectively.

[0030] An end-to-end intelligent video coding device, the device includes: a processor and a memory, and program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to enable the device to execute any of the method steps.

[0031] The beneficial effects of the technical solution provided by the present invention are:

[0032] 1. By exploring the long-term temporal relationship of video sequences, the present invention improves the accuracy of motion compensation prediction in video coding;

[0033] 2. The present invention uses the encoded temporal context information to supplement the high-frequency information loss in the motion vector and residual coding processes, improves the coding efficiency of the motion vector and residual, and thus improves the video coding performance.

[0034] 3. Compared with the video coding standard HEVC reference software HM-16.21, the method of the present invention can effectively save the bit rate and improve the coding performance. Description of the Drawings

[0035] Figure 1 It is a flowchart of an end-to-end intelligent video coding method;

[0036] Figure 2 It is a schematic diagram of the global temporal reference feature generation module;

[0037] Figure 3 It is a schematic diagram of the temporal prior encoder compressing the motion vector;

[0038] Figure 4 It is a schematic diagram of the temporal prior encoder compressing the residual;

[0039] Figure 5 Schematic diagram of the bit consumption comparison between the proposed method and the video coding standard HEVC reference software HM-16.21. Detailed Embodiments

[0040] To make the objectives, technical solutions, and advantages of the present invention clearer, the following further describes the embodiments of the present invention in detail.

[0041] I. Constructing an end-to-end intelligent video coding network framework

[0042] The input of the end-to-end intelligent video coding network is the original video sequence, and the output is the compressed video sequence. When compressing the current coded frame, first, the feature extraction module is used to map the current coded frame to the feature space to obtain the current coded frame features. Subsequently, the global time domain reference feature generation module is used to obtain accurate global reference features. Then, the motion estimation module is used to estimate the motion vector between the global reference feature and the current coded frame feature. After that, the time domain prior encoder is used to compress the motion vector. Based on the compressed motion vector and the global reference feature, the motion compensation module is used to obtain the predicted feature of the current coded frame. Then, the time domain prior encoder is used to compress the residual between the predicted feature and the current coded frame feature. Finally, the compressed residual is cascaded and superimposed with the predicted feature to generate a reconstructed feature. After the reconstructed feature passes through the reconstruction module, the final reconstructed frame is obtained.

[0043] 2. Building a global time domain reference feature generation module

[0044] Given a reference frame And the time domain context state information As input, the global temporal reference feature generation module generates accurate global reference features f for the motion estimation module and the motion compensation module. t r , and dynamically update the temporal context state information once The subscript t indicates the current time t, and the superscript l i Represents the i-th level context state information.

[0045] First, a feature extraction module is used to extract Extracting short-term temporal context The calculation formula is expressed as:

[0046]

[0047] Where FE(·) represents the feature extraction module.

[0048] Then, long-term temporal context is generated by exploring the long-term temporal relationship of the video sequence. Specifically, first, the short-term temporal context With time domain context state information Feed it into three stacked Conv-LSTM units to update the time domain context state information

[0049]

[0050] Here, h(·) represents the stacked Conv-LSTM unit.

[0051] Subsequently, the updated temporal context state information is cascaded through two convolutional layers and then fused through two more convolutional layers to obtain the long-term temporal context. The calculation is as follows:

[0052]

[0053] Among them, Fusion(·), h0(·), h1(·), h2(·) all represent two convolutional layers. represents the concatenation of the channel dimensions.

[0054] Finally, the short-term temporal context and the long-term temporal context are aggregated to generate the final global temporal reference feature f. t r :

[0055]

[0056] Among them, g(·) represents one convolutional layer.

[0057] The global temporal reference feature generation module not only utilizes the short-term temporal context information but also fully explores the long-term temporal context information to generate the accurate global reference feature f. t r . Since f t r aggregates the reference information from the long-term temporal context of the video, the designed end-to-end intelligent video coding network can effectively improve the coding efficiency in complex motion scenarios.

[0058] III. Design a temporal prior encoder to compress motion vectors

[0059] Given the motion vector v t and the temporal reference information as inputs, the temporal prior encoder outputs the compressed motion vector whose structure mainly includes: a temporal context generator, an encoder, a decoder, and a conditional entropy codec.

[0060] The main purpose of the temporal context generator is to extract the multi-level temporal context information {m1, m2, m3} in the temporal reference information . First, the discrete wavelet transform (DWT) is used to transform the temporal reference information Decompose into four sub-bands, namely one low-frequency sub-band LL1 and three high-frequency sub-bands {HL1, LH1, HH1}. Subsequently, decompose the low-frequency sub-band LL1 into four lower-resolution sub-bands {LL2, HL2, LH2, HH2}, and similarly decompose the low-frequency sub-band LL2 into four even lower-resolution sub-bands {LL3, HL3, LH3, HH3}. Finally, the obtained multi-scale frequency sub-bands can be expressed as:

[0061]

[0062] Subsequently, aggregate the frequency sub-bands from the same level through concatenation and convolution operations to generate multi-level temporal context information {m1, m2, m3}:

[0063]

[0064] where g(·) represents a convolutional layer, represents the concatenation of the channel dimension.

[0065] The purpose of the encoder is to compress the input motion vector v t into a compact latent representation y3. The main structure of the encoder consists of three stacked encoding units. Each encoding unit includes: two residual units and a downsampling convolutional layer. To effectively alleviate the loss of high-frequency detail information in the latent representation y3 during the downsampling process, add the output of each encoding unit to the corresponding multi-level temporal context information generated by the temporal context generator, and the calculation formula is expressed as:

[0066]

[0067] where enc(·) represents the encoding unit.

[0068] The main task of the conditional entropy codec is to encode the latent representation y3 into a binary bitstream b and decode and recover the quantized latent variable from the binary bitstream b

[0069] In the conditional entropy codec, first round and quantize the latent representation y3 to generate the quantized latent variable Subsequently, use two stacked encoding units and a convolutional layer to extract the temporal prior from the temporal reference information and use a convolutional layer to fuse the hyperprior φ and the autoregressive prior ψ to generate the spatial prior Fuse the temporal prior and the spatial prior through a 3-layer 1×1 convolution to obtain the mean μ and variance σ of the Gaussian distribution of the latent variable , and the calculation formula is expressed as:

[0070]

[0071] Among them, convs(·) represents a three-layer 1×1 convolutional layer, and conv(·) represents a single convolutional layer. Finally, an arithmetic encoder is used to encode the latent variable into a binary bitstream b according to the Gaussian distribution mean μ and variance σ. In the entropy decoder, the latent variable is also decoded and recovered from the binary bitstream b according to the Gaussian distribution mean μ and variance σ

[0072] The decoder decodes the quantized latent variable through an inverse transformation to obtain the compressed motion vector The main structure of the decoder consists of three stacked decoding units. Each decoding unit includes: two residual units and an upsampling convolutional layer.

[0073] To effectively alleviate the loss of high-frequency detail information of the latent representation y3 during quantization, the output of each decoding unit is added to the corresponding multi-level temporal context information generated by the temporal context generator, and finally the compressed motion vector is obtained The calculation formula is expressed as:

[0074]

[0075] where dec(·) represents the decoding unit.

[0076] IV. Design a Temporal Prior Encoder to Compress Residuals

[0077] Similar to compressing the motion vector v t , a temporal prior encoder is used to efficiently compress the residual r t . The temporal prior encoder used for compressing the residual r t has the same structure as the temporal prior encoder used for compressing the motion vector v t , so its structure will not be elaborated here. Given the residual r t and the temporal reference information as inputs, the temporal prior encoder outputs the compressed residual The calculation formula is expressed as:

[0078]

[0079] where tcc(·) represents the temporal prior encoder.

[0080] V. Train the End-to-End Intelligent Video Coding Network

[0081] The end-to-end intelligent video coding network includes: a feature extraction module, a global temporal reference feature generation module, a motion estimation module, a motion compensation module, a temporal prior encoder, and a reconstruction module. Among them, the feature extraction module, the motion estimation module, the motion compensation module, and the reconstruction module adopt the neural network structure of excellent end-to-end intelligent video coding methods. In addition, a multi-stage training strategy is designed to asymptotically train the end-to-end intelligent video coding network.

[0082] The first stage: Training on a single video frame.

[0083] First, use the rate-distortion loss function and the prediction distortion loss function The sum of is used as the overall loss function \(L = L_{}\) r + \(L_{}\) p to train the neural network model for 5 epochs, where \(D(\cdot)\) represents the mean square error (MSE), \(R\) represents the bitrate, \(x_{}\) t represents the predicted frame, \(x_{}\) t represents the current encoded frame, represents the reconstructed frame, \(\lambda\) represents the hyperparameter for adjusting the bitrate, and its value is set to \(\{256, 512, 1024, 2048\}\). Subsequently, use the rate-distortion loss function \(L_{}\) r to optimize for another 5 epochs.

[0084] The second stage: Training on multiple video frames.

[0085] In the multi-frame training stage, the end-to-end intelligent video coding network is continuously optimized on three video frames. First, use the overall loss function \(L\) to train for 14 epochs, and then use the rate-distortion loss function \(L_{}\) r to train for another 6 epochs. Finally, in order to alleviate the problem of reference frame error accumulation when the network model encodes consecutive multiple frames, the cumulative rate-distortion loss function \(L_{}\) r is calculated by accumulating the rate-distortion loss functions \(L_{}\) of three frames. * And use the cumulative rate-distortion loss function \(L_{}\) * to train the neural network model for 5 epochs.

[0086] After training the video coding network, an end-to-end intelligent video coding model is obtained. This model takes a video sequence as input and finally outputs a compressed video sequence.

[0087] In the embodiments of the present invention, the video coding standard HEVC reference software HM-16.21 is compared with the method proposed in the present invention. See Figure 5, on the premise of the same reconstructed video quality, the present invention only needs to consume 95.79% of the bits of the HM-16.21 method. That is to say, compared with the HM-16.21 method, this method realizes a 4.21% bit saving, indicating that the proposed scheme of the present invention can effectively improve the video coding performance.

[0088] An end-to-end intelligent video coding device, the device includes: a processor and a memory, and program instructions are stored in the memory. The processor calls the program instructions stored in the memory to enable the device to execute the following method steps:

[0089] Construct an end-to-end intelligent video coding network framework composed of a feature extraction module, a global temporal reference feature generation module, a motion estimation module, a temporal prior encoder, a motion compensation module, and a reconstruction module;

[0090] The feature extraction module maps the current coding frame to the feature space to obtain the current coding frame feature, obtains the global reference feature through the global temporal reference feature generation module, and estimates the motion vector between the global reference feature and the current coding frame feature through the motion estimation module;

[0091] Compress the motion vector through the temporal prior encoder, based on the compressed motion vector and the global reference feature, use the motion compensation module to obtain the predicted feature of the current coding frame, use the temporal prior encoder to compress the residual between the predicted feature and the current coding frame feature, generate the reconstruction feature based on the residual, and obtain the final reconstructed frame through the reconstruction module.

[0092] Further, the temporal prior encoder takes the motion vector v t and the temporal reference information as inputs, and outputs the compressed motion vector including: a temporal context generator, an encoder, a decoder, and a conditional entropy codec.

[0093] Among them, the temporal context generator is used to extract the multi-level temporal context information {m1, m2, m3} in the temporal reference information , and aggregates the frequency subbands from the same level through concatenation and convolution operations to generate the multi-level temporal context information {m1, m2, m3}:

[0094]

[0095] Among them, g(·) represents a convolutional layer, represents the concatenation of the channel dimensions.

[0096] Among them, the encoder takes the input motion vector v tCompressed into a compact latent representation y3, the encoder consists of three stacked encoding units, and each encoding unit includes: two residual units and a downsampling convolutional layer. The output of each encoding unit is added to the multi-level temporal context information generated by the temporal context generator, and the calculation formula is expressed as:

[0097]

[0098] Among them, enc(·) represents the encoding unit.

[0099] Furthermore, the conditional entropy codec rounds and quantizes the latent representation y3 to generate a quantized latent variable Two stacked encoding units and one convolutional layer are used to extract the temporal prior from the temporal reference information A spatial prior is generated by using one convolutional layer to fuse the hyperprior and the autoregressive prior, and the temporal prior and the spatial prior are fused through three 1×1 convolutions to obtain the mean μ and variance σ of the Gaussian distribution of the latent variable ;

[0100] The quantized latent variable is decoded through an inverse transformation to obtain the compressed motion vector The decoder consists of three stacked decoding units, and each decoding unit includes: two residual units and an upsampling convolutional layer;

[0101] The output of each decoding unit is added to the multi-level temporal context information generated by the temporal context generator, and finally the compressed motion vector is obtained.

[0102] In the embodiments of the present invention, except for those with special specifications for the models of each device, the models of other devices are not limited, as long as the devices can perform the above functions.

[0103] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred embodiment, and the serial numbers of the above embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments.

[0104] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. An end-to-end intelligent video coding method, characterized in that, The method includes: Constructing an end-to-end intelligent video coding network framework composed of a feature extraction module, a global temporal reference feature generation module, a motion estimation module, a temporal prior encoder, a motion compensation module, and a reconstruction module; The feature extraction module maps the current coding frame to the feature space to obtain the current coding frame feature, obtains the global reference feature through the global temporal reference feature generation module, and estimates the motion vector between the global reference feature and the current coding frame feature through the motion estimation module; Compressing the motion vector through the temporal prior encoder, obtaining the predicted feature of the current coding frame based on the compressed motion vector and the global reference feature, compressing the residual between the predicted feature and the current coding frame feature using the temporal prior encoder, generating the reconstructed feature based on the compressed residual and the predicted feature, and obtaining the final reconstructed frame through the reconstruction module; Among them, the global time-domain reference feature generation module generates a global reference feature f for the motion estimation module and the motion compensation module t r , and dynamically updates the time-domain context state information once where the subscript t represents the current time t, and the superscript l i represents the context state information of the i-th level; The feature extraction module extracts short-term temporal context from the reference frame and represents it as follows: After cascading the updated time-domain context state information through two convolutional layers and then fusing it through two more convolutional layers, the long-term time-domain context is obtained. The calculation is as follows: Among them, Fusion(·), h0(·), h1(·), and h2(·) all represent two convolutional layers, represent the concatenation of the channel dimensions; Aggregate the short-term temporal context and the long-term temporal context to generate the final global temporal reference feature f t r : Wherein, g(·) represents a convolutional layer; The time-domain prior encoder takes the motion vector v t and the time-domain reference information as inputs, and outputs the compressed motion vector comprising: a time-domain context generator, an encoder, a decoder, and a conditional entropy codec; The time-domain context generator is used to extract time-domain reference information The multi-level time-domain context information {m1, m2, m3} in is aggregated for frequency subbands from the same level through concatenation and convolution operations to generate the multi-level time-domain context information {m1, m2, m3}: Among them, g(·) represents a convolutional layer, representing the concatenation of the channel dimensions; The encoder compresses the input motion vector v t into a compact latent representation y3. The encoder consists of three stacked encoding units, and each encoding unit includes: two residual units and a downsampling convolutional layer. The output of each encoding unit is added to the multi-level temporal context information generated by the temporal context generator, and the calculation formula is expressed as: Wherein, enc(·) represents an encoding unit; The conditional entropy codec rounds and quantizes the latent representation y3 to generate a quantized latent variable Two stacked encoding units and one convolutional layer are used to extract the temporal prior from the temporal reference information and one convolutional layer is used to fuse the hyperprior and the autoregressive prior to generate the spatial prior. The temporal prior and the spatial prior are fused through three 1×1 convolutions to obtain the mean μ and variance σ of the Gaussian distribution of the latent variable ; Decode the quantized latent variable through inverse transformation Obtain the compressed motion vector The decoder consists of three stacked decoding units, and each decoding unit includes: two residual units and an upsampling convolutional layer; Adding the output of each decoding unit to the multi-level temporal context information generated by the temporal context generator, and finally obtaining the compressed motion vector.

2. The end-to-end intelligent video encoding method according to claim 1, wherein The method further includes: training the end-to-end intelligent video coding network framework on single videos and multi-videos respectively.

3. An end-to-end intelligent video encoding device, characterized in that, The device includes: a processor and a memory, wherein program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to enable the device to execute the method according to any one of claims 1-2.

Citation Information

Patent Citations

  • Visual SLAM front-end pose estimation method based on deep learning

    CN111127557A

  • Multi-level image compression method using Transform

    CN113709455A