Learnable Video Coding Method, System, Device and Storage Medium
By performing spatial decomposition and motion estimation of video frames, combined with joint encoding and decoding and motion compensation, the problem of motion in the prior art that cannot accurately estimate the motion of objects with inconsistent motion is solved, and more efficient video encoding performance is achieved.
Patent Information
- Application Number
- CN202311760229.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-19
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-12-19
AI Technical Summary
Existing learning video encoding methods cannot accurately estimate the motion when processing objects with inconsistent motion, resulting in limited encoding performance.
By spatially decomposing the video frames, motion estimation is performed on the low-frequency structure and high-frequency details, and joint encoding and decoding is performed. The reference features are then spatially decomposed, motion compensation is performed, and multi-scale time domain context features are fused for more accurate inter-frame prediction.
Improve the performance of video encoding, enable more accurate estimate of motion vectors, reduce time domain redundancy, and thus improve encoding performance.
Smart Images

Figure CN117750020B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video coding technology, and in particular to a learnable video coding method, system, device and storage medium. Background Art
[0002] As a form of multimedia data, video is widely used in broadcasting, mobile live broadcasting, road monitoring, smart cities and other fields. For a video with a resolution of 1080p and 30 frames per second, the data volume can reach 180Mbytes per second. The huge amount of data has caused a huge cost for video transmission and storage. Therefore, before transmission and storage, it is usually necessary to compress the size of the video and encode the video into a more compact bit stream to reduce its transmission and storage cost.
[0003] Traditional video coding standards, such as H.264 / AVC, H.265 / HEVC, and H.266 / VVC, mostly adopt a block-based hybrid coding framework, which includes block-based motion prediction, motion compensation, transform, quantization, entropy coding and other modules. Although traditional video coding standards have achieved great success, their coding performance has also reached a bottleneck, and it is becoming increasingly difficult to achieve greater coding performance. In recent years, neural network-based learnable video coding methods have opened up a new direction, bringing hope for achieving greater coding performance. The learnable video coding method uses neural networks to implement each coding module in the traditional hybrid coding framework, and uses the rate-distortion (RDO) function to jointly train all coding modules.
[0004] Existing learnable conditional coding methods can be mainly divided into two categories, including residual coding-based methods and conditional coding-based methods.
[0005] The common point of these two methods is that they both require motion prediction and motion compensation. Motion prediction usually sends the current frame to be encoded and the reference frame into a motion estimation network, such as an optical flow network, to obtain the motion vector between the current frame and the reference frame, such as an optical flow (containing the motion vector of each pixel in the current frame). The predicted motion vector needs to be encoded and decoded. In the learnable video coding method, an autoencoder is often used to implement the encoding and decoding of the motion vector. The motion encoder compresses the predicted motion vector into a bitstream, and the motion decoder decodes the bitstream into a reconstructed motion vector. Motion compensation means that after obtaining the reconstructed motion vector, the reference frame needs to be used to obtain a prediction of the current frame to be encoded.
[0006] The main difference between these two methods is that after motion prediction and motion compensation, the residual coding method (Lu, G., Ouyang, W., Xu, D., Zhang, X., Cai, C., & Gao, Z. (2019). Dvc: An end-to-end deep video compression framework. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (pp. 11006-11015).) subtracts the current video frame to be encoded from the predicted frame to obtain the residual to reduce time domain redundancy, and then uses another autoencoder's encoding network to encode the residual to obtain the residual latent variable, which is then entropy encoded to obtain the bitstream. In the decoder, the entropy decoder re-decodes the bitstream into the residual latent variable, and the autoencoder's decoding network decodes the latent variable into the residual and then adds the predicted frame to obtain the reconstructed frame. In addition to residual coding in the pixel domain, Hu et al. (Hu, Z., Lu, G., & Xu, D. (2021). FVC: A new framework towards deep video compression in feature space. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (pp. 1502-1511).) also proposed residual coding in the feature domain. They first extracted deep features from the original video frame to be encoded and the reference frame, then performed motion prediction and motion compensation in the feature domain, and then encoded the residual of the depth features of the current frame and the depth features of the predicted frame.
[0007] For conditional coding methods, Li (Li, J., Li, B., & Lu, Y. (2021). Deep contextual video compression [DCVC]. Advances in Neural Information Processing Systems, 34, 18114-18125.) et al. proposed a DCVC learnable video coding method. In this method, after obtaining the predicted frame, the predicted frame is sent to the neural network to extract deep features as context features, and sent together with the frame to be encoded (the common method is to cascade concatenate in the channel dimension) to the encoding network of the autoencoder. Instead of explicitly calculating the residual, the encoding network automatically learns to reduce time domain redundancy. The encoding network encodes the input frame into latent variables, and then uses the entropy encoder to losslessly encode the latent variables into a bitstream. At the decoding end, the entropy decoder decodes the bitstream losslessly into latent variables, and the decoding network of the autoencoder decodes the latent variables into a reconstructed frame. Before the decoding network obtains the reconstructed frame, the context features are sent to the decoding network (the common method is to cascade concatenate in the channel dimension). Sheng et al. (Sheng, X., Li, J., Li, B., Li, L., Liu, D., & Lu, Y. (2022). Temporal context mining for learned video compression. IEEE Transactions on Multimedia.) also proposed a DCVC-TCM learnable video coding method based on DCVC. This method proposed motion compensation in the feature domain, and used the intermediate features of the decoding network before obtaining the reconstructed frame of the previous frame as the reference features for encoding the next frame. The reference features were motion compensated in the feature domain using the reconstructed optical flow to obtain the predicted features, and then multi-scale context features were extracted from the predicted features. In the process of encoding in the encoding network and decoding in the decoding network, the multi-scale context features are sent to the encoding network and the decoding network in a conditional coding manner, so as to utilize the temporal correlation and reduce the temporal redundancy.Li et al. (Li, J., Li, B., & Lu, Y. (2022, October). Hybrid spatial-temporal entropy modelling for neural video compression. In Proceedings of the 30th ACM International Conference on Multimedia (pp. 1503-1511) proposed a DCVC-HEM learnable video coding method, which uses the feature domain motion compensation and multi-scale context feature technology of DCVC-TCM, and further adds a hybrid spatiotemporal entropy model on this basis. Li et al. (Li, J., Li, B., & Lu, Y. (2023). Neural video compression with diverse contexts. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (pp.22616-22626).) further proposed the DCVC-DC learnable video coding method based on DCVC-HEM. This method proposed a hybrid spatiotemporal entropy model based on quadtree partitioning, which greatly improved the coding performance of the learnable video coding method and made its coding performance surpass the reference software of the traditional video coding standard H.266 / VVC.
[0008] Among the above schemes, the DCVC-DC learnable video coding method is most relevant to the present invention. However, its defect is that in a video frame, different moving objects often have different motion modes (such as non-uniform motion, rotation, and scaling), resulting in inconsistent motion in different regions of the video frame. For example, a local region may contain foreground and background objects at the same time, and their motions may be different. The inconsistent motion characteristics of objects in the region pose a huge challenge to motion estimation. However, the DCVC-DC learnable video coding method does not explicitly distinguish between objects with inconsistent motion. For some regions with inconsistently moving objects, it will reduce the inconsistency of motion of different objects, and cannot accurately estimate motion, thereby restricting the coding performance. Summary of the invention
[0009] The purpose of the present invention is to provide a learnable video coding method, system, device and storage medium, which can achieve more accurate inter-frame prediction and effectively improve the performance of video coding.
[0010] The objective of the present invention is achieved through the following technical solutions:
[0011] A learnable video encoding method, comprising:
[0012] Step 1: spatially decompose the current frame to be encoded and the reference frame, and perform motion estimation to obtain motion vectors of low-frequency structures and motion vectors of high-frequency details;
[0013] Step 2: jointly encode and jointly decode the motion vector of the low-frequency structure and the motion vector of the high-frequency detail;
[0014] Step 3: Decompose the reference features of the current frame to be encoded spatially, and use the reconstructed low-frequency structure motion vector and high-frequency detail motion vector obtained by joint decoding to perform motion compensation on the low-frequency structure part and high-frequency detail part of the reference features obtained by spatial decomposition, and then obtain multi-scale time domain context features through feature fusion;
[0015] Step 4: Encode and decode the current frame to be encoded by combining the multi-scale time domain context features;
[0016] Step 5: transform the decoded features outputted in step 4 to obtain a reconstructed frame of the current frame to be encoded and reference features for the next frame to be encoded.
[0017] A learnable video coding system includes a learnable video coding model, and video coding is performed by the learnable video coding model. The learnable video coding model includes:
[0018] The motion estimation module based on structure and detail decomposition is used to spatially decompose the current frame to be encoded and the reference frame, and perform motion estimation to obtain the motion vector of the low-frequency structure and the motion vector of the high-frequency detail;
[0019] A motion vector coding network based on structure and detail decomposition is used to jointly encode and decode motion vectors of low-frequency structures and motion vectors of high-frequency details;
[0020] The temporal context mining module based on structure and detail decomposition is used to spatially decompose the reference features of the current frame to be encoded, and use the reconstructed low-frequency structure motion vector and high-frequency detail motion vector obtained by joint decoding to perform motion compensation on the low-frequency structure part and high-frequency detail part of the reference features obtained by spatial decomposition, and then obtain multi-scale temporal context features through feature fusion;
[0021] A context coding network is used to encode and decode the current frame to be encoded by combining multi-scale temporal context features;
[0022] The frame generator is used to transform the decoded features output by the context decoder to obtain a reconstructed frame of the current frame to be encoded and a reference feature for the next frame to be encoded.
[0023] A processing device, comprising: one or more processors; a memory for storing one or more programs;
[0024] When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method.
[0025] A readable storage medium stores a computer program, which implements the above method when the computer program is executed by a processor.
[0026] It can be seen from the technical solution provided by the present invention that the current frame to be encoded and the reference frame are spatially decomposed to obtain a low-frequency structure part and a high-frequency detail part, and motion estimation is performed on the low-frequency structure part and the high-frequency detail part of the video. After spatial decomposition, the motion of the low-frequency structure of the two frames (i.e., the current frame to be encoded and the reference frame) includes the more consistent motion of the original two frames, and the motion difference in the local area is reduced, while the motion of the high-frequency detail part of the two frames includes the residual of the original inconsistent motion, and the reference features are also spatially decomposed to obtain the low-frequency structure part and the high-frequency detail part of the reference features; by performing spatial decomposition first and then motion estimation, the motion vector can be estimated more accurately, and by performing motion compensation on the decomposed low-frequency structure and high-frequency details respectively, more accurate time domain context features can be predicted, thereby reducing time domain redundancy to improve coding performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0028] Figure 1 A schematic diagram of a learnable video encoding method provided by an embodiment of the present invention;
[0029] Figure 2 A schematic diagram of a model architecture of a learnable video encoding method provided by an embodiment of the present invention;
[0030] Figure 3 A schematic diagram of spatial decomposition provided by an embodiment of the present invention;
[0031] Figure 4 A schematic diagram of motion estimation and motion compression based on spatial decomposition provided by an embodiment of the present invention;
[0032] Figure 5 A schematic diagram of motion compensation based on spatial decomposition provided by an embodiment of the present invention;
[0033] Figure 6 A schematic diagram of a first subjective quality comparison result provided by an embodiment of the present invention;
[0034] Figure 7 A schematic diagram of a second subjective quality comparison result provided by an embodiment of the present invention;
[0035] Figure 8 A schematic diagram of a processing device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0036] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the protection scope of the present invention.
[0037] First, the terms that may be used in this article are explained as follows:
[0038] The terms "include", "comprises", "contains", "has" or other descriptions with similar semantics should be interpreted as non-exclusive inclusion. For example, including certain technical feature elements (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, procedures, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products or products, etc.) should be interpreted as including not only certain technical feature elements explicitly listed, but also other technical feature elements known in the art that are not explicitly listed.
[0039] The following is a detailed description of a learnable video encoding method, system, device and storage medium provided by the present invention. The contents not described in detail in the embodiments of the present invention belong to the prior art known to professional and technical personnel in the field. If no specific conditions are specified in the embodiments of the present invention, the conventional conditions in the field or the conditions recommended by the manufacturer shall be followed.
[0040] Embodiment 1
[0041] The embodiment of the present invention provides a learnable video encoding method, such as Figure 1 As shown, it mainly includes the following steps:
[0042] Step 1: perform spatial decomposition on the current frame to be encoded and the reference frame, respectively, and perform motion estimation to obtain motion vectors of low-frequency structures and motion vectors of high-frequency details.
[0043] In this step: the current frame to be encoded x t and reference frame Perform spatial decomposition respectively to obtain the low-frequency structural parts of the current frame to be encoded and the reference frame And the high-frequency details Among them, s is the identifier of the low-frequency structure part, d is the identifier of the high-frequency detail part, and t is the video frame number; High frequency details Perform motion estimation separately to obtain the motion vector of the low-frequency structure and motion vectors for high-frequency details
[0044] Those skilled in the art can understand that the low-frequency structure generally refers to an image obtained by performing low-pass filtering on the image, and the high-frequency detail generally refers to an image obtained by performing high-pass filtering on the image.
[0045] Exemplarily, the image may be downsampled and upsampled in sequence to obtain a low-frequency structure, and then the low-frequency structure may be subtracted from the image to obtain high-frequency details.
[0046] Step 2: Jointly encode and jointly decode the motion vector of the low-frequency structure and the motion vector of the high-frequency detail.
[0047] In this step: the motion vector of the combined low-frequency structure and motion vectors of high-frequency details Quantized into latent variables [m t ]; for latent variables [m t ] to estimate the probability distribution parameters, and according to the estimated probability distribution parameters, the latent variable [m t ] lossless entropy encoding into motion vector bitstream; combined with probability distribution parameters, the motion vector bitstream is losslessly entropy decoded into latent variables [m t ], and then the hidden variable [m t ] Joint decoding to reconstruct low-frequency structure motion vector and high frequency detail motion vector
[0048] Step 3: spatially decompose the reference features of the current frame to be encoded, and use the reconstructed low-frequency structure motion vector and high-frequency detail motion vector obtained by joint decoding to perform motion compensation on the low-frequency structure part and high-frequency detail part of the reference features obtained by spatial decomposition, respectively, and then obtain multi-scale time domain context features through feature fusion.
[0049] In this step: the reference feature of the current frame to be encoded Perform spatial decomposition to obtain the low-frequency structure part of the reference feature and high frequency details Using the reconstructed low-frequency structure motion vector The low-frequency structure of the reference feature Compensate and obtain multi-scale context features of the low-frequency structure part and Utilize reconstructed high-frequency detail motion vectors High-frequency details of the reference feature Compensate to obtain multi-scale context features of high-frequency details and Among them, 0, 1, and 2 are the three scale identifiers. The larger the value of the identifier, the smaller the scale. The multi-scale context features of the low-frequency structure part and Multi-scale contextual features with high-frequency details and Corresponding fusion to obtain multi-scale temporal context features and
[0050] Step 4: Encode and decode the current frame to be encoded by combining multi-scale time-domain context features.
[0051] In this step: In the multi-scale time domain context features and With the help of t Quantized into latent variables [y t ]; for latent variables [y t ] to estimate the probability distribution parameters, and according to the estimated probability distribution parameters, the latent variable [y t ] is losslessly entropy encoded into a video bitstream; where 0, 1, and 2 are the identifiers of three scales; combined with the probability distribution parameters, the video bitstream is losslessly entropy decoded into a hidden variable [y t ], combined with multi-scale temporal context features and For the latent variable [y t ] to decode and obtain the decoding features.
[0052] Step 5: transform the decoded features outputted in step 4 to obtain a reconstructed frame of the current frame to be encoded and reference features for the next frame to be encoded.
[0053] In order to more clearly demonstrate the technical solution and technical effects provided by the present invention, the method provided by the embodiment of the present invention is described in detail with reference to specific embodiments below.
[0054] 1. Principle description
[0055] In a video frame, different moving objects often have different motion modes, such as non-uniform motion, rotation, and scaling, which leads to inconsistent motion in different regions of the video frame. For example, a local region may contain foreground and background objects at the same time, and their motions may be different. The inconsistent motion characteristics of objects in the region pose a huge challenge to motion estimation. In the existing learnable video coding methods, during training, the reconstructed optical flow is often used to perform motion compensation on the reference frame in the pixel domain to obtain a predicted frame, and then the mean square error between the predicted frame and the current frame to be encoded is calculated to train the optical flow estimation network and the optical flow autoencoder. However, this training method does not explicitly distinguish objects with inconsistent motion, and can only obtain the average minimum prediction error of all regions. For some regions with inconsistently moving objects, this training method will reduce the inconsistency of motion of different objects and cannot accurately estimate motion. Compared with the prior art, the present invention spatially decomposes the video to obtain a low-frequency structure part and a high-frequency detail part, and performs motion estimation on the low-frequency structure part and the high-frequency detail part of the video respectively. After spatial decomposition, the motion of the low-frequency structure of the two frames contains the more consistent motion of the original two frames, the motion difference of the local area is reduced, and the motion of the high-frequency detail parts of the two frames contains the residual of the original inconsistent motion. Then the reference feature is also spatially decomposed to obtain the low-frequency structure part and high-frequency detail part of the reference feature. Using the estimated motion of the low-frequency structure part and the high-frequency detail part, the low-frequency structure part and the high-frequency detail part of the reference feature are motion compensated respectively, and then the low-frequency structure prediction feature and the high-frequency detail feature after motion compensation are fused to obtain the fused prediction feature (multi-scale temporal context feature). ). The present invention can estimate the motion vector more accurately by performing spatial decomposition first and then motion estimation. By performing motion compensation on the decomposed low-frequency structure and high-frequency details respectively, more accurate time domain context features can be predicted, thereby reducing time domain redundancy and improving coding performance.
[0056] 2. Introduction of the plan.
[0057] Comparison Figure 1 As shown in the process, step 1 is implemented by a motion estimation module based on structure and detail decomposition, step 2 is implemented by a motion vector coding network based on structure and detail decomposition, step 3 is implemented by a temporal context mining module based on structure and detail decomposition, step 4 is implemented by a context coding network, and step 5 is implemented by a frame generator; together they form a learnable video coding model, as shown in Figure 2 shown.
[0058] 1. SDD-based Motion Estimation module.
[0059] The input of this module is the current frame to be encoded x t and reference frame The output is the motion vector of the low-frequency structure of both sides and the motion vectors of the high-frequency details like Figure 3 As shown in the figure, it is a schematic diagram of spatial decomposition. First, the current frame to be encoded x t and reference frame Perform spatial structure and detail decomposition to obtain the low-frequency structural parts of the current frame and the reference frame and high frequency details
[0060] Exemplarily, the video frame may be first bilinearly downsampled and then bilinearly upsampled to obtain a low-frequency structure portion, and then the low-frequency structure portion may be subtracted from the original video frame to obtain a high-frequency detail portion.
[0061]
[0062] Among them, Down represents bilinear downsampling, and Up represents bilinear upsampling.
[0063] like Figure 4 As shown in the figure, it is a schematic diagram of motion estimation and motion compression (Compression) based on spatial decomposition. The dotted box on the left is the motion estimation part based on spatial decomposition, and the dotted box on the right is the motion compression (Compression) based on spatial decomposition, which includes the joint encoding and joint decoding process implemented by the motion vector encoder and motion vector decoder based on structure and detail decomposition, which will be specifically introduced later.
[0064] like Figure 4 As shown in the dotted box on the left, after obtaining the low-frequency structure part and high-frequency detail part of the current frame to be encoded and the reference frame, the low-frequency structure part and high frequency details Perform motion estimation to obtain the respective motion vectors (MV) Exemplarily, SpyNet can be selected as the optical flow estimation network to estimate pixel-level motion vectors, that is, a motion vector is estimated for each pixel position.
[0065] In practical applications, users can independently design high- and low-frequency decomposition methods. The present invention emphasizes that before motion estimation, the video frame is first spatially decomposed to obtain the low-frequency structure part and the high-frequency detail part of the video frame, and then motion estimation is performed on them respectively to obtain their respective motion information.
[0066] 2. Motion vector coding network based on structure and detail decomposition (SDD).
[0067] In an embodiment of the present invention, a motion vector coding network based on structure and detail decomposition (SDD) mainly includes: a motion vector encoder based on structure and detail decomposition (SDD) (SDD-based MV Encoder), a motion vector decoder based on structure and detail decomposition (SDD) (SDD-based MV Decoder) and a motion entropy model (MV Entropy Model).
[0068] like Figure 4 As shown, the input of the motion vector encoder based on structure and detail decomposition is the current frame to be encoded x t and reference frame The motion vector of the low-frequency structure part and the motion vectors of the high-frequency details The output is a bitstream of motion vectors. The function of this module is to jointly encode the motion vectors of the low-frequency structure and the motion vectors of the high-frequency details is the quantized latent variable [m t Then, the motion entropy model is used to calculate [m t ] is used for probability distribution, and the arithmetic entropy encoder (AE) is used to convert [m t ]Lossless entropy encoding into bitstream.
[0069] The input of the motion entropy model is the motion latent variable [m t ], the output is the probability distribution parameters of motion latent variables. The function of the motion entropy model is to estimate the probability distribution parameters of motion latent variables and perform entropy coding on the latent variables.
[0070] like Figure 4 As shown in the figure, the input of the motion vector decoder based on structure and detail decomposition is the motion vector code stream transmitted from the encoder to the decoder, and the output is the reconstructed motion vector of the low-frequency structure. and the motion vectors of the high-frequency details The motion vector bitstream is firstly decoded by the arithmetic decoder (AD) into a latent variable [m t ], and then the motion vector decoder converts [m t ] Joint decoding to reconstruct motion vectors of low-frequency structures and the motion vectors of the high-frequency details
[0071] In the embodiment of the present invention, the motion vector encoder and the motion vector decoder based on structure and detail decomposition form an autoencoder structure, and the user can independently design the network structure. The present invention emphasizes that the motion vector decoder can reconstruct the motion vector of the low-frequency structure part. and the motion vectors of the high-frequency details Perform joint encoding and joint decoding.
[0072] 3. SDD-based Temporal Context Mining module.
[0073] like Figure 5 As shown in the figure, it is a schematic diagram of motion compensation based on spatial decomposition. The input of this module is the reference feature Motion vectors of reconstructed low-frequency structures Motion vectors of reconstructed high-frequency details The output is a multi-scale temporal context First, the reference feature Decompose the structure and details in space to obtain the low-frequency structure part of the reference feature and high frequency details
[0074] Exemplarily, the spatial decomposition solution introduced above is also adopted.
[0075]
[0076] Then, the motion vector of the reconstructed low-frequency structure is used Motion vectors of reconstructed high-frequency details right and Motion compensation (warp operation) is performed in the feature domain, and the predicted features are fused step by step from small scale to large scale to obtain the multi-scale context of the low-frequency structure part. and multi-scale context of high-frequency details Then and Add together to get the fused time domain context
[0077] 4. Context encoding network.
[0078] In the embodiment of the present invention, the contextual coding network includes: a contextual encoder (Contextual Encoder), a contextual decoder (Contextual Decoder) and a contextual entropy model (Contextual Entropy Model).
[0079] The input of the context encoder is the current video frame to be encoded x t and multi-scale temporal context The output is a video code stream. The context encoder uses a multi-scale temporal context Encode the current video frame to be encoded with the help of xt is the quantized latent variable [y t Then use the context entropy model to analyze [y t ] is used for probability distribution, and the arithmetic entropy encoder (AE) is used to convert [y t ]Lossless entropy encoding into video bitstream.
[0080] The input of the context entropy model is the video latent variable [y t ] and multi-scale temporal context The output is the video latent variable [y t The function of the context entropy model is to estimate the probability distribution parameters of the video bitstream and perform entropy coding.
[0081] The input of the context decoder is the video bitstream and the multi-scale temporal context The output is an incompletely decoded feature. The video bitstream is first decoded by the arithmetic decoder (AD) according to the common probability distribution of the codec end, and the lossless entropy decoding is converted into a latent variable [y t ], the context decoder transforms the hidden variable [y t ] is decoded into a feature.
[0082] In the embodiment of the present invention, the context encoder and decoder form a common autoencoder structure, and the user can independently design the network structure.
[0083] 5. Frame Generator:
[0084] The input of the frame generator is the features output by the context decoder, and the output is the reconstructed frame in the pixel domain. and the reconstructed features as reference features for the next frame The function of the frame generator is to transform the reconstructed features into the pixel domain to obtain the reconstructed video in the pixel domain. In getting Previously, the input features before the last convolution layer of the frame generator were Used for encoding the next frame as a reference feature for the next frame.
[0085] At the same time, the reconstructed visual frame and reference features shown by the frame generator are placed in a picture and feature buffer unit (Picture & Feature Buffer) for use in encoding the next frame.
[0086] 3. Model training plan.
[0087] The above model provided in the embodiment of the present invention needs to be trained in advance, and the training method is as follows:
[0088] Step 11: Obtain training data and input it into the learnable video coding model, and execute steps 1 to 2.
[0089] Step 12: Calculate the loss function L using the error between the output predicted frame of the current frame to be encoded and the current frame to be encoded 1 , and use this to train the motion estimation module based on structure and detail decomposition, the motion vector encoder based on structure and detail decomposition, and the motion vector decoder based on structure and detail decomposition.
[0090] For example, the mean square error can be used to calculate the loss function L 1 , expressed as:
[0091] The predicted frame of the current frame to be encoded here It is obtained by motion compensating the reference frame using the motion vector of the low-frequency structure and the motion vector of the high-frequency detail. It is an estimate of the current frame to be encoded.
[0092] Step 13: In the loss function L 1 The motion vector latent variable bit rate term estimated by the motion vector entropy model is added, and the Lagrange multiplier λ is used to control the balance between the mean square error and bit rate of the predicted frame in the pixel domain and the current frame to be encoded, and the loss function L is obtained. 2 , in order to jointly train the motion estimation module based on structure and detail decomposition, the motion vector encoder based on structure and detail decomposition, the motion vector decoder based on structure and detail decomposition, and the motion vector entropy model.
[0093] Among them, the motion vector latent variable rate term uses the estimated motion vector latent variable [m t ] is obtained, and the probability value at each position in the probability distribution parameter is recorded as p, then the bit rate is -log2(p), and the motion vector latent variable bit rate term is obtained by integrating all positions, recorded as R([m t ]), then the loss function
[0094] Step 14: After completing step 13, fix these modules and execute steps 1 to 5 again. Use the error between the reconstructed frame of the current frame to be encoded and the current frame to be encoded to calculate the loss function L 3 , and use it to train the time domain context mining module, context encoder, context decoder and frame generator based on structure and detail decomposition.
[0095] Similarly, the loss function L is calculated using the mean square error 3 , expressed as:
[0096] Step 15: The context entropy model is also used in the encoding and decoding of step 4. After completing step 14, the loss function L 3Add the video latent variable bit rate term estimated by the context entropy model to obtain the loss function L 4 , in order to jointly train the temporal context mining module, context encoder, context decoder, context entropy model and frame generator based on structure and detail decomposition.
[0097] Similarly, the video latent variable bit rate term is also calculated using the estimated latent variable [y t ] is obtained, and the video latent variable bit rate term is recorded as R([y t ]), then the loss function
[0098] Step 16: After completing step 15, in the loss function L 4 Add the motion vector latent variable rate term described in step 13 to obtain the loss function L t , thereby jointly training the entire learnable video coding model.
[0099] Here the loss function
[0100] 4. Experimental description.
[0101] The learnable video coding method provided by the embodiment of the present invention has achieved the best coding performance compared with the current video coding method. Specifically, under the condition that the intra-frame spacing is 32, the coding gain is measured using BD-rate in the RGB color space, and the reference software VTM-13.2 of the H.266 / VVC coding standard is used as the baseline, and is configured as encoder_lowdelay_vtm. Negative values represent the percentage of improvement in coding performance, and positive values represent the percentage of decrease in coding performance. The results are shown in Tables 1 and 2. The HM in the table is the reference software of the H.265 / HEVC coding standard, configured as encoder_lowdelay_main_rext. RLVC, M-LVC, DVC_Pro, DCVC, CANF-VC, TCMVC, and HEM in the table are existing end-to-end video coding methods. Ours is the method proposed by the present invention.
[0102] Table 1: Performance gain relative to VTM, the difference between the reconstructed video and the original video is measured by PSNR
[0103]
[0104] Table 2: Performance gain relative to VTM. The difference between the reconstructed video and the original video is measured by MS-SSIM.
[0105]
[0106] In addition to bringing about coding performance gains in objective indicators, the method proposed in the present invention can also achieve better subjective quality. Figure 6 As shown, the first subjective quality comparison result, specifically, it is a subjective quality comparison of the uncoded video, the VTM reconstructed video and the reconstructed video obtained by the method proposed in the present invention. It can be seen that this method can retain more texture details.
[0107] Through the inter-frame prediction technology based on structure and detail decomposition proposed by the present invention, more accurate motion information can be obtained. Figure 7 As shown, the second subjective quality comparison result, specifically, it is a subjective quality comparison of an uncoded video, a predicted video obtained using the inter-frame prediction technology without structure and detail decomposition, and a predicted video obtained using the inter-frame prediction technology with structure and detail decomposition, as well as a comparison of the residual sizes between the predicted video and the coded video; it can be seen that the motion vector obtained by using the inter-frame prediction technology based on structure and detail decomposition proposed in the present invention to perform motion compensation on the reference frame, the edge of the obtained predicted frame is smoother and the prediction error is smaller.
[0108] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above embodiments can be implemented by software, or by means of software plus necessary general hardware platforms. Based on such understanding, the technical solutions of the above embodiments can be embodied in the form of software products, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0109] Embodiment 2
[0110] The present invention also provides a learnable video coding system, which is mainly used to implement the method provided in the above embodiment. The system includes a learnable video coding model. Video coding is performed by the learnable video coding model. The learnable video coding model includes:
[0111] The motion estimation module based on structure and detail decomposition is used to spatially decompose the current frame to be encoded and the reference frame, and perform motion estimation to obtain the motion vector of the low-frequency structure and the motion vector of the high-frequency detail;
[0112] A motion vector coding network based on structure and detail decomposition is used to jointly encode and decode motion vectors of low-frequency structures and motion vectors of high-frequency details;
[0113] The temporal context mining module based on structure and detail decomposition is used to spatially decompose the reference features of the current frame to be encoded, and use the reconstructed low-frequency structure motion vector and high-frequency detail motion vector obtained by joint decoding to perform motion compensation on the low-frequency structure part and high-frequency detail part of the reference features obtained by spatial decomposition, and then obtain multi-scale temporal context features through feature fusion;
[0114] A context coding network is used to encode and decode the current frame to be encoded by combining multi-scale temporal context features;
[0115] The frame generator is used to transform the decoded features output by the context decoder to obtain a reconstructed frame of the current frame to be encoded and a reference feature for the next frame to be encoded.
[0116] Considering that the specific technical details of the above model have been explained in detail in the previous embodiments, they will not be repeated here.
[0117] Embodiment 3
[0118] The present invention also provides a processing device, such as Figure 8 As shown, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the methods provided in the aforementioned embodiments.
[0119] Furthermore, the processing device also includes at least one input device and at least one output device; in the processing device, the processor, memory, input device, and output device are connected via a bus.
[0120] In the embodiment of the present invention, the specific types of the memory, input device and output device are not limited; for example:
[0121] The input device may be a touch screen, an image acquisition device, a physical button or a mouse, etc.;
[0122] The output device may be a display terminal;
[0123] The memory may be a random access memory (RAM) or a non-volatile memory, such as a disk memory.
[0124] Embodiment 4
[0125] The present invention also provides a readable storage medium storing a computer program, which implements the method provided in the above embodiment when the computer program is executed by a processor.
[0126] In the embodiment of the present invention, the readable storage medium is a computer-readable storage medium and can be set in the aforementioned processing device, for example, as a memory in the processing device. In addition, the readable storage medium can also be a U disk, a mobile hard disk, a read-only memory (ROM), a disk or an optical disk, etc., which can store program codes.
[0127] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A learnable video coding method, It is characterized in that include: Step 1: spatially decompose the current frame to be encoded and the reference frame, and perform motion estimation to obtain motion vectors of low-frequency structures and motion vectors of high-frequency details; Step 2: jointly encode and jointly decode the motion vector of the low-frequency structure and the motion vector of the high-frequency detail; Step 3: Decompose the reference features of the current frame to be encoded spatially, and use the reconstructed low-frequency structure motion vector and high-frequency detail motion vector obtained by joint decoding to perform motion compensation on the low-frequency structure part and high-frequency detail part of the reference features obtained by spatial decomposition, and then obtain multi-scale time domain context features through feature fusion; Step 4: Encode and decode the current frame to be encoded by combining the multi-scale time domain context features; Step 5: transform the decoded features outputted in step 4 to obtain a reconstructed frame of the current frame to be encoded and reference features for the next frame to be encoded.
2. A learnable video encoding method according to claim 1, It is characterized in that In step 1: For the current frame to be encoded x t and reference frame Perform spatial decomposition respectively to obtain the low-frequency structural parts of the current frame to be encoded and the reference frame And the high-frequency details Among them, s is the identifier of the low-frequency structure part, d is the identifier of the high-frequency detail part, and t is the video frame number; For low frequency structural parts High frequency details Perform motion estimation separately to obtain the motion vector of the low-frequency structure and motion vectors for high-frequency details 3. A learnable video encoding method according to claim 1, It is characterized in that In step 2: Motion vectors associated with low-frequency structures and motion vectors of high-frequency details Quantized into latent variables [m t ]; For hidden variables [m t ] to estimate the probability distribution parameters, and according to the estimated probability distribution parameters, the latent variable [m t ]Lossless entropy encoding into motion vector code stream; Combined with the probability distribution parameters, the motion vector bitstream is losslessly entropy decoded into latent variables [m t ], and then the hidden variable [m t ] Joint decoding to reconstruct low-frequency structure motion vector and high frequency detail motion vector 4. A learnable video encoding method according to claim 1, It is characterized in that In step 3: The reference feature of the current frame to be encoded Perform spatial decomposition to obtain the low-frequency structure part of the reference feature and high frequency details Using the reconstructed low-frequency structure motion vector The low-frequency structure of the reference feature Compensate and obtain multi-scale context features of the low-frequency structure part and Utilize reconstructed high-frequency detail motion vectors High-frequency details of the reference feature Compensate to obtain multi-scale context features of high-frequency details and Among them, 0, 1, and 2 are the identifiers of the three scales; The multi-scale context features of the low-frequency structure part and Multi-scale contextual features with high-frequency details and Corresponding fusion to obtain multi-scale temporal context features and 5. A learnable video encoding method according to claim 1, It is characterized in that In step 4: Multi-scale temporal context features and With the help of t Quantized into latent variables [y t ]; for latent variables [y t ] to estimate the probability distribution parameters, and according to the estimated probability distribution parameters, the latent variable [y t ] Lossless entropy coding is used to encode the video bitstream; 0, 1, and 2 are the identifiers of three scales; Combined with the probability distribution parameters, the video stream is losslessly entropy decoded into a hidden variable [y t ], combined with multi-scale temporal context features and For the latent variable [y t ] to decode and obtain the decoding features.
6. A learnable video encoding method according to any one of claims 1 to 5, It is characterized in that Step 1 is implemented through a motion estimation module based on structure and detail decomposition, step 2 is implemented through a motion vector coding network based on structure and detail decomposition, step 3 is implemented through a time domain context mining module based on structure and detail decomposition, step 4 is implemented through a context coding network, and step 5 is implemented through a frame generator; together they form a learnable video coding model, and the learnable video coding model is pre-trained.
7. A learnable video encoding method according to claim 6, It is characterized in that The learnable video encoding model training method is as follows: Step 11: Obtain training data and input it into the learnable video coding model, and execute steps 1 to 2; Step 12: Calculate the loss function L using the error between the output predicted frame of the current frame to be encoded and the current frame to be encoded 1 , wherein the prediction frame of the current frame to be encoded uses the motion vector of the low-frequency structure and the motion vector of the high-frequency detail to perform motion compensation on the reference frame, and the motion vector coding network based on structure and detail decomposition includes: a motion vector encoder based on structure and detail decomposition, a motion vector decoder based on structure and detail decomposition, and a motion vector entropy model, and the motion vector entropy model is used to estimate the probability distribution parameters used in joint encoding and joint decoding; using the loss function L 1 Training a motion estimation module based on structure and detail decomposition, a motion vector encoder based on structure and detail decomposition, and a motion vector decoder based on structure and detail decomposition; Step 13: In the loss function L 1 Add the motion vector latent variable rate term estimated by the motion vector entropy model to obtain the loss function L 2 , thereby jointly training a motion estimation module based on structure and detail decomposition, a motion vector encoder based on structure and detail decomposition, a motion vector decoder based on structure and detail decomposition, and a motion vector entropy model; wherein the motion vector latent variable bit rate term is determined by the corresponding probability distribution parameter; Step 14: After completing step 13, fix the modules trained in steps 12 and 13, and execute steps 1 to 5 again. Use the error between the reconstructed frame of the current frame to be encoded and the current frame to be encoded to calculate the loss function L 3 , where the context coding network includes: a context encoder, a context decoder and a context entropy model. The context entropy model is used to estimate the probability distribution parameters used when encoding and decoding the current frame to be encoded; using the loss function L 3 Train the temporal context mining module, context encoder, context decoder and frame generator based on structure and detail decomposition; Step 15: In the loss function L 3 Add the video latent variable bit rate term estimated by the context entropy model to obtain the loss function L 4 , in order to jointly train the temporal context mining module, context encoder, context decoder, context entropy model and frame generator based on structure and detail decomposition; wherein the video latent variable bit rate term is determined by the corresponding probability distribution parameters; Step 16: After completing step 15, in the loss function L 4 Add the motion vector latent variable rate term described in step 13 to obtain the loss function L 5 , thereby jointly training the entire learnable video coding model.
8. A learnable video coding system, It is characterized in that A learnable video coding model is included, and video coding is performed by the learnable video coding model, wherein the learnable video coding model includes: The motion estimation module based on structure and detail decomposition is used to spatially decompose the current frame to be encoded and the reference frame, and perform motion estimation to obtain the motion vector of the low-frequency structure and the motion vector of the high-frequency detail; A motion vector coding network based on structure and detail decomposition is used to jointly encode and decode motion vectors of low-frequency structures and motion vectors of high-frequency details; The temporal context mining module based on structure and detail decomposition is used to spatially decompose the reference features of the current frame to be encoded, and use the reconstructed low-frequency structure motion vector and high-frequency detail motion vector obtained by joint decoding to perform motion compensation on the low-frequency structure part and high-frequency detail part of the reference features obtained by spatial decomposition, and then obtain multi-scale temporal context features through feature fusion; A context coding network is used to encode and decode the current frame to be encoded by combining multi-scale temporal context features; The frame generator is used to transform the decoded features output by the context decoder to obtain a reconstructed frame of the current frame to be encoded and a reference feature for the next frame to be encoded.
9. A processing device, It is characterized in that include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
10. A readable storage medium storing a computer program, It is characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
End-to-end intelligent video coding method and device
CN115278262A
Depth high dynamic range imaging based on wavelet transform
CN116113978A