Neural network based time-space adaptive video compression

By using an end-to-end neural network-based video codec, combined with STAC and BPLC components, adaptive video compression was achieved, solving the problem of low video encoding and decoding efficiency in existing technologies and improving the rate-distortion performance of video compression.

CN115442618BActive Publication Date: 2026-02-10FACE CUTE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210625548.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-06-03
Filing Date
2022-06-02
Publication Date
2026-02-10
Estimated Expiration
2042-06-02

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies struggle to effectively utilize temporal and spatial correlations for adaptive compression when processing digital video, especially in random access scenarios, resulting in high bandwidth requirements and low compression efficiency.

Method used

An end-to-end neural network-based video codec is used, which combines a spatial-temporal adaptive compression (STAC) component and a bilateral predictive learning compression (BPLC) component to adaptively compress video frames. The optimal compression scheme is selected through frame extrapolation compression (FEC) and image compression branches.

Benefits of technology

It improves the rate-distortion performance of video compression, optimizes encoding and decoding efficiency, adapts to the characteristics of different types of video frames, and achieves more efficient bandwidth utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115442618B_ABST
    Figure CN115442618B_ABST
Patent Text Reader

Abstract

The present disclosure relates to neural network based spatio-temporal adaptive video compression. A mechanism of processing video data is disclosed. It is determined to apply an end-to-end neural network based video codec to a current video unit of a video. The end-to-end neural network based video codec comprises a spatio-temporal adaptive compression (STAC) component comprising a frame extrapolation compression (FEC) branch and an image compression branch. A conversion between the current video unit and a bitstream of the video is performed by the end-to-end neural network based video codec.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This patent application claims the benefit of U.S. Provisional Patent Application No. 63 / 196,332, filed June 3, 2021, by Zhaobin Zhang et al., entitled “Neural Network-Based Temporal-Spatial Adaptive Video Compression,” which is incorporated herein by reference. Technical Field

[0003] This patent document relates to digital video processing. Background Technology

[0004] Digital video accounts for the largest share of bandwidth usage in the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention

[0005] The first aspect relates to a method for processing video data, comprising: determining a current video unit of a video to which an end-to-end neural network-based video codec will be applied, wherein the end-to-end neural network-based video codec includes a spatial-temporal adaptive compression (STAC) component, the STAC component including a frame extrapolation compression (FEC) branch and an image compression branch; and performing a conversion between the current video unit and the bitstream of the video via the end-to-end neural network-based video codec.

[0006] Alternatively, in any of the foregoing aspects, another implementation of this aspect provides that the current video unit is assigned for conversion via one of the FEC branch and the image compression branch based on the motion information of the current video unit, the entropy of the current video unit, or a combination thereof.

[0007] Alternatively, in any of the foregoing aspects, another embodiment of this aspect provides that the FEC branch uses multiple codec frames as reference frames to predict the current video unit, and wherein the bitstream includes an indication of motion information between the reference frames and the current video unit.

[0008] Alternatively, in any of the foregoing aspects, another embodiment of this aspect provides that the FEC branch uses multiple codec frames as reference frames to predict the current video unit, and wherein the bitstream does not include motion information between the reference frames and the current video unit.

[0009] Optionally, in any of the foregoing aspects, another embodiment of the end-to-end neural network-based video codec further includes a bilateral predictive learning compression (BPLC) component, wherein the STAC component performs the transformation of the current video unit when the current video unit is a keyframe, and wherein the BPLC component performs the transformation of the current video unit when the current video unit is not a keyframe.

[0010] Alternatively, in any of the foregoing aspects, another embodiment of that aspect provides that the BPLC component interpolates the current video unit based on at least one preceding reference frame and at least one following reference frame.

[0011] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides prediction of the current video unit based on a previously reconstructed keyframe, a subsequently reconstructed keyframe, motion information, multiple reference frames, or a combination thereof.

[0012] Alternatively, in any of the foregoing aspects, another implementation of this aspect provides that the end-to-end neural network-based video codec selects one of a plurality of inter-frame predictive compression networks to apply to the current video unit.

[0013] Alternatively, in any of the foregoing aspects, another embodiment of that aspect provides that the plurality of inter-frame predictive compression networks include a stream-based network that encodes and decodes the current video unit by deriving an optical flow.

[0014] Alternatively, in any of the foregoing aspects, another embodiment of this aspect provides that the plurality of inter-frame prediction compression networks include a kernel-based network that encodes and decodes the current video unit by convolving a learned kernel with one or more reference frames to derive a prediction frame.

[0015] Alternatively, in any of the foregoing aspects, another implementation of this aspect provides that the selection of one of the plurality of inter-frame predictive compression networks is included in the bitstream.

[0016] Alternatively, in any of the foregoing aspects, another implementation of this aspect provides that the end-to-end neural network-based video codec selects one of a plurality of intra-frame predictive compression networks to apply to the current video unit.

[0017] Alternatively, in any of the foregoing aspects, another embodiment of this aspect provides that the end-to-end neural network-based video codec uses a combination of block prediction and frame prediction to encode and decode the current video unit.

[0018] Alternatively, in any of the foregoing aspects, another implementation of this aspect provides that the end-to-end neural network-based video codec selects one of a plurality of motion compression networks to apply to the current video unit.

[0019] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the conversion includes encoding the current video unit into the bitstream.

[0020] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the conversion includes decoding the current video unit from the bitstream.

[0021] The second aspect relates to an apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: determine a current video unit of video to be applied to the video using an end-to-end neural network-based video codec, wherein the end-to-end neural network-based video codec includes a spatial-temporal adaptive compression (STAC) component, the STAC component including a frame extrapolation compression (FEC) branch and an image compression branch; and perform a conversion between the current video unit and the bitstream of the video using the end-to-end neural network-based video codec.

[0022] Alternatively, in any of the foregoing aspects, another implementation of this aspect provides that the current video unit is assigned for conversion via one of the FEC branch and the image compression branch based on the motion information of the current video unit, the entropy of the current video unit, or a combination thereof.

[0023] Optionally, in any of the foregoing aspects, another embodiment of the end-to-end neural network-based video codec further includes a bilateral predictive learning compression (BPLC) component, wherein the STAC component performs the transformation of the current video unit when the current video unit is a keyframe, and wherein the BPLC component performs the transformation of the current video unit when the current video unit is not a keyframe.

[0024] The third aspect relates to a non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: determining a current video unit to which an end-to-end neural network-based video codec will be applied, wherein the end-to-end neural network-based video codec includes a spatial-temporal adaptive compression (STAC) component, the STAC component including a frame extrapolation compression (FEC) branch and an image compression branch; and generating the bitstream using the end-to-end neural network-based video codec.

[0025] For clarity, any of the above embodiments may be combined with any one or more of the other embodiments described above to form new embodiments within the scope of this disclosure.

[0026] These and other features will become clearer from the following detailed description taken in conjunction with the accompanying drawings and claims. Attached Figure Description

[0027] To gain a more complete understanding of this disclosure, reference is now made to the following brief description in conjunction with the accompanying drawings and detailed description, wherein the same or similar reference numerals denote the same or similar parts.

[0028] Figure 1 This is a schematic diagram illustrating an example transform encoding / decoding scheme.

[0029] Figure 2 This is a schematic diagram illustrating the comparison of compression schemes.

[0030] Figure 3 This is a schematic diagram illustrating an example neural network framework.

[0031] Figure 4 This is a schematic diagram illustrating the direct synthesis scheme for inter-frame predictive compression.

[0032] Figure 5 This is a schematic diagram illustrating a kernel-based scheme for inter-frame predictive compression.

[0033] Figure 6 This is a schematic diagram illustrating an example residual refinement network.

[0034] Figure 7 This is a block diagram of an example video processing system.

[0035] Figure 8 This is a block diagram of an example video processing device.

[0036] Figure 9 This is a flowchart illustrating an example video processing method.

[0037] Figure 10 A block diagram of an example video codec system is shown.

[0038] Figure 11 A block diagram of an example encoder is shown.

[0039] Figure 12 A block diagram of an example decoder is shown.

[0040] Figure 13 This is a schematic diagram of an example encoder. Detailed Implementation

[0041] It should be understood first and foremost that, although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or existing. This disclosure should not be limited in any way to the illustrative embodiments, drawings, and techniques shown below, including the exemplary designs and implementations shown and described herein, but modifications can be made within the scope of the appended claims, together with their full equivalents.

[0042] This patent document relates to video compression using neural networks, and more specifically, to spatial-temporal adaptive compression of neural network-based (NN-based) video compression for random access scenarios. The disclosed mechanism can also be applied to hybrid compression frameworks, in which neural network-based encoding and decoding tools are integrated into the framework of video codec standards, such as high-efficiency video coding (HEVC), versatile video coding (VVC), etc.

[0043] The described techniques involve improved algorithms and methods for optimizing rate-distortion (RD) performance. Typically, these techniques support temporal adaptive compression within neural network-based video codec frameworks, for example, by providing inter-frame predictive compression and / or intra-frame predictive compression modes for keyframes and / or providing content-adaptive compression mode selection mechanisms. For instance, spatial-temporal adaptive compression can remove temporal dependencies in keyframes, leading to better RD performance. Furthermore, spatial-temporal adaptive compression can be performed at the basic unit (BU) level. For example, keyframes can select a more optimized compression scheme between block-level inter-frame prediction and intra-frame prediction schemes.

[0044] Deep learning is developing in multiple fields, such as computer vision and image processing. Inspired by the successful application of deep learning technology in computer vision, neural image / video compression technology is being researched for application in image / video compression. Neural networks are designed based on interdisciplinary research in neuroscience and mathematics. Neural networks have demonstrated powerful capabilities in nonlinear transformations and classification. Examples of neural network-based image compression algorithms have achieved comparable performance to the general video codec (VVC), a video codec standard developed by the Joint Video Experts Team (JVET) in collaboration with experts from the Moving Picture Experts Group (MPEG) and the Video Codec Experts Group (VCEG). Neural network-based video compression is an actively developing research area, continuously improving the performance of neural image compression. However, due to the inherent difficulties in solving problems with neural networks, neural network-based video codec remains a largely unexplored discipline.

[0045] This discussion now focuses on image and / or video compression. Image / video compression generally refers to a computational technique that compresses video images into binary codecs for easier storage and transmission. Binary codecs may or may not support lossless reconstruction of the original image / video. Codecs that do not lose data are called lossless compression, while those that allow targeted data loss are called lossy compression. Most codec systems use lossy compression because lossless reconstruction is not always necessary. Typically, the performance of an image / video compression algorithm is evaluated based on the resulting compression ratio and reconstruction quality. The compression ratio is directly related to the number of binary codecs produced by compression; fewer binary codecs result in better compression. Reconstruction quality is measured by comparing the reconstructed image / video with the original image / video; higher similarity indicates better reconstruction quality.

[0046] Image / video compression techniques can be categorized into video encoding / decoding methods and neural network-based video compression methods. Video encoding / decoding schemes employ transform-based solutions, where statistical correlations in latent variables (such as discrete cosine transform (DCT) and wavelet coefficients) are carefully hand-designed to entropy encoding / decoding, modeling the correlations in the quantization mechanism. Neural network-based video compression can be further divided into neural network-based encoding / decoding tools and end-to-end neural network-based video compression. The former is embedded as an encoding / decoding tool within existing video codecs, serving only as part of the framework, while the latter is an independent framework developed based on neural networks, independent of the video codec.

[0047] A range of video codec standards have been developed to meet the growing demand for visual content delivery. The International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC) has two expert groups: the Joint Photographic Experts Group (JPEG) and the Moving Picture Experts Group (MPEG). The International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) also has a Video Codecs Expert Group (VCEG), responsible for standardizing image / video codec technologies. Influential video codec standards published by these organizations include JPEG, JPEG 2000, H.262, H.264 / Advanced Video Coding (AVC), and H.265 / High-Efficiency Video Coding (HEVC). The Joint Video Experts Team (JVET), comprised of MPEG and VCEG, developed the Universal Video Codec (VVC) standard. Compared to HEVC, VVC reports an average bitrate reduction of 50% at the same visual quality.

[0048] Neural network-based image / video compression / encoding / decoding is also under development. The example neural network encoding / decoding architectures are relatively shallow, and the performance of such networks is not satisfactory. Neural network-based methods benefit from the support of abundant data and powerful computing resources, thus finding better utilization in a variety of applications. Neural network-based image / video compression has shown promising improvements and has been proven feasible. However, this technology is far from mature and many challenges need to be overcome.

[0049] Neural networks, also known as artificial neural networks (ANNs), are computational models used in machine learning techniques. Neural networks typically consist of multiple processing layers, each composed of several simple but non-linear basic computational units. One advantage of these deep networks is their ability to process data with multiple levels of abstraction and transform data into different types of representations. The representations created by neural networks are not hand-designed. Instead, deep networks containing processing layers learn from large amounts of data using general machine learning procedures. Deep learning eliminates the need for hand-crafted representations. Therefore, deep learning is considered particularly suitable for processing natively unstructured data, such as acoustic and visual signals. Processing such data has been a long-standing challenge in the field of artificial intelligence.

[0050] Neural networks used for image compression can be divided into two categories: pixel probabilistic models and autoencoder models. Pixel probabilistic models employ predictive encoding / decoding strategies. Autoencoder models use transform-based solutions. Sometimes, these two approaches are combined.

[0051] Now we discuss pixel probability modeling. According to Shannon's information theory, the optimal lossless encoding / decoding method achieves the minimum encoding / decoding rate, denoted as -log2p(x), where p(x) is the probability of the symbol x. Arithmetic encoding / decoding is a lossless encoding / decoding method and is considered one of the optimal methods. Given a probability distribution p(x), arithmetic encoding / decoding makes the encoding / decoding rate as close as possible to the theoretical limit -log2p(x) without considering rounding errors. Therefore, the remaining problem is to determine the probabilities, which is very challenging for natural images / videos due to the curse of dimensionality. The curse of dimensionality refers to the problem that increasing dimensionality leads to a sparse dataset; therefore, as dimensionality increases, a rapidly increasing amount of data is needed to effectively analyze and organize the data.

[0052] Following a predictive encoding / decoding strategy, one approach to modeling p(x) is based on previous observations, predicting pixel probabilities one by one in raster scan order, where x is the image, which can be represented as follows:

[0053] p(x) = p(x1)p(x2|x1)…p(x) i |x1,…,x i-1 )…p(x m×n |x1,…,x m×n-1 (1)

[0054] Where m and n are the height and width of the image, respectively. Previous observations are also called the context of the current pixel. When the image is large, estimating the conditional probability can be difficult. Therefore, a simplifying approach is to limit the context of the current pixel as follows:

[0055] p(x) = p(x1)p(x2|x1)…p(x) i |x i-k ,…,x i-1 )…p(x m×n |x m×n-k ,…,x m×n-1 (2)

[0056] Where k is a predefined constant that controls the scope of the context.

[0057] It should be noted that this condition can also consider the sample values ​​of other color components. For example, when encoding and decoding the red (R), green (G), and blue (B) color components (RGB), the R sample depends on previously encoded pixels (including R, G, and / or B samples), and the current G sample can be encoded and decoded based on previously encoded pixels and the current R sample. Furthermore, when encoding and decoding the current B sample, previously encoded pixels as well as the current R and G samples can also be considered.

[0058] Neural networks can be designed for computer vision tasks and can effectively solve regression and classification problems. Therefore, neural networks can be used to solve problems given a context x1, x2, ..., x... i-1 Estimate p(x) in the case of i The probability of a pixel is determined by x. In the example neural network design, the pixel probability is determined by x. i ∈{-1,+1} is used for binary images. The neural autoregressive distribution estimator (NADE) is designed to model pixel probabilities. NADE is a feedforward network with a single hidden layer. In another example, the feedforward network may include connections that skip hidden layers. Furthermore, parameters can be shared. This neural network was used to experiment on the Modified National Institute of Standards and Technology (MNIST) dataset with binarization. In one example, NADE was extended to a real-valued NADE (RNADE) model, where the probability p(x) is given by the model. i |x1,…,x i-1 The Gaussian mixture model is derived from this. The RNADE model's feedforward network also has a single hidden layer, but the hidden layer is rescaled to avoid saturation and uses a rectified linear unit (ReLU) instead of a sigmoid. In another example, NADE and RNADE are improved by reorganizing the order of pixels and using a deeper neural network.

[0059] Designing advanced neural networks plays a crucial role in improving pixel probabilistic modeling. In the example neural network, multidimensional long short-term memory (LSTM) is used. LSTM is combined with a conditional Gaussian scaling mixture for probabilistic modeling. LSTM is a special type of recurrent neural network (RNN) used to model sequential data. Spatial variants of LSTM can also be used for images. Several different neural networks can be employed, including recurrent neural networks (RNNs) and convolutional neural networks (CNNs), such as PixelRNN and PixelCNN. In PixelRNN, two variants of LSTM are used, denoted as row LSTM and diagonal bidirectional LSTM (BiLSTM). Diagonal BiLSTM is specifically designed for images. PixelRNN incorporates residual connections to aid in training deep neural networks with up to 12 layers. In PixelCNN, masked convolutions are used to adjust the shape of the context. PixelRNN and PixelCNN focus more on natural images. For example, PixelRNN and PixelCNN treat pixels as discrete values ​​(e.g., 0, 1, ..., 255) and predict a multinomial distribution over these discrete values. Furthermore, PixelRNN and PixelCNN process color images in the RGB color space. They also perform well on large-scale image datasets like ImageNet. In one example, gated PixelCNN is used to improve upon PixelCNN. Gated PixelCNN achieves comparable performance to PixelRNN but with significantly lower complexity. In another example, PixelCNN++ improves upon PixelCNN by using discrete logistic mixture likelihood instead of a 256-way multinomial distribution; downsampling is used to capture structures of varying precision; additional shortcut connections are introduced to speed up training; dropout is used for regularization; and RGB values ​​are combined into a single pixel. In yet another example, PixelSNAIL combines causal convolution with self-attention.

[0060] Most of the methods described above directly model the probability distribution in the pixel domain. Some designs also model the probability distribution as conditions based on explicit or latent representations. Such models can be expressed as:

[0061]

[0062] Here, h is an additional condition, and p(x) = p(h)p(x|h) indicates that the modeling is divided into unconditional and conditional models. The additional condition can be image label information or a high-level representation.

[0063] Now, let's describe the autoencoder. An autoencoder is trained for dimensionality reduction and consists of an encoding component and a decoding component. The encoding component transforms a high-dimensional input signal into a low-dimensional representation. The low-dimensional representation can have a reduced spatial size but more channels. The decoding component recovers the high-dimensional input from the low-dimensional representation. The autoencoder's ability to automatically learn representations and eliminate the need for hand-crafted features is considered one of the most significant advantages of neural networks.

[0064] Figure 1 This is a schematic diagram illustrating an example transform encoding / decoding scheme 100. The analysis network g... a The original image x is transformed to obtain a latent representation y. The latent representation y is quantized (q) and compressed into bits. The number of bits R is used to measure the encoding / decoding rate. Then, it is synthesized through a network g. s Potential representation of quantization Perform inverse transform to obtain the reconstructed image Distortion (D) in the perceptual space is achieved by using the function g p Transform x and To calculate, we obtain z and Compare them to obtain D.

[0065] Autoencoder networks can be applied to lossy image compression. The learned latent representations can be encoded from a trained neural network. However, adapting an autoencoder to image compression is not easy, as the original autoencoder is not optimized for compression, making it inefficient for direct use as a trained autoencoder. Furthermore, other significant challenges exist. First, the low-dimensional representation should be quantized before encoding. However, quantization is non-differentiable and necessary for backpropagation during neural network training. Second, the objectives differ in compression scenarios, as both distortion and rate are considered simultaneously. Rate estimation is challenging. Third, practical image encoding / decoding schemes should support variable rates, scalability, encoding / decoding speeds, and interoperability. Various solutions are being developed to address these challenges.

[0066] An example autoencoder for image compression using the example transform encoding / decoding scheme 100 can be considered as a transform encoding / decoding strategy. The original image x is processed by the analysis network y = g. a (x) is transformed, where y is the latent representation to be quantized and encoded / decoded. The synthetic network then transforms the quantized latent representation. Inverse transformation is performed to obtain the reconstructed image. The framework uses a rate distortion loss function. Training is performed, where D is the sum of x and y. The distortion between them, R is based on the quantization representation The rate is calculated or estimated, and λ is the Lagrange multiplier. D can be calculated in the pixel domain or the receptive domain. Most example systems follow this prototype, and the differences between such systems may lie only in the network architecture or the loss function.

[0067] In terms of network architecture, RNNs and CNNs are the most widely used architectures. Within the RNN-related category, a general-purpose framework for variable-rate image compression uses an RNN. This example uses binary quantization to generate the code and does not consider the rate during training. This framework provides scalable encoding and decoding capabilities, with RNNs with both convolutional and deconvolutional layers performing well. Another example provides an improved version by compressing binary code using a neural network similar to PixelRNN, thus providing an improved version. On the Kodak image dataset, using multi-scale structural similarity (MS-SSIM) as the evaluation metric, it outperforms JPEG. Another example further improves the RNN-based solution by introducing hidden-state start-up. Furthermore, an SSIM-weighted loss function is designed, and a spatial adaptive bitrate mechanism is included. This example achieves better results than Better Portable Graphics (BPG) on the Kodak image dataset using MS-SSIM as the evaluation metric. Another example system supports spatial adaptive bitrate by training a stopping code-tolerant RNN.

[0068] Another example proposes a general framework for rate-distortion optimized image compression. The example system uses multivariate quantization to generate integer codes and considers the rate during training. The loss is the joint rate-distortion cost, which can be mean square error (MSE) or other metrics. The example system adds random noise during training to stimulate quantization and uses the differential entropy of the noisy codes as a surrogate for the rate. The example system uses generalized divisive normalization (GDN) as the network structure, which includes linear mappings and nonlinear parameter normalization. The effectiveness of GDN for image encoding and decoding is validated. Another example system includes an improved version using three convolutional layers, each followed by a downsampling layer and a GDN layer as the forward transform. Correspondingly, this example version uses a three-layer inverse GDN, each followed by an upsampling layer and a convolutional layer to stimulate the inverse transform. Furthermore, an arithmetic encoding / decoding method is designed to compress integer codes. It is reported that, in terms of MSE, performance on the Kodak dataset is superior to JPEG and JPEG 2000. Another example improves the method by designing a super-prior scaling in the autoencoder. The system will have a subnetwork h a The latent representation y is transformed into z = h a (y), and z is quantized and transmitted as side information. Accordingly, the inverse transform is performed using the subnetwork h. s Implemented, subnetwork h s Quantify the edge information Decoding to quantization The standard deviation, in The standard deviation is further used during arithmetic encoding. On the Kodak image set, this method performs slightly worse than BGP in terms of peak signal-to-noise ratio (PSNR). Another example system further utilizes the structure in the residual space by introducing an autoregressive model to estimate the standard deviation and mean. This example uses a Gaussian mixture model to further remove redundancy in the residuals. On the Kodak image set using PSNR as the evaluation metric, its performance is comparable to VVC.

[0069] The use of neural networks in video compression will now be discussed. Similar to video codecs, neural image compression is the foundation of intra-frame compression in video compression based on neural networks. Therefore, the development of neural network-based video compression technology has lagged behind that of neural network-based image compression, as neural network-based video compression is more complex and requires greater effort to address its challenges. Compared to image compression, video compression requires efficient methods to remove inter-frame image redundancy. Inter-frame image prediction is then a key step in these example systems. Motion estimation and compensation are widely used in video codecs, but are typically not implemented by trained neural networks.

[0070] Neural network-based video compression can be categorized into two types based on the target scenario: random access and low latency. In the case of random access, the system allows decoding to begin at any point in the sequence, typically dividing the entire sequence into multiple individual segments and allowing each segment to be decoded independently. In the case of low latency, the system aims to reduce decoding time, allowing previous frames to be used as reference frames in the temporal domain for decoding subsequent frames.

[0071] Now we discuss low-latency systems. An example system employs a video compression scheme with a trained neural network. The system first divides the video sequence frames into blocks, each block being encoded or decoded according to either an intra-frame or inter-frame encoding / decoding mode. If intra-frame encoding / decoding is selected, the associated autoencoder compresses the block. If inter-frame encoding / decoding is selected, motion estimation and compensation are performed, and residual compression is performed using the trained neural network. The output of the autoencoder is directly quantized and encoded / decoded using the Huffman method.

[0072] Another neural network-based video encoding / decoding scheme employs PixelMotionCNN. Frames are compressed sequentially in the temporal domain, and each frame is divided into blocks, which are compressed in raster scan order. First, each frame is extrapolated using the first two reconstructed frames. When compressing a block, the extrapolated frame, along with the context of the current block, is input into PixelMotionCNN to derive the latent representation. The residuals are then compressed using a variable-rate imaging scheme. This scheme achieves performance comparable to H.264.

[0073] Another example system employs an end-to-end neural network-based video compression framework, where all modules are implemented using neural networks. This scheme accepts the current frame and a previously reconstructed frame as input. A pre-trained neural network is used to derive optical flow as motion information. The motion information is warped along with a reference frame, and then the neural network generates motion-compensated frames. Two independent neural autoencoders are used to compress the residual and motion information. The entire framework is trained using a single rate-distortion loss function. The example system achieves better performance than H.264.

[0074] Another example system employs an advanced neural network-based video compression scheme. This system inherits from and extends neural network video encoding / decoding schemes with the following key features: First, the system uses only one autoencoder to compress motion information and residuals. Second, the system utilizes multi-frame and multi-optical flow motion compensation. Third, the system employs an online state that is learned and propagated over time through subsequent frames. This scheme achieves better performance than the HEVC reference software in MS-SSIM.

[0075] Another example system uses a video compression framework based on an extended end-to-end neural network. In this example, multiple frames are used as references. This allows the example system to provide more accurate predictions of the current frame by using multiple reference frames and associated motion information. Furthermore, motion field prediction is configured to remove motion redundancy along the temporal channels. A post-processing network is also used to remove reconstruction artifacts from previous processes. This system outperforms H.265 by a significant margin in both PSNR and MS-SSIM.

[0076] Another example system replaces optical flow with scale-space flow by adding a frame-based scale parameter. This example system achieves better performance than H.264. Yet another example system uses a multi-precision representation based on optical flow. Specifically, the motion estimation network generates multiple optical flows with different precisions, and the network learns which one to choose under a loss function. Its performance is slightly better than H.265.

[0077] We now discuss systems employing random access. Another example system uses a neural network-based video compression scheme with frame interpolation. Keyframes are first compressed using a neural image compressor, and the remaining frames are compressed hierarchically. The system performs motion compensation in the perceptual domain by deriving feature maps at multiple spatial scales of the original frames and warping these feature maps using motion. The result is used in the image compressor. This method is comparable to H.264.

[0078] One example system uses an interpolation-based video compression method. The interpolation model combines motion information compression and image synthesis. The same autoencoder is used for both the images and the residuals. Another example system employs a neural network-based video compression method based on a variational autoencoder with a deterministic encoder. Specifically, this model includes an autoencoder and an autoregressive prior. Unlike previous methods, this system accepts a group of pictures (GOP) as input and incorporates a three-dimensional (3D) autoregressive prior by considering temporal correlations when encoding and decoding the latent representation. This system delivers performance comparable to H.265.

[0079] Now let's discuss the preliminaries. Almost all natural images and / or videos are in digital format. Grayscale digital images can be generated by... It means that, among them, It is a set of pixel values, where m is the image height and n is the image width. For example, This is an example setting, and in this case... Therefore, a pixel can be represented by an 8-bit integer. An uncompressed grayscale digital image has 8 bits per pixel (bpp), while compressed images certainly have fewer bits.

[0080] Color images typically record color information using multiple channels. For example, in the RGB color space, an image can be represented by... This means that three separate channels store red, green, and blue information. Similar to an 8-bit grayscale image, an uncompressed 8-bit RGB image has 24 bpp. Digital images / videos can be represented using different color spaces. Most neural network-based video compression schemes are developed in the RGB color space, while video codecs typically use the YUV color space to represent video sequences. In the YUV color space, an image is decomposed into three channels: luminance (Y), blue difference chrominance (Cb), and red difference chrominance (Cr). Y is the luminance component, and Cb and Cr are the chrominance components. The compression advantage of YUV arises because Cb and Cr are often downsampled for pre-compression, as the human visual system is less sensitive to the chrominance component.

[0081] A color video sequence consists of multiple color images (also called frames) used to record scenes at different timestamps. For example, in the RGB color space, a color video can be represented as X = {x0, x1, ..., x...} t ,…,x T-1}, where T is the number of frames in the video sequence, and If m = 1080 and n = 1920, And if the video has 50 frames per second (fps), then the data rate of this uncompressed video is 1920 × 1080 × 8 × 3 × 50 = 2,488,320,000 bits per second (bps). This results in approximately 2.32 gigabits per second (Gbps), which consumes a lot of storage space and should be compressed before being transmitted over the network.

[0082] For natural images, lossless methods typically achieve compression ratios of around 1.5 to 3, which is significantly lower than the requirements of streaming media. Therefore, lossy compression is employed to achieve better compression ratios, but at the cost of distortion. Distortion can be measured by calculating the mean variance between the original and reconstructed images, for example, based on MSE. For grayscale images, the MSE can be calculated using the following formula.

[0083]

[0084] Accordingly, the quality of the reconstructed image compared to the original image can be measured by the peak signal-to-noise ratio (PSNR):

[0085]

[0086] in, yes The maximum value in the value is, for example, 255 for an 8-bit grayscale image. Other quality evaluation metrics include structural similarity (SSIM) and multi-scale SSIM (MS-SSIM). To compare different lossless compression schemes, compression ratios at a given result rate can be compared, and vice versa. However, to compare different lossy compression methods, the comparison must consider both rate and reconstruction quality. This can be achieved, for example, by calculating the relative rates at several different quality levels and then averaging these rates. The average relative rate is called Bjontegaard's Delta rate (BD rate). Other aspects for evaluating image and / or video encoding / decoding schemes include encoding / decoding complexity, scalability, robustness, etc.

[0087] Figure 2 To illustrate the comparison of compression scheme 200, a scheme using intra-frame prediction methods (e.g., image compression methods) to compress keyframes is compared to a scheme using inter-frame prediction methods (e.g., extrapolative compression methods). These frames are from the HoneyBee sequence in the Ultra Video Group (UVG) dataset. The results show that video content with simple motion but rich texture is more suitable for inter-frame prediction compression.

[0088] The following are exemplary technical problems solved by the disclosed technical solutions. For random access scenarios, the example system uses image compression methods to compress the first and last frames, also known as keyframes. The remaining frames are interpolated from previously reconstructed frames. However, using image compression only on keyframes cannot account for temporal dependencies. Figure 2 As shown, frame extrapolative compression (FEC) and image compression can be used to compress the current frame x. t The results show that the frame extrapolation compression method produces better RD performance than the image compression method. Therefore, sequences with rich texture and simple motion are more suitable for inter-frame prediction compression schemes.

[0089] This document discloses a mechanism for addressing one or more of the aforementioned problems. This disclosure includes a video codec based on an end-to-end neural network. The codec includes multiple compression networks. The codec can receive video comprising multiple frames. The codec can then encode and decode different frames using different compression networks. This can be achieved by selecting a compression network for each frame based on predetermined and / or learned characteristics (e.g., based on motion and / or image texture between frames). In another example, this can be achieved by using multiple image compression networks to encode and decode each frame and selecting the image compression network that provides the best combination of compression and image distortion. In the example, the end-to-end neural network-based video codec includes a bilateral predictive learning compression (BPLC) component and a spatial-temporal adaptive compression (STAC) component. Keyframes can be encoded and decoded by the STAC component, and non-keyframes can be encoded and decoded by the BPLC component. The BPLC component employs a bidirectional predictive network, a residual autoencoder, and a bidirectional predictive residual refinement network to obtain compressed frames. The STAC component may also include a frame extrapolative compression (FEC) branch and an image compression branch. The FEC branch uses a combination of motion estimation (ME), motion vector compression, motion compensation, residual compression, and residual thinning to obtain compressed frames. The image compression branch takes latent representations from the frames and quantizes these latent representations to obtain compressed frames. In one example, frames classified as having rich texture and simple motion are compressed by the FEC branch, while frames with low texture and / or complex motion are compressed by the image compression branch. Such classification can be performed via a neural network and / or by comparison with predetermined and / or learned parameters.

[0090] To address the aforementioned and other issues, methods summarized below are disclosed. These items should be considered as examples to illustrate general concepts and should not be interpreted narrowly. Furthermore, these items can be applied individually or in combination in any way. The techniques described in this disclosure provide a neural network-based video compression method with spatial-temporal adaptation. More specifically, for video units to be encoded / decoded (e.g., frames, pictures, and / or stripes), an extrapolation process (using reconstructed samples from another video unit) or an image compressor (using information from the current video unit) can be selected. In the following discussion, frames, pictures, and images can have the same meaning.

[0091] Example 1

[0092] In one example, the extrapolation process is implemented by incorporating an FEC module into a video codec based on an end-to-end neural network. By including the FEC module, the current video unit can be compressed by either an image compressor or the FEC module. In one example, the FEC module uses at least one previously encoded frame as a reference to predict the current video unit. Indications regarding motion information between the current video unit and the reference frame are signaled in the bitstream or implicitly inferred by the decoder. In one example, the FEC module uses two or more previously encoded frames as references to predict the current video unit. Indications regarding motion information between the current video unit and the reference frames can be signaled in the bitstream or implicitly inferred by the decoder. In one example, the FEC module uses two or more previously encoded frames to extrapolate the current video unit without signaling motion information. In one example, the reference image used in the FEC should precede the current image in the display order.

[0093] Example 2

[0094] In one example, the reconstructed frame with minimum cost (e.g., RD ​​loss) is stored in the reconstructed frame buffer on the encoder side. In one example, an indication of the compression method used (e.g., FEC or an image compressor) is present in the bitstream. In one example, the decoding and / or reconstruction processes depend on the indication. In one example, a video unit can be a keyframe and / or a frame immediately following another keyframe, such as an intra-frame encoded frame or a frame encoded with FEC. In one example, the same or two different end-to-end NN-based methods can be used to encode and decode frames using either an FEC compressor or an image compressor. In one example, an end-to-end NN-based method can be used to compress motion information.

[0095] Example 3

[0096] In another example, a video codec based on an end-to-end trainable neural network is used to provide improved RD performance. This codec includes a Spatial-Temporal Adaptive Compression (STAC) module and a Bilateral Predictive Learning Compression (BPLC) module dedicated to keyframe and non-keyframe compression. The Spatial-Temporal Adaptive Compression includes multiple branches capable of removing spatial and temporal redundancy in keyframes. In one example, STAC may include at least an image compressor and an extrapolation-based compression method. In one example, BPLC may use at least two reference frames, such as a preceding reference frame and a following reference frame (in terms of display order), to interpolate the current frame. In one example, BPLC may use more than two reference frames to predict the current frame. In one example, BPLC may insert frames following a hierarchical order.

[0097] Example 4

[0098] In another example, besides using image compression methods alone to compress keyframes, spatial-temporal adaptive compression can also be used. In one example, for the video sequence to be compressed, an inter-frame predictive compression method is combined with an image compressor to compress keyframes. In one example, previously reconstructed keyframes can be used to predict the current keyframe. In one example, motion information can be used explicitly or implicitly to derive the predicted frame. In one example, a preceding or subsequent reconstructed frame can be used as a reference frame. In one example, multiple reference frames can be used to derive the predicted frame.

[0099] Example 5

[0100] In one example, more than one inter-frame prediction compression method can be employed. In another example, a stream-based method can be used in conjunction with a kernel-based method to derive the predicted frame. The stream-based method can explicitly derive the optical flow and coding. The kernel-based method can derive the predicted frame using a learned kernel convolved with a reference frame, without deriving the optical flow. In one example, forward and backward prediction can be combined in some form to form the final predicted frame. In one example, the choice of inter-frame prediction compression method can be indicated by signaling in the bitstream or derived by the decoder.

[0101] Example 6

[0102] In one example, more than one intra-predictive compression method can be employed. In one example, different image compression networks / implementations can be provided as intra-predictive compression methods. In one example, different block levels can be used together. For example, block-level and frame-level prediction can be combined. In one example, the selection of the intra-predictive compression method can be notified by signaling in the bitstream or deduced by the decoder.

[0103] Example 7

[0104] In one example, more than one motion compression method may be employed. In one example, different motion compression networks and / or implementations may be provided as motion compression methods. In one example, different block levels may be used together. For example, block-level and frame-level prediction may be combined. In one example, the selection of the motion compression method may be notified by signaling in the bitstream or deduced by the decoder.

[0105] Example 8

[0106] In one example, instead of performing multiple compressions, the compression mode decision can be made before encoding and decoding the frame, reducing runtime complexity. In another example, motion information between two specific frames can be used to determine whether an inter-frame prediction scheme should be chosen. For example, motion amplitude below a threshold should be considered when using an inter-frame prediction scheme. In another example, the frame entropy can be used to determine whether an intra-frame prediction scheme should be used to compress the current keyframe. For example, entropy below a threshold should be considered when using an intra-frame prediction scheme. In yet another example, multiple criteria can be combined to determine the compression mode.

[0107] Example 9

[0108] In one example, the extrapolated frame is used as additional information to enhance the residual encoding / decoding. In another example, the extrapolated frame is concatenated with the residual frame and then used as input at both the encoder and decoder. In yet another example, features are first extracted from the interpolated and residual frames, and then fused together at both the encoder and decoder. In yet another example, the extrapolated frame is used only on the decoder side to help improve the quality of the reconstructed residual frame.

[0109] Example 10

[0110] In the above example, a video unit can be an image, strip, slice, sub-image, coding tree unit (CTU) row, CTU, coding tree block (CTB), coding unit (CU), prediction unit (PU), transform unit (TU), coding block (CB), transform block (TB), virtual pipeline data unit (VPDU), region within an image, region within a strip, region within a slice, region within a sub-image, one or more pixels and / or samples within a CTU, or a combination thereof.

[0111] Example 11

[0112] In one example, the end-to-end motion compression network is designed as a motion vector (MV) encoder-decoder, as follows: Figure 3 As shown. In one example, the end-to-end FEC network is designed as follows: Figure 3 As shown.

[0113] An example embodiment will now be described. There are two common types of redundancy in video signals: spatial redundancy and temporal redundancy. Some neural network-based video compression methods apply image compression only to keyframes in random access scenarios. This approach may degrade RD performance by ignoring temporal redundancy. This neural network embodiment uses intra-frame prediction compression and inter-frame prediction compression for keyframes in random access scenarios. The optimal solution is selected based on the RD loss of the reconstructed frame and the original frame from each scheme. An example is provided in the following subsection.

[0114] Figure 3 This is a schematic diagram illustrating an example neural network framework 300. In the neural network framework 300, the original video sequence is divided into groups of pictures (GOPs). Within each GOP, keyframes are compressed using a spatial-temporal adaptive compression (STAC) component that includes an image compression branch 303 and a frame extrapolation compression (FEC) branch 305. The remaining frames are compressed using a bilateral predictive learning compression (BPLC) component 301.

[0115] Figure 3 The symbols used are as follows. The original video sequence (V) is represented as follows. Where, x t This is a frame at time t. Every Nth frame is set as a keyframe, and interpolation is performed on the remaining N-1 frames between two keyframes. These frames are organized into GoPs, with two consecutive GoPs sharing the same boundary frame. In this document, superscripts I, E, and B represent the variables used in branch image compression, frame extrapolation compression, and bidirectional predictive learning compression, respectively, as shown in neural network framework 300. V t and These represent the original and reconstructed motion vector (MV) fields, respectively. and These are prediction frames from the Bi-predictionNet and the motion compensation network (MC Net), respectively. t and These are the original residual and the reconstructed residual of the residual autoencoder, respectively. The final decoding residual after the residual refinement network is... z t m t and y t These represent the latent representations of image compression, MV, and residuals, respectively. and It is the quantized potential value. The final decoded frame is represented as...

[0116] The neural network framework 300 includes a STAC component and a BPLC component 301. The STAC component is used to compress keyframes, while the BPLC component 301 is used to compress the remaining frames between two keyframes. Frames between keyframes can be referred to as non-keyframes and / or interpolated frames. The STAC component includes an image compression branch 303 and an FEC branch 305. For video sequences... Since no reference frame is available, the first frame x0 is compressed using image compression branch 303. Subsequent keyframes are compressed by either of these two branches (image compression branch 303 and / or FEC branch 305). During training, the first frame in the GoP is used to train image compression branch 303, and the last frame is used to train FEC branch 305. During inference, keyframes are compressed twice. The reconstructed frame with the minimum RD loss is selected and stored in the reconstructed frame buffer for future use. After the keyframes are compressed, the remaining frames are compressed using BPLC component 301.

[0117] Now we discuss spatial-temporal adaptation. There are two common types of redundancy in video signals: spatial redundancy and temporal redundancy. In most video codecs, intra-frame coding and inter-frame coding are the two main techniques for removing spatial and temporal redundancy, respectively. Motion and texture are two factors that determine whether inter-frame or intra-frame predictive coding is more suitable for a sequence. Empirically, video content can be categorized based on motion and texture characteristics. Motion can be simple or complex, and texture can be simple or rich. Complex motion has two types and / or sometimes mixes. First, motion information in complex motion may be difficult to estimate using motion estimation (ME) techniques. Second, encoding and decoding motion information for complex motion may require excessive bits, leading to poorer RD performance compared to intra-frame predictive coding. For example, some video sequences may contain only specific motions, such as translation, while other video sequences may contain combinations of multiple motions, such as translation, rotation, and zoom in / out. Similarly, some video sequences contain simple textures, such as mostly sky and / or monochromatic objects, while other video sequences may contain rich textures, such as bushes, grass, and buildings.

[0118] Intra-frame predictive coding and decoding are suitable for video content with simple textures but complex motion, while inter-frame predictive coding and decoding are suitable for video content with simple motions but rich textures. For example... Figure 2As shown, the HoneyBee sequence contains very small motions between frames, but each frame contains rich textures. Introducing the FEC branch 305 with optical flow can effectively remove correlations. Since the image compression branch 303 only requires 21% of the bits, the FEC branch 305 achieves even better reconstruction quality. However, using only inter-frame prediction encoding and decoding is not optimal. Some video sequences have simple textures but very complex motions. In this case, the image compression branch 303 outperforms the FEC branch 305. Therefore, the neural network framework 300 combines the image compression branch 303 and the FEC branch 305 to build STAC components for learning keyframes in the video compression paradigm.

[0119] Accordingly, the STAC component is used for keyframes. The objective of the STAC component is stated as follows:

[0120]

[0121] Where λ is the Lagrange multiplier; and These represent the decoded frames from image compression branch 303 and FEC branch 305, respectively; and These represent the bits used to encode the residual and motion vector, respectively. For image compression, Set to zero. and The derivation can be expressed as:

[0122]

[0123] Where θ and ψ are optimized parameters. A depth autoencoder is used as the image compression network in image compression branch 303. Image compression branch 303 includes an analysis transform component and a synthesis transform component. The analysis transform component starts from keyframe x. t Obtain the latent representation z t The potential representation z t Quantized to obtain the quantized potential value Synthetic transformation component for quantized latent values Perform transformation to obtain image compression decoding frames

[0124] Now we will discuss FEC. This disclosure considers three methods for obtaining the predicted keyframes. These methods are as follows: Figure 3 , Figure 4 and Figure 5 As shown.

[0125] Figure 4This is a schematic diagram illustrating a direct synthesis scheme 400 for acquiring predicted keyframes. The direct synthesis scheme 400 forwards previously decoded keyframes via a CNN. To obtain the prediction frame From the prediction frame Subtract the current frame x from the middle t To obtain residual r t Residual r t The residual from the final decoding is obtained by refining the network transmission. Then predict the frame With the final decoding residual Add them together to get the final decoded frame.

[0126] Figure 5 This is a schematic diagram illustrating a kernel-based scheme 500 for acquiring predicted keyframes. The kernel-based scheme 500 is substantially similar to the direct synthesis scheme 400. However, the kernel-based scheme 500 utilizes multiple previously decoded keyframes. ...to generate prediction frames

[0127] Accordingly, methods for obtaining prediction keyframes include, for example: Figure 3 The flow-based scheme shown in FEC branch 305, such as... Figure 5 The direct synthesis scheme 400 shown and as follows Figure 5 The kernel-based scheme 500 is shown. Among these options, only the flow-based scheme in FEC branch 305 explicitly derives and encodes motion information. In one example, the flow-based solution of FEC branch 305 is used as the FEC model for the following reasons. First, keyframe prediction is more challenging due to larger temporal distances, often accompanied by larger motion. Kernel-based methods struggle to capture larger motions when the kernel size is smaller than the motion. Furthermore, determining the optimal kernel size is not easy, considering the trade-off between model size and performance. Second, accurate motion estimation is a major factor in capturing temporal correlations. The effectiveness of explicitly deriving optical flow has been validated in practice. Figure 3 As shown, FEC branch 305 uses the following main steps to derive the decoded frame. Motion estimation (ME), MV compression, motion compensation (MC), residual compression, and residual refinement.

[0128] In FEC branch 305, motion estimation and compression are implemented using a pre-trained pyramid, warping, cost volume network (PWC-Net) as the ME network. Since PWC-Net is trained using two consecutive frames, this disclosure fine-tunes the PWC-Net on the training data using the first and last frames from the GoP. Due to the scarcity of optical flow labels, the L2 loss between the warped and original frames is directly used to deploy the fine-tuning. An attention model is used as the MV compression autoencoder. In one example, the number of feature maps is set to 128. FEC branch 305 represents the ME Net as h... me And representing the MV autoencoder as Motion estimation and compression are represented as follows:

[0129]

[0130] Now let's discuss motion compensation. The MC network consists of warping layers based on bilateral linear interpolation, followed by a convolutional neural network (CNN) to refine the warped frames. The MC network is represented as h. mc .

[0131]

[0132] Now let's discuss residual compression and refinement. Figure 6 To illustrate the example residual refinement network 600, which can be used to refine the residuals in FEC branch 305, the diagram shows a convolutional layer with a kernel size of 3×3, 64 output channels, and a stride of 1. The convolutional layers use leaky rectified linear units (ReLU) for activation. Each residual block consists of two convolutional layers.

[0133] Back Figure 3 After obtaining the prediction frames from MC Net, a residual autoencoder is used to compress the residuals. The residual autoencoder can share the same architecture as the MV autoencoder, but has 192 channels. To compensate for errors from the previous stage, Figure 6 The residual refinement network 600 in the model is used to enhance the quality of the reconstructed residuals. The final decoded residual is represented as... Output from the residual decoder in FEC branch 305. The residual autoencoder is represented as... The residual refinement network is represented as This process is represented as:

[0134]

[0135] Now we discuss BPLC component 301. While keyframes are compressed by the STAC component, the remaining non-keyframes are compressed by BPLC component 301. BPLC component 301 includes a bidirectional prediction network, a residual autoencoder, and a bidirectional prediction residual refinement network. Frames are interpolated hierarchically to support random access. Based on the current frame x... t Identify two reference frames Where, d t x represents t The temporal distance between the GoP and the corresponding reference frame. When the GoP size N is an integer exponent of 2, d t According to x t The associated time-domain layer is determined as follows:

[0136]

[0137] Wherein, τ(x) t ) is x t The time-domain ID, where k is the index in GOP.

[0138] Video frame interpolation via adaptive separable convolution (SepConv) was used as the bidirectional prediction residual refinement network. The pre-trained model was fine-tuned on the dataset before joint training. Intuitively, different temporal layers should use separate models. However, experiments showed no significant performance improvement was observed using multiple bidirectional prediction networks. Therefore, this example uses a single network for all temporal layers. The bidirectional prediction residual refinement network was fine-tuned using three consecutive frames. The residual autoencoder can share the same architecture as the MV autoencoder with 192 channels. The residual refinement network can also share the same architecture as the extrapolation residual refinement network in FEC branch 305. The bidirectional prediction learning compression component can be described as follows:

[0139]

[0140] Among them, h bp , and These represent the bidirectional predictive residual refinement network, the residual autoencoder, and the bidirectional predictive residual refinement network, respectively.

[0141] Figure 7A block diagram of an example video processing system 4000 that can implement various techniques of this disclosure is shown. Various implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8 or 10-bit multi-component pixel values), or in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.

[0142] System 4000 may include codec component 4004, which may implement the various codec or encoding methods described in this document. Codec component 4004 may reduce the average bit rate of the video from input 4002 to the output of codec component 4004 to produce a codec representation of the video. Therefore, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of codec component 4004 may be stored or transmitted via a communication connection, as shown in component 4006. The stored or communicated bitstream (or codec) representation of the video received at input 4002 may be used by component 4008 to generate pixel values ​​or displayable video sent to display interface 4010. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “codec” operations or tools, it should be understood that codec tools or operations are used by the encoder, and corresponding decoding tools or operations that invert the codec results will be performed by the decoder.

[0143] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), PCI, IDE, etc. The technologies described in this document can be found in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0144] Figure 8This is a block diagram of an example video processing apparatus 4100. Apparatus 4100 can be used to implement one or more methods described in this document. Apparatus 4100 can be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. Apparatus 4100 may include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. Processor 4102 can be configured to implement one or more methods described in this document. Memory 4104 can be used to store data and code for implementing the methods and techniques described herein. Video processing circuitry 4106 can be used to implement some of the techniques described in this document in hardware circuitry. In some embodiments, video processing circuitry 4106 may be at least partially included in processor 4102 (e.g., a graphics coprocessor).

[0145] Figure 9 A flowchart of an example method 4200 for video processing. Method 4200 includes determining, in step 4202, the current video unit to which an end-to-end neural network-based video codec will be applied to the video. The current video unit can be a picture, strip, slice, subpicture, CTU line, CTU, CTB, CU, PU, ​​TU, CB, TB, VPDU, region within a picture, region within a strip, region within a slice, region within a subpicture, one or more pixels and / or samples within a CTU, or a combination thereof. In a specific example, the current video unit is a frame (also referred to as a picture). The video includes keyframes that can be used as random access points and non-keyframes that may not be used as random access points due to dependencies in inter-frame prediction correlation between frames. The end-to-end neural network-based video codec includes multiple components / branches to process different frames. For example, the end-to-end neural network-based video codec may include a STAC component, which includes an FEC branch and an image compression branch. The end-to-end neural network-based video codec may also include a BPLC component. Depending on the nature of the current video unit, the current video unit may be passed to different components / branches.

[0146] In step 4204, a conversion between the current video unit and the video bitstream is performed using a video codec based on an end-to-end neural network. In one example, the conversion includes encoding the current video unit into a bitstream. In another example, the conversion includes decoding the current video unit from the bitstream. For example, when the current video unit is not a keyframe, the BPLC component can perform the conversion on the current video unit. Furthermore, when the current video unit is a keyframe, the STAC component can perform the conversion on the current video unit.

[0147] For example, when the current video unit is a keyframe, it can be transformed by one of the FEC branch and the image compression branch based on its motion information, entropy, or a combination thereof. For instance, the FEC branch can use multiple codec frames as reference frames to predict the current video unit. In one example, an indication of motion information between the reference frames and the current video unit is included in the bitstream. In another example, motion information between the reference frames and the current video unit is not included in the bitstream. For example, when the current video unit is not a keyframe, the BPLC component can interpolate it based on at least one preceding reference frame and at least one following reference frame.

[0148] In some examples, the current video unit is predicted based on preceding reconstruction keyframes, subsequent reconstruction keyframes, motion information, multiple reference frames, or a combination thereof. In this disclosure, a preceding frame is a frame in the video sequence that precedes the current frame, and a subsequent frame is a frame in the video sequence that follows the current frame.

[0149] In some examples, an end-to-end neural network-based video codec selects one of several inter-frame predictive compression networks to apply to the current video unit. Inter-frame predictive compression networks compress video by determining motion between frames. For example, an end-to-end neural network-based video codec may include one or more BPLC components, one or more FEC branches, and one or more image compression branches. The end-to-end neural network-based video codec can send the current video unit to the optimal component / branch. In one example, this can be achieved by encoding and decoding the current video unit on each relevant component / branch and selecting the component / branch that results in the best combination of compression and distortion. In another example, the end-to-end neural network-based video codec may use one or more learned or predefined parameters to classify the current video unit for compression by the appropriate component / branch. In this example, the bitstream includes the selection of one of several inter-frame predictive compression networks. This instructs the decoder which inter-frame predictive compression network should be applied to correctly decode the current video unit.

[0150] In one example, multiple inter-frame predictive compression networks include a flow-based network that encodes and decodes the current video unit by deriving optical flow. Optical flow is a visible pattern of motion of objects, surfaces, and edges in a visual scene caused by the relative motion between the observer and the scene. Optical flow can be estimated by calculating an approximation of the motion field based on the time-varying image intensity.

[0151] In one example, multiple inter-frame prediction compression networks include kernel-based networks that encode and decode the current video unit by convolving a learned kernel with one or more reference frames to derive a prediction frame. The learned kernel is one of a set of kernel functions that have been tuned to this set of kernel functions based on a machine learning algorithm by applying test data. Mathematical convolution operations can be performed to convolve the learned kernel with one or more reference frames to generate a prediction for the current video unit.

[0152] In one example, a video codec based on an end-to-end neural network can encode and decode the current video unit using a combination of block prediction and frame prediction. Block prediction uses samples from a reference block to predict samples in the current block. Frame prediction uses samples from a reference frame to predict samples in the current frame.

[0153] In one example, an end-to-end neural network-based video codec selects one of several motion compression networks to apply to the current video unit. Motion compression is a mechanism that uses machine learning to predict object motion across multiple frames based on spatial-temporal patterns and encodes the object's shape and direction of travel.

[0154] In one example, an end-to-end neural network-based video codec selects one of several intra-frame prediction compression networks to apply to the current video unit. The intra-frame prediction compression network compresses frames based on the similarity between regions within the same frame.

[0155] It should be noted that method 4200 can be implemented in an apparatus for processing video data, the apparatus including a processor and a non-transitory memory having instructions thereon, such as a video encoder 4400, a video decoder 4500, and / or an encoder 4600. In this case, the instructions, when executed by the processor, cause the processor to perform method 4200. Furthermore, method 4200 can be executed by a non-transitory computer-readable medium including a computer program product for use by a video encoding / decoding device. The computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium, such that, when executed by a processor, the instructions cause the video encoding / decoding device to perform method 4200. Also, a non-transitory computer-readable recording medium can store a bitstream of video generated by method 4200 executed by a video processing apparatus. Additionally, method 4200 can be executed by an apparatus for processing video data, the apparatus including a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform method 4200.

[0156] Figure 10This is a block diagram illustrating an example video encoding / decoding system 4300 from which the techniques of this disclosure can be utilized. The video encoding / decoding system 4300 may include a source device 4310 and a target device 4320. The source device 4310 generates encoded video data and may be referred to as a video encoding device. The target device 4320 can decode the encoded video data generated by the source device 4310 and may be referred to as a video decoding device.

[0157] Source device 4310 may include video source 4312, video encoder 4314, and input / output (I / O) interface 4316. Video source 4312 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations thereof. Video data may include one or more pictures. Video encoder 4314 encodes the video data from video source 4312 to produce a bitstream. The bitstream may include bit sequences forming a codec representation of the video data. The bitstream may include codec pictures and associated data. Codec pictures are codec representations of pictures. Associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 4316 may include a modulator / demodulator (modem) and / or transmitter. Encoded video data may be transmitted directly to target device 4320 via network 4330 through I / O interface 4316. Encoded video data may also be stored on storage medium / server 4340 for access by target device 4320.

[0158] Target device 4320 may include I / O interface 4326, video decoder 4324, and display device 4322. I / O interface 4326 may include a receiver and / or a modem. I / O interface 4326 may acquire encoded video data from source device 4310 or storage medium / server 4340. Video decoder 4324 may decode the encoded video data. Display device 4322 may display the decoded video data to a user. Display device 4322 may be integrated with target device 4320, or it may be external to target device 4320, which may be configured to interface with an external display device.

[0159] The video encoder 4314 and the video decoder 4324 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Universal Video Codec (VVC) standard, and other current and / or further standards.

[0160] Figure 11 To illustrate an example of a video encoder 4400, a block diagram is provided. The video encoder 4400 can be... Figure 10The system 4300 shown includes a video encoder 4314. The video encoder 4400 can be configured to perform any or all of the techniques disclosed herein. The video encoder 4400 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 4400. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0161] The functional components of the video encoder 4400 may include a segmentation unit 4401, a prediction unit 4402 (which may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, and an intra-frame prediction unit 4406), a residual generation unit 4407, a transform unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a buffer 4413, and an entropy coding unit 4414.

[0162] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In one example, the prediction unit 4402 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture containing the current video block.

[0163] In addition, some components (such as motion estimation unit 4404 and motion compensation unit 4405) can be highly integrated, but are shown separately in the example of video encoder 4400 for illustrative purposes.

[0164] The segmentation unit 4401 can segment an image into one or more video blocks. The video encoder 4400 and the video decoder 4500 can support various video block sizes.

[0165] The mode selection unit 4403 can select one of the encoding / decoding modes (intra-frame or inter-frame, e.g., based on error results) and provide the resulting intra-frame or inter-frame encoded / decoded block to the residual generation unit 4407 to generate residual block data, and then provide it to the reconstruction unit 4412 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 4403 can select a combination of intra-frame and inter-frame prediction (CIIP) modes, where prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 4403 can also select the precision of the motion vector for the block (e.g., sub-pixel or integer pixel precision).

[0166] To perform inter-frame prediction on the current video block, motion estimation unit 4404 can generate motion information for the current video block by comparing one or more reference frames from buffer 4413 with the current video block. Motion compensation unit 4405 can determine the predicted video block for the current video block based on motion information and decoded samples from images other than those associated with the current video block from buffer 4413.

[0167] The motion estimation unit 4404 and the motion compensation unit 4405 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.

[0168] In some examples, motion estimation unit 4404 can perform unidirectional prediction on the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Then, motion estimation unit 4404 can generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial shift between the current video block and the reference video block. Motion estimation unit 4404 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0169] In other examples, motion estimation unit 4404 can perform bidirectional prediction on the current video block. Motion estimation unit 4404 can search for a reference video block for the current video block in the reference images in list 0, and can also search for another reference video block for the current video block in the reference images in list 1. Motion estimation unit 4404 can then generate a reference index and a motion vector, where the reference index indicates the reference images containing the reference video blocks in lists 0 and 1, and the motion vector indicates the spatial shift between the reference video blocks and the current video block. Motion estimation unit 4404 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.

[0170] In some examples, the motion estimation unit 4404 may output a complete set of motion information for the decoder's decoding process. In some examples, the motion estimation unit 4404 may not output a complete set of motion information for the current video. Instead, the motion estimation unit 4404 may signal the motion information of the current video block by referencing the motion information of another video block. For example, the motion estimation unit 4404 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0171] In one example, the motion estimation unit 4404 may indicate a value in the syntactic structure associated with the current video block that indicates to the video decoder 4500 that the current video block has the same motion information as another video block.

[0172] In another example, motion estimation unit 4404 can identify another video block and motion vector difference (MVD) within the syntactic structure associated with the current video block. The motion vector difference represents the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 4500 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0173] As discussed above, the video encoder 4400 can predictively signal motion vectors. Two examples of predictive signaling notification techniques that can be implemented by the video encoder 4400 include Advanced Motion Vector Prediction (AMVP) and merge pattern signaling notification.

[0174] Intra-prediction unit 4406 can perform intra-prediction on the current video block. When intra-prediction unit 4406 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block may include the predicted video block and various syntax elements.

[0175] The residual generation unit 4407 can generate residual data for the current video block by subtracting the predicted video block from the current video block. The residual data for the current video block may include residual video blocks that correspond to different sample components of the samples in the current video block.

[0176] In other examples, residual data for the current video block may not exist, such as in skip mode, and residual generation unit 4407 may not perform the subtraction operation.

[0177] The transformation (processing) unit 4408 can generate one or more transform coefficient video blocks of the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0178] After the transform unit 4408 generates a transform coefficient video block associated with the current video block, the quantization unit 4409 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0179] The inverse quantization unit 4410 and the inverse transform unit 4411 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 4412 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 4402 to produce a reconstructed video block associated with the current block for storage in the buffer 4413.

[0180] After the video block is reconstructed by the reconstruction unit 4412, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0181] The entropy encoding unit 4414 can receive data from other functional components of the video encoder 4400. When the entropy encoding unit 4414 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bit stream including the entropy-encoded data.

[0182] Figure 12 To illustrate an example block diagram of a video decoder 4500, the video decoder 4500 may be... Figure 10 The system 4300 shown includes a video decoder 4324. The video decoder 4500 can be configured to perform any or all of the techniques disclosed herein. In the example shown, the video decoder 4500 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 4500. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0183] In the example shown, the video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra-frame prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, the video decoder 4500 can perform decoding channels that typically correspond to the encoding passes described with respect to the video encoder 4400.

[0184] The entropy decoding unit 4501 can retrieve the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded video data blocks). The entropy decoding unit 4501 can decode the entropy-coded video data, and based on the entropy-coded video data, the motion compensation unit 4502 can determine motion information including motion vectors, motion vector precision, reference image list index, and other motion information. For example, the motion compensation unit 4502 can determine this information by executing AMVP and merge modes.

[0185] The motion compensation unit 4502 can generate motion compensation blocks and can perform interpolation based on an interpolation filter. The identifier of the interpolation filter used at sub-pixel precision can be included in the syntax element.

[0186] The motion compensation unit 4502 can use the interpolation filter used by the video encoder 4400 during the encoding of the video block to calculate the interpolation of the sub-integer pixels of the reference block. The motion compensation unit 4502 can determine the interpolation filter used by the video encoder 4400 based on the received syntax information and use the interpolation filter to generate the prediction block.

[0187] The motion compensation unit 4502 can use some syntax information to determine the size of the blocks used to encode the frames and / or stripes of the encoded video sequence, segmentation information describing how each macroblock of the picture of the encoded video sequence is segmented, the mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame codec block, and other information for decoding the encoded video sequence.

[0188] Intra-prediction unit 4503 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 4503 performs inverse quantization, i.e., dequantization, on the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 4501. Inverse transform unit 4505 applies an inverse transform.

[0189] The reconstruction unit 4506 can add the residual block to the corresponding prediction block generated by the motion compensation unit 4502 or the intra-frame prediction unit 4503 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove blocky artifacts. The decoded video block is then stored in a buffer 4507, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates the decoded video for presentation on the display device.

[0190] Figure 13This is a schematic diagram of an example encoder 4600. Encoder 4600 is suitable for implementing VVC technology. Encoder 4600 includes three loop filters: a deblocking filter (DF) 4602, a sample adaptive offset (SAO) 4604, and an adaptive loop filter (ALF) 4606. Unlike DF 4602, which uses predefined filters, SAO 4604 and ALF 4606 utilize the original samples of the current image, reducing the mean square error between the original and reconstructed samples by adding an offset and applying a finite impulse response (FIR) filter, respectively. The offset and filter coefficients are notified by the side information signaling of the encoding and decoding. ALF 4606 is located in the last processing stage of each image and can be viewed as a tool attempting to capture and repair artifacts produced in previous stages.

[0191] The encoder 4600 also includes an intra-frame prediction component 4608 and a motion estimation / motion compensation (ME / MC) component 4610 configured to receive input video. The intra-frame prediction component 4608 is configured to perform intra-frame prediction, while the ME / MC component 4610 is configured to perform inter-frame prediction using a reference image obtained from a reference image buffer 4612. Residual blocks from inter-frame or intra-frame prediction are provided to a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are then provided to an entropy encoding / decoding component 4618. The entropy encoding / decoding component 4618 entropy-encodes and decodes the prediction results and quantization transform coefficients, and transmits the same content to a video decoder (not shown). The quantization component output from the quantization component 4616 can be provided to an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. REC component 4624 can output images to DF 4602, SAO 4604 and ALF 4606 for filtering before these images are stored in reference image buffer 4612.

[0192] The following is a list of some preferred solutions.

[0193] The following solutions illustrate examples of the techniques discussed in this disclosure.

[0194] 1. A media data processing method (e.g., Figure 9 The method described in the text (4200) includes: for the conversion between a video comprising video units and a bitstream of video using neural network-based processing, determining, based on rules, which encoding / decoding process in the extrapolation process and the image compressor scheme to use for the current video unit; and performing the conversion according to the determination.

[0195] 2. As described in Solution 1, wherein the extrapolation process is based on frame extrapolation compression FEC.

[0196] 3. The method described in Solution 2, wherein FEC uses N previously encoded and decoded video units to predict the current video unit, where N is a positive integer.

[0197] 4. The method described in Solution 3, wherein N and / or the encoding / decoding process are indicated in the bitstream.

[0198] 5. A video processing method, comprising: performing a conversion between video and video bitstreams using a video processing device, wherein the video processing system includes a spatial-temporal adaptive compression (STAC) module configured to process keyframes of the video and a bilateral predictive learning compression (BPLC) module configured to process non-keyframes of the video, wherein the STAC module includes various techniques for removing spatial or temporal redundancy from the video.

[0199] 6. The method described in Solution 5 further includes: operating the STAC module to implement image compression or decompression techniques and extrapolation-based compression or decompression techniques.

[0200] 7. The method of any one of solutions 5-6 further includes: operating the BPLC module to use N reference frames, where N is an integer greater than 1.

[0201] 8. The method of any one of solutions 5-7 further includes: operating the BPLC module to perform video unit interpolation in a hierarchical manner.

[0202] 9. A video processing method, comprising: performing a conversion between a video and a video bitstream using neural network-based processing; wherein the video includes one or more key video units and one or more non-key video units, wherein the key video units are selectively encoded and decoded using a spatial-temporal adaptive compression tool according to rules.

[0203] 10. The method described in Solution 9, wherein the rule specifies the use of spatial-temporal adaptive compression tools and image processing tools based on rate-distortion minimization criteria.

[0204] 11. The method as described in Solution 9, wherein key video units are predictively encoded and decoded using previously processed keyframes.

[0205] 12. The method described in Solution 9, wherein the spatial-temporal adaptive compression tool includes an inter-frame predictive codec tool, which includes a stream-based codec tool.

[0206] 13. The method described in Solution 9, wherein the spatial-temporal adaptive compression tool includes an inter-frame prediction encoding / decoding tool, which includes a forward prediction tool or a backward prediction tool.

[0207] 14. The method described in Solution 9, wherein the rule is capable of using intra-frame prediction codec tools.

[0208] 15. The method described in Solution 9, wherein the rule is capable of using motion compression codecs.

[0209] 16. The method described in Solution 9, wherein the rule is capable of using multichannel codec tools.

[0210] 17. The method of any one of solutions 1-16, wherein the neural network-based processing includes, for the transformation of the current video unit, using extrapolated video units and residual video units.

[0211] 18. The method of solution 17, wherein the extrapolation video unit is connected to the residual video unit and used as input in neural network processing on the encoder side and decoder side.

[0212] 19. The method of any one of solutions 1-18, wherein the video unit comprises a video picture, a video strip, a video slice, or a video sub-picture.

[0213] 20. The method of any one of solutions 1-18, wherein the video unit includes a codec tree unit CTU row, CTU, codec tree block CTB, codec unit CU, prediction unit PU, transform unit TU, codec block CB, transform block TB, virtual pipeline data unit VPDU, or sample area.

[0214] 21. The method of any one of solutions 1-20, wherein the conversion includes generating a bitstream from the video.

[0215] 22. The method of any one of solutions 1-20, wherein the conversion includes generating video from a bitstream.

[0216] 23. A video decoding apparatus, comprising a processor configured to implement one or more of the methods described in solutions 1-21.

[0217] 24. A video encoding apparatus comprising a processor configured to implement the method described in one or more of solutions 1-21.

[0218] 25. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to perform the method described in any one of solutions 1-22.

[0219] 26. A video processing method, comprising: generating a bitstream according to any one or more of the methods in solutions 1-21, and storing the bitstream on a computer-readable medium.

[0220] 27. A method, apparatus or system described in this disclosure.

[0221] In the solution described in this paper, the encoder conforms to the format rules by generating a codec representation based on those rules. In the solution described in this paper, the decoder can use the format rules to parse the syntax elements in the codec representation, determining the presence or absence of syntax elements according to the format rules to generate the decoded video.

[0222] In this disclosure, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, and vice versa. The bitstream representation of the current video block may, for example, correspond to bits co-occurring or distributed at different locations within the bitstream, as defined by the syntax. For example, macroblocks may be encoded based on the error residual values ​​of the transform and encoding / decoding, and bits in the header and other fields of the bitstream may also be used. Furthermore, during the conversion, the decoder may parse the bitstream based on this determination, knowing that some fields may or may not be present, as described in the solutions above. Similarly, the encoder may determine whether to include or exclude certain syntax fields and generate the codec representation accordingly by including or excluding syntax fields from the codec representation.

[0223] The other solutions, examples, embodiments, modules, and functional operations disclosed in this document can be implemented in digital electronic circuits or computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or combinations thereof. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more computer program instruction modules encoded on a computer-readable medium, executed or controlled by a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of substances that influences machine-readable propagated signals, or a combination thereof. The term "data processing apparatus" encompasses all means, devices, and machines that process data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof. The propagated signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, generated to encode information for transmission to a suitable receiver device.

[0224] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language (including compiled or interpreted languages) and can be deployed in any form, including standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple harmonizing files (e.g., a file storing one or more modules, subroutines, or portions of code). Computer programs can be deployed to execute on a single computer or on multiple computers located in one location or distributed across multiple locations and interconnected via a communication network.

[0225] The processes and logic described in this document can be executed by one or more programmable processors to execute one or more computer programs, thereby performing functions by manipulating input data and producing outputs. The processes and logic can also be executed by dedicated logic circuits, and can be implemented as dedicated logic circuits, such as FPGAs (field-programmable gate arrays) or ASICs (application-specific integrated circuits).

[0226] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors in any kind of digital computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include, or be operatively coupled to, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, to receive data from or transfer data to one or more mass storage devices, or both. However, a computer need not have such means or devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.

[0227] Although this patent document contains numerous details, these details should not be construed as limiting any invention or the scope of the claims, but rather as a description of features that may be specific to particular embodiments of a particular invention. Certain features described in this patent document in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases one or more features from the claimed combination may be removed from the combination, and the claimed combination may involve sub-combinations or variations of sub-combinations.

[0228] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the specific order shown or in a sequential order, or requiring all shown operations to achieve the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.

[0229] Only some implementation methods and examples are described, and other implementation methods, enhancements and variations can be made based on the content described and shown in this patent document.

[0230] When there are no intermediate components other than a line, trace, or other medium between the first and second components, the first component is directly coupled to the second component. When there are intermediate components other than a line, trace, or other medium between the first and second components, the first component is indirectly coupled to the second component. The term "coupling" and its variations include direct coupling and indirect coupling. Unless otherwise stated, the term "approximately" is used to refer to a range including ±10% of the subsequent value.

[0231] While several embodiments have been provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of this disclosure. The present examples are to be considered illustrative rather than restrictive, and are intended to be limited to the details given in this disclosure. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.

[0232] Furthermore, without departing from the scope of this disclosure, the separate or individual technologies, systems, subsystems, and methods described and illustrated in the various embodiments may be combined or integrated with other systems, modules, technologies, or methods. Other items shown or discussed as coupled may be directly connected or indirectly coupled or communicated via interfaces, devices, or intermediate components in an electrical, mechanical, or other manner. Those skilled in the art can identify other examples of changes, substitutions, and modifications, and these changes, substitutions, and modifications may be made without departing from the spirit and scope of this disclosure.

Claims

1. A method for processing video data, comprising: The method determines to apply an end-to-end neural network-based video codec to the current video unit of the video, wherein the end-to-end neural network-based video codec includes a spatial-temporal adaptive compression (STAC) component, and the STAC component includes a frame extrapolation compression (FEC) branch and an image compression branch; and The conversion between the current video unit and the video bitstream is performed by the video codec based on the end-to-end neural network. The FEC branch uses multiple codec frames as reference frames to predict the current video unit, and the bitstream includes an indication of motion information between the reference frames and the current video unit.

2. The method according to claim 1, wherein, Based on the motion information of the current video unit, the entropy of the current video unit, or a combination thereof, the current video unit is allocated for conversion through one of the FEC branch and the image compression branch.

3. The method according to claim 1, wherein, The bitstream does not include motion information between the reference frame and the current video unit.

4. The method according to claim 1, wherein, The video codec based on an end-to-end neural network further includes a bilateral predictive learning compression (BPLC) component, wherein when the current video unit is a keyframe, the STAC component performs the transformation of the current video unit, and wherein when the current video unit is not a keyframe, the BPLC component performs the transformation of the current video unit.

5. The method according to claim 4, wherein, The BPLC component interpolates the current video unit based on at least one preceding reference frame and at least one following reference frame.

6. The method according to claim 1, wherein, The current video unit is predicted based on the preceding reconstructed keyframes, the following reconstructed keyframes, motion information, multiple reference frames, or a combination thereof.

7. The method according to claim 1, wherein, The end-to-end neural network-based video codec selects one of multiple inter-frame predictive compression networks to apply to the current video unit.

8. The method according to claim 7, wherein, The plurality of inter-frame predictive compression networks include a stream-based network, which encodes and decodes the current video unit by deriving optical streams.

9. The method according to claim 8, wherein, The plurality of inter-frame prediction compression networks include a kernel-based network that encodes and decodes the current video unit by convolving a learned kernel with one or more reference frames to derive a prediction frame.

10. The method according to claim 7, wherein, The selection of one of the plurality of inter-frame predictive compression networks is included in the bitstream.

11. The method according to claim 1, wherein, The end-to-end neural network-based video codec selects one of multiple intra-frame predictive compression networks to apply to the current video unit.

12. The method according to claim 1, wherein, The video codec based on an end-to-end neural network uses a combination of block prediction and frame prediction to encode and decode the current video unit.

13. The method according to claim 1, wherein, The end-to-end neural network-based video codec selects one of several motion compression networks to apply to the current video unit.

14. The method according to claim 1, wherein, The conversion includes encoding the current video unit into the bitstream.

15. The method according to claim 1, wherein, The conversion includes decoding the current video unit from the bitstream.

16. An apparatus for processing video data, comprising: processor; as well as It has a non-transitory memory for instructions, wherein the instructions, when executed by the processor, cause the processor to: The method determines to apply an end-to-end neural network-based video codec to the current video unit of the video, wherein the end-to-end neural network-based video codec includes a spatial-temporal adaptive compression (STAC) component, and the STAC component includes a frame extrapolation compression (FEC) branch and an image compression branch; and The conversion between the current video unit and the video bitstream is performed by the video codec based on the end-to-end neural network. The FEC branch uses multiple codec frames as reference frames to predict the current video unit, and the bitstream includes an indication of motion information between the reference frames and the current video unit.

17. The apparatus according to claim 16, wherein, Based on the motion information of the current video unit, the entropy of the current video unit, or a combination thereof, the current video unit is allocated for conversion through one of the FEC branch and the image compression branch.

18. The apparatus according to claim 16, wherein, The video codec based on an end-to-end neural network further includes a bilateral predictive learning compression (BPLC) component, wherein when the current video unit is a keyframe, the STAC component performs the transformation of the current video unit, and wherein when the current video unit is not a keyframe, the BPLC component performs the transformation of the current video unit.

19. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein, The method includes: The method determines to apply an end-to-end neural network-based video codec to the current video unit of the video, wherein the end-to-end neural network-based video codec includes a spatial-temporal adaptive compression (STAC) component, and the STAC component includes a frame extrapolation compression (FEC) branch and an image compression branch; and The bitstream is generated by the video codec based on the end-to-end neural network. The FEC branch uses multiple codec frames as reference frames to predict the current video unit, and the bitstream includes an indication of motion information between the reference frames and the current video unit.

Citation Information

Patent Citations

  • Machine learning based video compression

    US20200053388A1