Rate control for end-to-end neural network-based video compression based on reinforcement learning
By using the gain vector of the reinforcement learning agent in video compression to optimize the encoding and decoding process of video, the problem of difficulty in realizing the highest quality video encoding within the bit budget constraints in the prior art is solved, and efficient and excellent quality video compression effect is achieved.
Patent Information
- Application Number
- CN202380068142.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-23
- Filing Date
- 2023-09-22
- Publication Date
- 2025-05-02
AI Technical Summary
The prior art is difficult to achieve the highest quality video encoding within bit budget constraints in video compression, especially in the context of end-to-end artificial neural network (ANN)-based video compression.
The gain vector based on the reinforcement learning agent is used to determine the number of bits allocated by the video, and the encoding and decoding process of the video is optimized by using a potential representation. The specific steps include parsing the video data to obtain the Ramda value, determining the vector index with the closest target Ramda value, and optimizing the decoding process by vector value interpolation.
It realizes efficient encoding and decoding of video content within the bit budget constraints, improving the quality and efficiency of video compression.
Smart Images

Figure CN119923651A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of U.S. application serial number 63 / 409,271, filed on September 23, 2022, which is incorporated herein by reference in its entirety. Technical Field
[0003] At least one of the present embodiments generally relates to a method or apparatus for compressing images and videos using neural network based tools. Background Art
[0004] The use of novel Artificial Neural Network (ANN) based tools to compress video content is an area of ongoing research by the Joint Video Exploration Team (JVET) between ISO / MPEG and ITU to replace some modules of the latest standard H.266 / VVC and in the longer term to replace the entire structure with an end-to-end autoencoder approach. Encoding video content at the sequence or subsequence level with the highest quality possible within the bit budget constraints in the context of end-to-end ANN based video compression is one goal of such research. Summary of the invention
[0005] At least one of the present embodiments generally relates to a method or apparatus in the context of compressing images and videos using novel artificial neural network (ANN)-based tools. Specifically, an objective of the described embodiments is to encode video content at the highest quality possible within a bit budget constraint at a sequence or subsequence level in the context of end-to-end ANN-based video compression.
[0006] According to a first aspect, a method is provided. The method comprises the steps of: encoding a portion of a video using a determined number of bits; and determining the number of bits allocated to the encoded portion of the video based on a plurality of frames, wherein the determining comprises using a gain vector from a reinforcement learning agent using a latent representation determined from the encoding.
[0007] According to a second aspect, a method is provided. The method comprises the steps of parsing video data to obtain lambda values; determining the index of a vector having a corresponding lambda closest to a target lambda value, wherein the lambda defines a rate-distortion operation point; interpolating vector values between the determined index vector and the vector having consecutive indices using the target lambda value and the lambda values of the vectors between the determined index vector and the vector having consecutive indices; and decoding the video data using the interpolated vectors.
[0008] According to another aspect, a device is provided. The device includes a processor. The processor can be configured to implement the general aspects by executing any one of the methods.
[0009] According to another general aspect of at least one embodiment, there is provided an apparatus comprising: an apparatus according to any of the decoding embodiments; and at least one of: (i) an antenna configured to receive a signal comprising a video block; (ii) a band limiter configured to limit the received signal to a frequency band comprising the video block; and (iii) a display configured to display an output representing the video block.
[0010] According to another general aspect of at least one embodiment, there is provided a non-transitory computer-readable medium comprising data content generated according to any of the described encoding embodiments or variations.
[0011] According to another general aspect of at least one embodiment, there is provided a signal comprising video data generated according to any of the encoding embodiments or variations described.
[0012] According to another general aspect of at least one embodiment, a bitstream is formatted to include data content generated according to any of the described encoding embodiments or variations.
[0013] According to another general aspect of at least one embodiment, there is provided a computer program product comprising instructions, wherein when the program is executed by a computer, the instructions cause the computer to perform any of the decoding embodiments or variations.
[0014] These and other aspects, features and advantages of the general aspects will become apparent from the following detailed description of illustrative embodiments, which is to be read in conjunction with the accompanying drawings.
[0015] According to another general aspect of at least one embodiment, a non-transitory computer-readable medium is provided, which contains data content including instructions for performing any of the encoding or decoding methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A random access structure with an 8-frame GOP is shown.
[0017] Figure 2 The reinforcement learning framework is shown.
[0018] Figure 3 A basic autoencoder for image compression is shown.
[0019] Figure 4 The basic AG-VAE architecture is shown.
[0020] Figure 5 A basic end-to-end NN-based video compression framework is shown.
[0021] Figure 6 A video compression autoencoder with gain is shown.
[0022] Figure 7 The proposed RL-based rate-distortion optimization scheme is shown.
[0023] Figure 8 A system with another video compression model is shown.
[0024] Fig. 9 AG-VAE based on a hyperparameter autoencoder architecture is shown.
[0025] Fig.10 One embodiment of a method for encoding a video using the described embodiments is shown.
[0026] Fig.11 One embodiment of a method for decoding a video using the described embodiments is shown.
[0027] Fig.12 An embodiment of an apparatus for encoding or decoding using the described embodiments is shown.
[0028] Fig.13 A standard general purpose video compression scheme is shown.
[0029] Fig.14 A standard general purpose video decompression scheme is shown.
[0030] Fig.15 A processor-based system for encoding / decoding is shown under generally described aspects. DETAILED DESCRIPTION
[0031] In traditional video compression, a specific temporal structure enables the encoder to select reference frames among previously decoded pictures to optimize the encoding of each frame, while also managing the bit budget for each frame. Certain key frames are used as references to predict other frames. It is then relevant to ensure that they are encoded with the right quality.
[0032] The typical structure used in the broadcast ecosystem is called the random access structure. It consists of a periodic group of pictures (GOP) consisting of a repetitive minimum sequential time frame structure.
[0033] Figure 1shows this structure for the case of an 8-frame GOP. The first frame is an intra frame, or I-frame, meaning it has no dependencies on other frames to be decoded. This can then be used as a random access point from which the decoder can begin decoding the sequence. In broadcast, they are usually spaced out at one second of video intervals, which allows television viewers to switch channels and begin decoding the new channel of their choice without having to wait too long for the video to start showing. However, these frames usually cost a lot more bits to transmit because they are not predicted using previously decoded content. In between I-frames, the other frames are predicted using previously decoded frames. Figure 2 In the structure of , it can be noticed that the coding order is different from the display order. This enables the encoder to predict frames using previously reconstructed pictures from the past and future. Therefore, these frames are called B frames for bidirectional prediction. B0 of each GOP is the first frame to be encoded and it is predicted using the last key frame (I or B0) from the previous GOP, for example, frame 8 in display order is predicted from frame 0. Subsequent frames in coding order can be predicted using past and future frames, as shown by the arrows. Frame B1 can use frames of type I, B0, frame B2 can be predicted by frames I, B0 and B1, etc.
[0034] In order to manage the bit budget between different types of frames, Figure 1 A typical quantization parameter (QP) offset is shown, which corresponds to the default QP structure used in the HEVC / H.265 reference software. In other words, if the nominal QP selected for the sequence is 30, then I frames will be encoded at QP-27, B0 frames will be encoded at 31, and so on.
[0035] As can be seen, the QP assigned to each frame has nothing to do with the video content in the input signal (i.e. texture, motion, etc.), it only depends on the pre-fixed temporal structure. However, this is suboptimal, as better trade-offs can be found when considering compression efficiency for sub-parts of the sequence and the image. For example, if encoding a still sequence, it is better to allocate a larger bit budget to the first I-frame, because after the temporal prediction, the subsequent frames should occupy only few bits.
[0036] Traditional video compression standards can perform prediction to reduce redundancy, use transform and entropy coding to decorrelate the signal, but also degrade the video based on signal fidelity or visual quality to achieve a low bit rate. Compression performance is then evaluated by looking at the size of the bitstream (i.e., the number of bits required) used to store or transmit the video at a given reconstruction quality after decoding. Codecs are then characterized by the quality of the decoded content according to the bit rate.
[0037] To accommodate the so-called rate-distortion tradeoff, a conventional encoder may adjust its quantization parameter QP, which drives the quantization step size of the transmitted data and, in turn, the distortion introduced in the reconstructed content.
[0038] Rate-distortion optimization (RDO) refers to the algorithms used by encoders to minimize the transmitted bits for a given target reconstruction quality. Traditional encoders divide the image into non-overlapping blocks of different shapes and sizes. Then, intra- or inter-frame prediction and transforms are applied to reduce redundancy with previously encoded content. Finally, the transformed prediction residuals are quantized and entropy coded. In this process chain, only the quantization part is lossy. However, the encoder can optimize for bitrate reduction at all levels: choosing block size, prediction, transform, quantization. To do this, the optimization process aims to minimize the Lagrangian criterion.
[0039] J=R+λD,
[0040] Where R represents the rate, i.e. the number of bits required to represent a picture or a group of pictures, and D represents the distortion or quality of the reconstructed picture at the decoder. Lambda is a parameter that defines at which rate-distortion tradeoff or operating point the codec is used.
[0041] Note that D can be measured using different quality metrics such as MSE (mean squared error), SSIM (structural similarity), etc. However, in traditional video codecs, D is constrained to be summable over blocks. In fact, when the encoder decides whether to encode a block directly or split it into smaller blocks, it needs to calculate and compare the rate-distortion tradeoff in both cases, which requires adding the costs of the smaller blocks. This limitation usually forces traditional encoders to use MSE as the base criterion.
[0042] Rate control consists in maximizing the quality of the reconstructed video under a bitrate constraint. Since the quantization parameter can be chosen at the frame or block level, the encoder needs to plan the expected bitrate for the next frame as frames are encoded sequentially. In most applications, the encoder does not have time to perform multiple passes on a sequence or subsequence to make the best decision based on the content. Therefore, the rate control algorithm attempts to estimate the motion and cost of transmitting residual information for future frames based on the current and past reconstructed frames when adjusting the quantization of the current frame. However, most existing and deployed rate control methods rely heavily on empirically designed methods to adapt. Newer machine learning-oriented methods (such as [1]) use machine learning mechanisms to address the QP variation problem, but their operation on QP is still limited and cannot be easily adjusted to the video compression use case.
[0043] Reinforcement Learning for Rate-Distortion Optimization in the Context of Traditional Video Compression
[0044] Google’s DeepMind published a recent study[2] in which they proposed using an adaptation of the reinforcement learning algorithm MuZero to optimize the bit budget of a set of frames by adjusting the quantization parameters of the VP9 standard[3].
[0045] The reference software for VP9 includes a two-pass encoding strategy. The first pass extracts relevant information about the texture to be encoded (called first-pass statistics). These statistics help the second and final pass make smart encoding decisions that consider the entire frame, not just the current block to be encoded. In contrast to the final pass where blocks can have variable sizes from 4×4 to 64×64, the image is divided into 16×16 non-overlapping blocks to extract this information faster.
[0046] The MuZero reinforcement learning agent uses the first-pass statistics of the frame at the current time step, as well as other information (such as compression statistics of frames processed so far, the percentage of the bit budget used so far, etc.) to determine its action, namely, to determine the QP offset to use for the current frame to achieve the best rate-distortion tradeoff.
[0047] The algorithms that these reinforcement learning agents use to learn to take optimal actions (such as MuZero or PPO[4]) work based on a feedback loop. The agent decides what action to take based on the state of the environment. The chosen action is applied to the environment, causing it to transition to a new state. The environment then returns the new state and a reward indicator to the agent. Thus, during training, the agent receives a reward for each action it takes, which indicates how good that action was. The agent then adjusts its internal parameters to maximize the cumulative reward it receives.
[0048] End-to-end video compression for deep learning
[0049] In recent years, new types of image and video compression methods based on artificial neural networks have been developed. In contrast to traditional methods that apply predefined prediction patterns and transformations, ANN-based methods rely on parameters learned on large datasets during training by minimizing a loss function in an iterative manner. In the case of compression, the loss function describes both the bitrate estimate of the encoded bitstream and the performance of the decoded content, just like the Lagrangian criterion J×R+λD described earlier.
[0050] Figure 3 A basic exemplary autoencoder pipeline is shown.
[0051] The input X to the encoder part of the network can be composed as follows:
[0052] 1. An image or frame of a video;
[0053] 2. A part of an image;
[0054] 3. A tensor representing a set of images; or
[0055] 4. A tensor representing a portion (crop) of a set of images.
[0056] In each case, the input can have one or more components, for example: monochrome, RGB, or YCbCr components.
[0057] 1. The input tensor X is fed into the encoder network. The encoder network is usually a sequence of convolutional layers with activation functions. Convolution or spatial to depth 1 Large strides in the operation can be used to reduce the spatial resolution while increasing the number of channels. The encoder network can be viewed as a learned nonlinear transformation.
[0058] 2. The output of the encoder network is a “feature map” or “latent representation” Y. It is quantized as This results in a tensor that is entropy encoded (EC) into a binary stream (bitstream) for storage or transmission.
[0059] 3. The bitstream is entropy decoded (ED) to obtain at the decoder end
[0060] 4. Decoder network from potential representation generate That is, an approximation of the original X tensor. The decoder network is typically a series of upsampling convolutions (e.g., "deconvolution" or convolution followed by an upsampling filter) or depth-to-spatial operations. The decoder network can be viewed as a learned inverse transform, or a denoising and generative transform.
[0061] Note that more complex architectures exist, such as adding a "super autoencoder" (super prior) to the network in order to jointly learn the underlying distribution properties of the encoder output. The embodiments presented here are not limited to the use of autoencoders. Any end-to-end differentiable codec can be considered.
[0062] As described above, in the following, Represents the quantized version of X.
[0063] For a long time, state-of-the-art models were trained for each Lagrangian parameter λ, so multiple pre-trained models were needed to evaluate performance over a range of bitrates. Each training set corresponds to millions of parameters, which makes it impossible to use such codecs in real applications. For example, the granularity of H.265 / HEVC is 51 QPs. Adapting to a specific bitrate at such a granularity would result in the decoder having 51 pre-trained models in memory and being able to implement switching for rate control.
[0064] To address memory requirements and avoid having to switch models, the authors of [5] proposed AG-VAE (Asymmetric Gain Variational Autoencoder). Below, we consider a latent tensor Y with dimensions C×H×W, where H and W represent the height and width of the tensor, which are usually a fraction of the resolution of the input X, and C represents the number of channels. Before quantization, this latent tensor is element-wise multiplied by a gain vector of shape C×1×1, i.e., the element of the i-th channel in Y is multiplied by G e At the decoder, the entropy decoded (ED) tensor is also multiplied by the inverse gain vector G d , and then fed into the synthesis function g s To reconstruct the output Note that we use a basic version of this gain unit as an example in the following, but the proposed embodiments can be extended to more improved mechanisms, such as spatially adaptive gains instead of one value per channel.
[0065] The model does not require a full set of training parameters for each operating point (i.e., each lambda), but only requires the vector G e and its corresponding G d At training time, these vectors can be randomly selected so that g is trained for all selected lambdas a and g s , while each vector is optimized for only one lambda. At inference time, i.e. when the codec is actually used, the appropriate gain vectors can be selected from the list, or even interpolated to optimize the target bitrate.
[0066] In the above description, for simplicity, we decompose all transmitted information into tensors, as if the input always consists of only the image X. However, in the case of video compression, different tensors may be computed and transmitted depending on the frame type considered. In the following, we describe an exemplary video model that specifies different models for intra (I) pictures and predicted (P) pictures.
[0067] Figure 5 A video compression framework based on a state-of-the-art published model is shown. It uses two separate architectures to process I-frames and P-frames. I-frames are processed similarly to image compression, i.e. using the same process as the model described previously. P-frames can be further compressed using reconstructed information (i.e. past decoded frames). A warper is used to map a reference picture onto the current picture for encoding to form a predictor. The warper uses a motion model The model uses an autoencoder similar to those used for image compression for estimation, compression, and transmission, except that the input is the concatenation along the channel axis of the reference and current pictures, e.g., a 6-channel tensor in the case of a 3-channel RGB input frame.
[0068] Then, as in conventional compression, a residual is formed as the difference between the predictor and the source signal. This residual is compressed and transmitted in the same way as I-frames, i.e. using a dedicated autoencoder such as Figure 5 shown.
[0069] Figure 6 An exemplary gain-based version of the above video compression pipeline is shown. Each autoencoder now contains a gain unit at both the encoder and decoder ends. For a specific lambda of an I frame, the model uses a pair of (G e,I ,G d,I ), just as for images, where I marks the vector corresponding to an I frame. However, P frames now use a 4-tuple (G e,P,F ,G d,P,F ,G e,P,R ,G d,P,R ), where F and R represent motion flow and residual, respectively. Note that this method can be achieved by having a correlation gain (G e,B,F ,G d,B,F ,G e,B,R ,G d,B,R ) is easily extended to bidirectional prediction.
[0070] Embodiments described herein aim to optimize compression of a bitstream by efficiently allocating a bit budget to groups of pictures or subsequences of an input video.
[0071] Reference encoders typically organize the processing of video sequences in terms of Group of Pictures (GOP), where pictures can be encoded dependent on previously reconstructed pictures, following a predefined timing structure. This structure is often used to empirically assign a quantization parameter (QP) to each frame based on its position in the GOP and the interdependencies between frames.
[0072] Professional encoders build on this to develop rate control methods that extract features of the video content to derive a QP policy that meets a bit budget target. These methods often use engineering approaches that may still be suboptimal. The aforementioned DeepMind work addresses this problem by using a reinforcement learning (RL) based mechanism in the context of traditional video coding and relying on information computed from fast past encoder estimates.
[0073] The method proposed in this specification also exploits the advantages of RL algorithms when optimizing compression under bit budget constraints. However, it can be run on a novel end-to-end NN video compression system. It provides a means to adjust the bit rate from the encoder and the syntax and mechanism to pass the information to the decoder.
[0074] In the proposed embodiment, reinforcement learning is proposed to optimize the compression of a sequence of frames under a bit budget constraint. We propose to exploit the structure of certain compression autoencoders to make the RL algorithm dependent on relevant data.
[0075] Specifically, some variants of the proposed solution exploit the AG-VAE architecture described in Section 1, where the analysis transformation g at the encoder a The gain is then applied. This means that in this encoder architecture, for any lambda that the encoder is running, the latent tensor Y = g a (X) is always the same. This is because the rate control mechanism involves multiplication by a lambda-dependent gain vector and quantization, which happens after Y has been computed. Therefore, the latent tensor Y provides the RL agent with meaningful information about the compressibility of the input. In contrast, the DeepMind work relies on first-pass encoding, which considers a limited number of encoding choices to estimate the bit cost of a frame under a given encoder setting.
[0076] Main general methods
[0077] The main idea of the described embodiment is to optimize the rate-distortion tradeoff of a GoP (a sub-part of a video) by selecting the amount of bit allocation on a frame-by-frame basis. This can be done under a target bitrate or limit bitrate constraint for that GoP. At each frame, the RL agent decides how much of the bit budget to allocate. This corresponds to selecting a tradeoff point in the rate-distortion curve. Among other possibilities, one strategy for changing the rate-distortion tradeoff of a neural network based compression model is to use the AG-VAE architecture explained in Section 1.4. The following is a proposed approach using AG-VAE as a means to optimize the rate-distortion tradeoff. Note that this system can be used with any other rate control method.
[0078] Frames from the GoP are input into the system one at a time. For any given frame, the RL agent takes as input the encoded latent representation of the current frame, along with other information such as the bit budget used so far. The agent then outputs its decision about how many bits to allocate for the current frame. Using this decision, the rate control mechanism encodes the frame at the selected bit rate. In this case, the rate control mechanism selects the gain and inverse gain vectors for the target bit rate of the frame.
[0079] Internally, the proxy uses Figure 7The policy network shown at the bottom (in this case a deep convolutional neural network) decides the frame bit allocation. The agent is rewarded based on how well it optimizes the rate-distortion tradeoff for the GoP. For example, this reward can be calculated based on the total distortion of the frame sequence, with a penalty for exceeding the bit budget. Figure 7 In the example of , distortion is measured using PSNR and bitrate is expressed in bpp (bits per pixel). This reward is used to compute the policy network gradients needed to improve the agent's decision making. The RL training algorithm iteratively modifies the parameters of the policy network to maximize the reward it obtains for each GoP.
[0080] Different state spaces and action spaces
[0081] The input to an RL agent is the state of the environment. The set of all possible states that the environment can be in is called the state space. A well-designed state space for an RL system should contain all the information the agent needs to decide its actions.
[0082] Similar to the state space, the set of all possible actions that an RL agent can take in the environment is called the action space.
[0083] For our proposed method of rate-distortion optimization using RL at the GoP level, we can have multiple formulations of the state space and action space based on the specific case of the video compression model. For example, Figure 7 The state space of the system shown in is the coded latent representation of the frame plus the bit budget used so far in the GoP. Some systems use frames directly instead of coded latent representations, e.g. Figure 8 It also has another additional state-space feature that indicates the type of frame being processed. Another possible state-space representation is a concatenated tensor with all possible products of the gain vector and the encoded latent representation.
[0084] Similarly, the action space can be formulated in several different ways. One possibility is to have a discrete action space, where the agent simply chooses between predetermined bitrate points. Another possibility is to have a continuous action space, where the agent chooses from a range of bitrate points. This action space can be suitable for systems with rate control mechanisms that support continuous rate control, such as AG-VAE.
[0085] The system variants described in the following sections are illustrated using the specific state space and action space described above. However, the system is not limited to use with this specific example and is compatible with a variety of other state and action spaces.
[0086] Simple variant when processing video in still picture mode
[0087] In one embodiment, we consider the case where the video is processed in pure intra mode, that is, each picture is encoded independently. The above description can be directly applied.
[0088] Adapting the proposed method on top of the video architecture
[0089] As mentioned before, in most video compression models, different types of frames (e.g., I-frames and P-frames) are handled differently. Our proposed method is also compatible with such video compression models, such as Figure 8 shown.
[0090] The RL agent can now take as additional input the type of frame it is processing and its action space can be discrete, i.e. if it is an I frame, then the gain vector (G e,I ,G d,I ), if it is a P frame, then the gain vector (G e,P,F ,G d,P,F ,G e,P,R ,G d,P,R ), where F represents motion flow and R represents residual. Another possibility is to have a continuous action space.
[0091] Adapting the proposed method on top of the Hyperprior architecture
[0092] The above description of AG-VAE uses the most basic autoencoder chain, where we did not elaborate on the entropy bottleneck and the entropy encoder and decoder (EC / ED in the figure above). Improved methods use hyper-prior models to estimate the latent tensor The statistical distribution of each element of . In this case, the second tensor Z is also quantized and entropy coded. The exact same gain unit mechanism can be applied to the super-prior tensor Z, where G e h , G d h With dimension c h 11, where c h Indicates the number of channels of Z. Figure 4 Describes an architecture where By h a Analysis, it is The entropy encoder provides the entropy model parameters for each element. This side information also needs to be transmitted in the bitstream so that the super-prior decoder can use h s Reconstruct the entropy model to reconstruct
[0093] Each operating point lambda is now represented by a four-tuple {G e ,G e h ,Gd ,G d h Therefore, it does not change the indices and information to be transmitted to the decoder, which contains a list of gains of the same size as the encoder and can do the interpolation of the values of the two vectors in the same way.
[0094] Suggested Standard Syntax
[0095] In the case of traditional video compression, the quantization parameter (QP) is usually transmitted per frame or per frame slice. Note that a QP offset can be transmitted at the block level to adjust the target quality spatially. One symbol is transmitted per frame, encoding a value in the range 0-51, and the QP offset mechanism can even reduce the cost of this element.
[0096] There is no interpolation between gain vectors
[0097] In the case of the AG-VAE architecture without gain interpolation, a pair of vectors G can be selected from the list of n vectors used to train the model e and G d This would require the bitstream to contain the index of the selected pair among the n vectors, i.e. a relatively small positive integer, e.g. less than 256, and thus encoded in 8 bits. If the frame is partitioned and processed tile by tile using a corresponding autoencoder, this syntax element can be encoded per frame or per frame tile.
[0098] Interpolation between 2 pre-trained vectors
[0099] If there are fewer predefined gain vectors corresponding to lambda points, the target lambda value can be used to interpolate between the two vectors. To interpolate between two existing pre-trained vectors available at the encoder and decoder, the decoder needs the information required to reconstruct the appropriate gain vector.
[0100] In the first embodiment, we assume that the decoder has a mechanism to parse the lambda transmitted with the bitstream. Each predefined vector is associated with the lambda value it was trained for. The decoder then determines the index of the vector corresponding to the lower lambda and closest to the target lambda. It can then interpolate the vector values between the latter vector and the vector corresponding to the consecutive index number using the target lambda and the lambda values of the two surrounding vectors.
[0101] The lambda value can then be transmitted directly in the bitstream. Since lambda is usually a real number, it needs to be rounded down to the nearest value with a predetermined precision. Since codecs usually avoid sharing floating point numbers between encoder and decoder, the transmitted lambda should be represented as a fixed point value.
[0102] In the second variant, it is proposed to transfer 2 variables to specify the target vector:
[0103] -The index of the pre-trained vector whose lambda is closest to and smaller than the target lambda;
[0104] - Consider the interpolation ratio between two consecutive pre-trained vectors for the operation. The value of this interpolation ratio is between 0 and 1 and can be easily represented as an unsigned integer depending on the selected interpolation precision.
[0105] Since the interpolation ratios are transmitted directly and do not need to be calculated, this pair of syntax elements can be more efficient in terms of the number of bits required for transmission, and the interpolation operation at the decoder is also slightly less complex.
[0106] Fig.10 One embodiment of a method 1000 for encoding video data is shown. The method starts at start block 1001 and proceeds to block 1010 to encode a portion of a video using a determined number of bits. Control proceeds from block 1010 to block 1020 to determine the number of bits allocated to the encoded portion of the video based on a plurality of frames, wherein the determining includes using a gain vector from a reinforcement learning agent using a latent representation determined from the encoding.
[0107] Fig.11 One embodiment of a method 1100 for decoding video data is shown. The method starts at start block 1101 and proceeds to block 1110 to parse the video data to obtain lambda values. Control proceeds from block 1110 to block 1120 to determine the index of a vector having a corresponding lambda closest to a target lambda value, wherein lambda defines a rate-distortion operation point. Control proceeds from block 1120 to block 1130 to interpolate vector values between the determined index vector and the vectors having consecutive indices using the target lambda value and the lambda values of the vectors between the determined index vector and the vectors having consecutive indices. Control proceeds from block 1130 to block 1140 to decode the video data using the interpolated vectors.
[0108] Fig.12 An embodiment of an apparatus 1200 for compressing, encoding or decoding a video using the above method is shown. The apparatus includes a processor 1210, which may be interconnected with a memory 1220 via at least one port. The processor 1210 and the memory 1220 may also have one or more additional interconnections with external connections.
[0109] The processor 1210 is also configured to insert or receive information in a bit stream and perform compression, encoding or decoding using the above-mentioned methods.
[0110] The embodiments described herein include various aspects, including tools, features, embodiments, models, methods, etc. Many of these aspects are described as having specificity, and at least in order to show individual features, are usually described in a manner that may sound restrictive. However, this is only for the purpose of describing clearly and does not limit the application or scope of these aspects. In fact, all different aspects can be combined and interchanged to provide further aspects. In addition, these aspects can also be combined and interchanged with aspects described in previous applications.
[0111] The aspects described and contemplated in this application may be implemented in a variety of different forms. Fig.13 , Fig.14 and Fig.15 Some embodiments are provided, but other embodiments are contemplated, and Fig.13 , Fig.14 and Fig.15 The discussion does not limit the breadth of implementation. At least one aspect generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a generated or encoded bitstream. These and other aspects can be implemented as methods, apparatus, computer-readable storage media storing instructions for encoding or decoding video according to any of the methods described, and / or computer-readable storage media storing a bitstream generated according to any of the methods described.
[0112] In this application, the terms "reconstruction" and "decoding" are used interchangeably, the terms "pixel" and "sample" are used interchangeably, and the terms "image", "picture" and "frame" are used interchangeably. Usually, but not necessarily, the term "reconstruction" is used on the encoder side, while "decoding" or "reconstruction" is used on the decoder side.
[0113] Various methods are described herein, and each method includes one or more steps or actions for implementing the method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions can be modified or combined. In addition, terms such as "first", "second" and the like can be used in various embodiments to modify elements, components, steps, operations, etc., such as "first decoding" and "second decoding". Unless specifically required, the use of these terms does not mean that the modified operations are sorted. Therefore, in this example, the first decoding does not have to be performed before the second decoding, and for example, can occur before the second decoding, during the second decoding, or in a time period overlapping with the second decoding.
[0114] Various methods and other aspects described in this application can be used to modify modules, e.g. Fig.13 and Fig.14The intra prediction module, entropy coding module and / or decoding module (160, 360, 145, 330) of the video encoder 100 and decoder 200 shown. In addition, the present aspects are not limited to VVC or HEVC, and can be applied to, for example, other standards and recommendations (whether existing or developed in the future) and extensions of any such standards and recommendations (including VVC and HEVC). Unless otherwise specified or technically excluded, the aspects described in this application can be used alone or in combination.
[0115] Various numerical values are used in this application. Specific values are used for illustrative purposes, and the aspects are not limited to these specific values.
[0116] Fig.13 An encoder 100 is shown. Variations of the encoder 100 are contemplated, but for clarity, only the encoder 100 is described below, rather than all contemplated variations.
[0117] Prior to encoding, the video sequence may undergo pre-encoding processing (101), such as applying a color transform to an input color picture (e.g., converting from RGB 4:4:4 to YCbCr 4:2:0), or performing a remapping of input picture components to obtain a signal distribution more amenable to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with the pre-processing and appended to the bitstream.
[0118] In encoder 100, a picture is encoded by encoder elements, as described below. The picture to be encoded is divided (102) and processed, for example, in units of CUs. Each unit is encoded, for example, using intra or inter mode. When the unit is encoded in intra mode, it performs intra prediction (160). In inter mode, motion estimation (175) and compensation (170) are performed. The encoder decides (105) which of intra mode or inter mode to use to encode the unit, and indicates the intra / inter decision, for example, by a prediction mode flag. For example, the prediction residual is calculated by subtracting (110) the predicted block from the original image block.
[0119] The prediction residual is then transformed (125) and quantized (130). The quantized transform coefficients, along with motion vectors and other syntax elements, are entropy encoded (145) to output a bitstream. The encoder may skip the transform and apply quantization directly to the untransformed residual signal. The encoder may bypass the transform and quantization, i.e., encode the residual directly without applying a transform or quantization process.
[0120] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are inverse quantized (140) and inverse transformed (150) to decode the prediction residual. The decoded prediction residual and the predicted block are combined (155) to reconstruct the image block. A loop filter (165) is applied to the reconstructed picture to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce coding artifacts. The filtered image is stored in a reference image buffer (180).
[0121] Fig.14 2 shows a block diagram of a video decoder 200. In the decoder 200, the bitstream is decoded by decoder elements as described below. The video decoder 200 generally performs a decoding process that is the reverse of the encoding process, such as Fig.13 Encoder 100 also typically performs video decoding as part of encoding the video data.
[0122] Specifically, the input to the decoder includes a video bitstream, which may be generated by the video encoder 100. The bitstream is first entropy decoded (230) to obtain transform coefficients, motion vectors, and other encoding information. Picture partition information indicates how the picture is partitioned. Thus, the decoder may segment (235) the picture according to the decoded picture partition information. The transform coefficients are inversely quantized (240) and inversely transformed (250) to decode the prediction residual. The decoded prediction residual and the predicted block are combined (255) to reconstruct the image block. The predicted block may be obtained (270) by intra-frame prediction (260) or motion compensated prediction (i.e., inter-frame prediction) (275). An in-loop filter (265) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (280).
[0123] The decoded picture may be further subjected to post-decoding processing (285), such as an inverse color transform (e.g., from YCbCr 4:2:0 to RGB 4:4:4) or an inverse remapping that performs the inverse of the remapping process performed in the pre-encoding process (101). The post-decoding processing may use metadata derived in the pre-encoding process and signaled in the bitstream.
[0124] Fig.15A block diagram of an example of a system implementing various aspects and embodiments is shown. System 1000 may be embodied as a device including various components described below, and is configured to perform one or more aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, notebook computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, networked home appliances, and servers. The elements of system 1000 may be embodied in a single integrated circuit (IC), multiple ICs, and / or discrete components, either individually or in combination. For example, in at least one embodiment, the processing and encoder / decoder elements of system 1000 are distributed over multiple ICs and / or discrete components. In various embodiments, system 1000 is coupled to one or more other systems or other electronic devices in a communication manner, such as by a communication bus or by a dedicated input and / or output port. In various embodiments, system 1000 is configured to implement one or more aspects described in this document.
[0125] The system 1000 includes at least one processor 1010 configured to execute instructions loaded therein, for example, to implement the various aspects described in this document. The processor 1010 may include embedded memory, input-output interfaces, and various other circuits known in the art. The system 1000 includes at least one memory 1020 (e.g., a volatile memory device and / or a non-volatile memory device). The system 1000 includes a storage device 1040, which may include a non-volatile memory and / or a volatile memory, including but not limited to an electrically erasable programmable read-only memory (EEPROM), a read-only memory (ROM), a programmable read-only memory (PROM), a random access memory (RAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), a flash memory, a magnetic disk drive, and / or an optical disk drive. As non-limiting examples, the storage device 1040 may include an internal storage device, an additional storage device (including removable and non-removable storage devices), and / or a network accessible storage device.
[0126] The system 1000 includes an encoder / decoder module 1030, which is configured to process data to provide encoded video or decoded video, for example, and the encoder / decoder module 1030 may include its own processor and memory. The encoder / decoder module 1030 represents a module that can be included in a device to perform encoding and / or decoding functions. As is well known, a device may include one or both of the encoding and decoding modules. In addition, the encoder / decoder module 1030 may be implemented as a separate element of the system 1000, or may be incorporated into the processor 1010 as a combination of hardware and software known to those skilled in the art.
[0127] Program code to be loaded onto the processor 1010 or the encoder / decoder 1030 to perform various aspects described in this document may be stored in the storage device 1040 and subsequently loaded onto the memory 1020 for execution by the processor 1010. According to various embodiments, one or more of the processor 1010, the memory 1020, the storage device 1040, and the encoder / decoder module 1030 may store one or more of a variety of items during the execution of the processes described in this document. These stored items may include, but are not limited to, input video, decoded video or a portion of decoded video, bitstreams, matrices, variables, and intermediate or final results from equations, formulas, operations, and arithmetic logic processing.
[0128] In some embodiments, memory internal to the processor 1010 and / or the encoder / decoder module 1030 is used to store instructions and provide processing working memory required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory may be memory 1020 and / or a storage device 1040, such as a dynamic volatile memory and / or a non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store, for example, an operating system for a television. In at least one embodiment, a fast external dynamic volatile memory (e.g., RAM) is used as working memory for video encoding and decoding operations (e.g., MPEG-2 (MPEG refers to Moving Picture Experts Group, MPEG-2 is also known as ISO / IEC 13818, 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, a new standard being developed by the Joint Video Experts Group JVET)).
[0129] Input to the elements of system 1000 may be provided through a variety of input devices shown in block 1130. Such input devices include, but are not limited to: (i) a radio frequency (RF) section that receives an RF signal transmitted over the air by, for example, a broadcaster, (ii) a component (COMP) input terminal (or a set of COMP input terminals), (iii) a universal serial bus (USB) input terminal; and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Fig.15 Other examples not shown include composite video.
[0130] In various embodiments, the input device of block 1130 has associated corresponding input processing elements known in the art. For example, the RF portion may be associated with elements adapted to (i) select a desired frequency (also referred to as selecting a signal or bandwidth limiting a signal to a frequency band), (ii) down-convert the selected signal, (iii) again bandwidth limit a narrower frequency band to select (for example) a signal frequency band (which may be referred to as a channel in some embodiments), (iv) demodulate the down-converted and bandwidth limited signal, (v) perform error correction, and (vi) demultiplex to select a desired packet stream. The RF portion of various embodiments includes one or more elements for performing these functions, such as a frequency selector, a signal selector, a bandwidth limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF portion may include a tuner that performs a variety of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near baseband frequency) or baseband. In a set-top box embodiment, the RF part and its related input processing element receive the RF signal transmitted by wired (for example, cable) medium, and filter to the desired frequency band again by filtering, down-conversion and perform frequency selection. Various embodiments rearrange the order of above-mentioned (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding element can include inserting element between existing element, for example inserting amplifier and analog-to-digital converter. In various embodiments, the RF part includes antenna.
[0131] In addition, the USB and / or HDMI terminals may include respective interface processors for connecting the system 1000 to other electronic devices via USB and / or HDMI connections. It should be appreciated that aspects of input processing (e.g., Reed-Solomon error correction) may be implemented in a separate input processing IC or in the processor 1010 as desired. Similarly, aspects of USB or HDMI interface processing may be implemented in a separate interface IC or in the processor 1010 as desired. The demodulated, error corrected, and demultiplexed streams are provided to a variety of processing elements, including, for example, the processor 1010 and the encoder / decoder 1030, which operate in conjunction with memory and storage elements to process the data streams as desired for presentation on an output device.
[0132] The various elements of system 1000 may be disposed within an integrated housing. Within the integrated housing, the various elements may be interconnected and transmit data between each other using suitable connection means (e.g., internal buses known in the art, including inter-IC (I2C) buses, wiring, and printed circuit boards).
[0133] The system 1000 includes a communication interface 1050 that enables communication with other devices through a communication channel 1060. The communication interface 1050 may include, but is not limited to, a transceiver configured to send and receive data through the communication channel 1060. The communication interface 1050 may include, but is not limited to, a modem or a network card, and the communication channel 1060 may be implemented, for example, within a wired and / or wireless medium.
[0134] In various embodiments, data is provided to the system 1000 in a streaming or otherwise manner using a wireless network such as a Wi-Fi network (e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers)). The Wi-Fi signals of these embodiments are received through a communication channel 1060 and a communication interface 1050 suitable for Wi-Fi communication. The communication channel 1060 of these embodiments is typically connected to an access point or router that provides access to an external network (including the Internet) to allow streaming applications and other over-the-top communications. Other embodiments provide streaming data to the system 1000 using a set-top box that transmits data through an HDMI connection of the input box 1130. Still other embodiments provide streaming data to the system 1000 using an RF connection of the input box 1130. As described above, various embodiments provide data in a non-streaming manner. In addition, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.
[0135] The system 1000 can provide output signals to a variety of output devices, including a display 1100, a speaker 1110, and other peripherals 1120. The display 1100 of various embodiments includes, for example, one or more of a touch screen display, an organic light emitting diode (OLED) display, a curved display, and / or a foldable display. The display 1100 can be used for a television, a tablet computer, a laptop computer, a mobile phone, or other devices. The display 1100 can also be integrated with other components (for example, as in a smartphone) or be separate (for example, an external display for a laptop computer). In examples of various embodiments, other peripherals 1120 include one or more of a stand-alone digital video disc (or digital versatile disc) (DVR, for both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripherals 1120 that provide functions based on the output of the system 1000. For example, a disk player performs the function of playing the output of the system 1000.
[0136] In various embodiments, control signals are communicated between the system 1000 and the display 1100, speaker 1110, or other peripheral device 1120 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that allow device-to-device control with or without user intervention. Output devices may be communicatively coupled to the system 1000 via dedicated connections through respective interfaces 1070, 1080, and 1090. Alternatively, output devices may be connected to the system 1000 through the communication interface 1050 using the communication channel 1060. The display 1100 and speaker 1110 may be integrated with other components of the system 1000 in a single unit in an electronic device (e.g., a television). In various embodiments, the display interface 1070 includes a display driver, such as a timing controller (T Con) chip.
[0137] Display 1100 and speaker 1110 may alternatively be separate from one or more other components, for example, if the RF portion of input 1130 is part of a separate set-top box. In various embodiments where display 1100 and speaker 1110 are external components, output signals may be provided through dedicated output connections (including, for example, an HDMI port, a USB port, or a COMP output).
[0138] The embodiments may be implemented by computer software or hardware or a combination of hardware and software implemented by the processor 1010. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. The memory 1020 may be of any type suitable for the technical environment and may be implemented using any appropriate data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory as non-limiting examples. The processor 1010 may be of any type suitable for the technical environment and may include one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture as non-limiting examples.
[0139] Various implementations involve decoding. "Decoding" as used in this application may encompass, for example, all or part of a process performed on a received coded sequence to produce a final output suitable for display. In various embodiments, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such a process also includes or alternatively includes a process performed by a decoder of the various embodiments described in this application.
[0140] As a further example, in one embodiment, "decoding" refers only to entropy decoding, in another embodiment, "decoding" refers only to differential decoding, and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or to generally refer to a broader decoding process will become clear based on the context of the specific description, and is believed to be well understood by those skilled in the art.
[0141] Various implementations involve encoding. Similar to the above discussion about "decoding", "encoding" as used in this application can cover, for example, all or part of the processes performed on an input video sequence to produce an encoded bitstream. In various embodiments, such processes include one or more processes typically performed by an encoder, such as partitioning, differential encoding, transforms, quantization, and entropy encoding. In various embodiments, such processes also include or alternatively include processes performed by an encoder of the various embodiments described in this application.
[0142] As a further example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process" is intended to refer specifically to a subset of operations or generally to a broader encoding process will become clear based on the context of the specific description and is believed to be well understood by those of ordinary skill in the art.
[0143] It should be noted that the grammatical elements used in this article are descriptive terms. Therefore, they do not exclude the use of other grammatical element names.
[0144] When a diagram is presented as a flow chart, it should be understood that it also provides a block diagram of the corresponding device. Similarly, when a diagram is presented as a block diagram, it should be understood that it also provides a flow chart of the corresponding method / process.
[0145] A number of embodiments may relate to parameter models or rate-distortion optimization. Specifically, during the encoding process, a balance or trade-off between rate and distortion is usually considered, usually taking into account constraints on computational complexity. It can be measured by a rate-distortion optimization (RDO) metric, or by a least mean square (LMS), mean absolute error (MAE) or other such measurements. Rate-distortion optimization is usually expressed as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. There are different methods for solving the rate-distortion optimization problem. For example, these methods can be based on extensive testing of all coding options, including all considered modes or coding parameter values, and a complete evaluation of their coding costs and the associated distortion of the reconstructed signal after encoding and decoding. Faster methods can also be used to save coding complexity, in particular, to calculate approximate distortion based on a prediction or prediction residual signal rather than a reconstructed signal. The two methods can also be mixed, for example, using approximate distortion only for some possible coding options and using complete distortion for other coding options. Other methods only evaluate a subset of possible coding options. More generally, many approaches employ any of a variety of techniques to perform optimization, but the optimization is not necessarily a complete assessment of the coding cost and associated distortion.
[0146] The embodiments and aspects described herein can be implemented as, for example, methods or processes, devices, software programs, data streams or signals. Even if only discussed in the context of a single embodiment (e.g., discussed only as a method), the implementation of the discussed features can also be implemented in other forms (e.g., devices or programs). The device can be implemented as, for example, suitable hardware, software, and firmware. The method can be implemented as, for example, a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device, such as a computer, a mobile phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate information communication between end users.
[0147] References to "an embodiment" or "an embodiment" or "an implementation" or "an implementation" and other variations thereof mean that the particular features, structures, characteristics, etc. described in connection with the embodiment are included in at least one embodiment. Therefore, the phrases "in an embodiment" or "in an embodiment" or "in an implementation" or "in an implementation" and any other variations appearing in multiple places in the present application do not necessarily all refer to the same embodiment.
[0148] Furthermore, the present application may involve “determining” a variety of information. Determining information may include one or more of the following, such as estimating information, calculating information, predicting information, or retrieving information from memory.
[0149] Furthermore, the present application may involve "accessing" various information. Accessing information may include one or more of the following, such as receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0150] Furthermore, the present application may involve "receiving" a variety of information. Like "accessing," receiving is intended to be a broad term. Receiving information may include one or more of the following, such as accessing information or retrieving information (e.g., from a memory). Furthermore, "receiving" generally involves in some way during an operation such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0151] It should be understood that the use of any of the following " / ", "and / or", and "at least one", such as in the case of "A / B", "A and / or B", and "at least one of A and B", is intended to cover selecting only the first listed option (A), or only the second listed option (B), or selecting both options (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A and B and C). This can be extended to as many items as listed as will be apparent to one of ordinary skill in this and related arts.
[0152] In addition, as used herein, the term "signal" refers, among other things, to indicating something to a corresponding decoder. For example, in some embodiments, an encoder signals a specific one of a plurality of transforms, coding modes or flags. Thus, in one embodiment, the same transform, parameter or mode is used on the encoder side and the decoder side. Thus, for example, an encoder can transmit (explicitly signal) a specific parameter to a decoder so that the decoder can use the same specific parameter. On the contrary, if the decoder already has the specific parameter and other parameters, a signal can be made without transmission (implicit signaling) to simply allow the decoder to know and select the specific parameter. Bit savings are achieved in various embodiments by avoiding transmission of any actual function. It should be understood that signaling can be done in a variety of ways. For example, in various embodiments, one or more grammatical elements, flags, etc. are used to signal information to a corresponding decoder. Although the verb form of the word "signal" is mentioned above, the word "signal" can also be used as a noun here.
[0153] As will be appreciated by those of ordinary skill in the art, embodiments may generate a variety of formatted signals to carry information that may be, for example, stored or transmitted. The information may include, for example, instructions for executing a method or data generated by one of the embodiments. For example, a signal may be formatted to carry a bitstream of the embodiment. Such a signal may be formatted as, for example, an electromagnetic wave (e.g., using a radio frequency portion of a spectrum) or a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is well known, a signal may be transmitted over a variety of different wired or wireless links. The signal may be stored on a processor readable medium.
[0154] The preceding sections describe a number of embodiments across a variety of claim categories and types. Features of these embodiments may be provided individually or in any combination. In addition, embodiments may include one or more of the following features, devices, or aspects, individually or in any combination across a variety of claim categories and types. At least one embodiment includes the following features:
[0155] Use neural networks to encode and decode video information;
[0156] determining a number of bits to allocate to an encoded portion of the video based on the plurality of frames, using a gain vector from a reinforcement learning agent using a latent representation determined from the encoding;
[0157] Use asymmetric gain variational autoencoder for encoding and decoding;
[0158] Reinforcement agents implemented using deep neural networks;
[0159] The encoding and decoding further includes: data indicating the potential representation is included in the encoded bit stream;
[0160] a bitstream or signal comprising one or more of said syntax elements or variants thereof;
[0161] A bitstream or signal comprising syntax conveying information generated according to any of the described embodiments;
[0162] Creating and / or transmitting and / or receiving and / or decoding according to any of the embodiments described;
[0163] Parsing video data or bitstream to determine the operating point of the codec;
[0164] A method, process, device, medium storing instructions, medium storing data or signal provided according to any one of the embodiments;
[0165] Inserting syntax elements in the signaling that enable the decoder to determine decoding information in a manner corresponding to that used by the encoder;
[0166] Creating and / or transmitting and / or receiving and / or decoding a bitstream or signal comprising one or more of said syntax elements or variants thereof;
[0167] A television, set-top box, mobile phone, tablet computer or other electronic device that executes the transformation method according to any of the embodiments;
[0168] A television, set-top box, mobile phone, tablet computer or other electronic device that performs the transformation method according to any of the embodiments described above to determine and displays (e.g., using a monitor, screen or other type of display) a resulting image;
[0169] A television, set-top box, mobile phone, tablet or other electronic device that selects, bandwidth limits or tunes (e.g., using a tuner) a channel to receive a signal including an encoded image and performs a conversion method according to any of the embodiments described;
[0170] A television, set-top box, mobile phone, tablet or other electronic device that wirelessly receives (eg, using an antenna) a signal containing an encoded image and performs the transformation method.
Claims
1. A method comprising: encoding a portion of the video using the determined number of bits; as well as A number of bits allocated to an encoded portion of a video is determined based on a plurality of frames, wherein the determining includes using a gain vector from a reinforcement learning agent that uses a latent representation determined from the encoding.
2. A device configured to perform: encoding a portion of the video using the determined number of bits; and A number of bits allocated to an encoded portion of a video is determined based on a plurality of frames, wherein the determining includes using a gain vector from a reinforcement learning agent that uses a latent representation determined from the encoding.
3. A method comprising: Parse the video data to obtain the lambda value; determining an index of a vector having a corresponding lambda that is closest to a target lambda value, wherein the lambda defines a rate-distortion operating point; interpolating vector values between the determined index vector and the vectors with consecutive indices using the target lambda value and lambda values of vectors between the determined index vector and the vectors with consecutive indices; and The video data is decoded using the interpolated vectors.
4. A device configured to perform: Parse the video data to obtain the lambda value; determining an index of a vector having a corresponding lambda that is closest to a target lambda value, wherein the lambda defines a rate-distortion operating point; interpolating vector values between the determined index vector and the vectors with consecutive indices using the target lambda value and lambda values of vectors between the determined index vector and the vectors with consecutive indices; and The video data is decoded using the interpolated vectors.
5. The method according to any one of claims 1 and 3, or the apparatus according to any one of claims 2 and 4, wherein an asymmetric gain variational autoencoder (AG-VAE) architecture is used.
6. The method according to claim 1 or the apparatus according to claim 2, wherein the reinforcement learning is implemented by a deep neural network (DNN).
7. The method of claim 1 or the apparatus of claim 2, wherein the plurality of frames is a group of pictures (GOP).
8. The method of claim 1 or the apparatus of claim 2, wherein the reinforcement learning agent determines a range of bit rate points.
9. The method according to any one of claims 1 and 3, or the apparatus according to any one of claims 2 and 4, wherein the video processed is pure intra mode prediction.
10. The method of claim 1 or the apparatus of claim 2, wherein the reinforcement learning agent receives a frame type as input.
11. A device comprising: The device according to claim 4; as well as At least one of: (i) an antenna configured to receive a signal, the signal comprising a video block; (ii) a frequency band limiter configured to limit the received signal to a frequency band comprising the video block; and (iii) a display configured to display an output representing the video block.
12. A non-transitory computer-readable medium comprising data content generated by the method according to claim 1 or by the apparatus according to claim 2 for playback using a processor.
13. A signal comprising video data generated according to the method of claim 1 or by the apparatus of claim 2, for playback using a processor.
14. A computer program product comprising instructions, wherein when the program is executed by a computer, the instructions cause the computer to perform the method according to claim 1 or claim 3.
15. A non-transitory computer-readable medium containing data content including instructions for executing the method according to any one of claims 1, 3, and 5 to 10.