Bitrate estimation for video coding with machine learning enhancement

By employing codec proxy and backpropagation techniques, the problem of neural networks' inability to effectively integrate bit rate and distortion performance derivatives in video decoding is solved, thereby improving the efficiency and performance of video decoding.

CN119487524BActive Publication Date: 2025-11-21QUALCOMM INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380051514.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-09-09
Filing Date
2023-06-12
Publication Date
2025-11-21
Estimated Expiration
2043-06-12

AI Technical Summary

Technical Problem

Existing video decoding technologies, when using neural networks for performance optimization, cannot effectively integrate the derivatives of bit rate and distortion performance, resulting in suboptimal performance.

Method used

By employing a codec proxy and combining it with backpropagation technology, we can achieve accurate estimation of codec performance and optimization of partial derivatives, and improve video decoding efficiency by leveraging machine learning.

Benefits of technology

It improves the efficiency and performance of video decoding, and enables effective estimation and optimization of codec performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119487524B_ABST
    Figure CN119487524B_ABST
Patent Text Reader

Abstract

Techniques for processing video data are described herein. For example, a technique can include encoding one or more frames of video data using a video encoder, the video encoder including at least a quantization process; determining an actual bitrate for the encoded one or more frames; predicting an estimated bitrate using an encoder agent, the encoder agent including a statistical model for estimating a bitrate for the encoded one or more frames; determining a gradient of the estimated bitrate using the encoder agent; and training the encoder agent to predict the estimated bitrate based on the actual bitrate, the estimated bitrate, and the gradient.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure generally relates to video coding (e.g., encoding and / or decoding video data). For example, aspects of the present disclosure relate to systems and techniques for performing bitrate estimation to enhance video coding using machine learning. BACKGROUND

[0002] Many devices and systems allow video data to be processed and output for consumption. Digital video data includes a large amount of data to satisfy the needs of consumers and video providers. For example, consumers of video data desire high quality video including high fidelity, high resolution, high frame rate, and the like. As a result, the large amount of video data required to meet these demands places a burden on communication networks and devices that process and store the video data.

[0003] Video coding techniques can be used to compress video data. One goal of video coding is to compress video data into a form that uses a lower bit rate, while avoiding or minimizing a degradation in video quality. As ever-evolving video services become available, there is a need for coding techniques with better coding efficiency. SUMMARY

[0004] The following presents a simplified summary relating to one or more aspects disclosed herein. Thus, the following summary should not be considered an extensive overview relating to all contemplated aspects, nor should the following summary be deemed to identify key or critical elements relating to all contemplated aspects or to delineate the scope associated with any particular aspect. The sole purpose of the following summary is to present some concepts relating to one or more aspects disclosed herein in a simplified form to precede the detailed description herein below.

[0005] Systems and techniques are described for coding (e.g., encoding and / or decoding) image and / or video content. In one illustrative example, a method of processing video data is provided. The method includes encoding one or more frames of video data using a video encoder, the video encoder including at least a quantization process; determining an actual bitrate of the encoded one or more frames; predicting an estimated bitrate using an encoder agent, the encoder agent including a statistical model for estimating a bitrate of the encoded one or more frames; determining a gradient of the estimated bitrate using the encoder agent; and training the encoder agent to predict the estimated bitrate based on the actual bitrate, the estimated bitrate, and the gradient.

[0006] In another example, an apparatus for processing video data is provided. The apparatus comprises at least one memory; and at least one processor coupled to the at least one memory, the at least one processor configured to encode one or more frames of video data using a video encoder, the video encoder comprising at least a quantization process; determine an actual bitrate of the encoded one or more frames; predict an estimated bitrate using an encoder agent, the encoder agent comprising a statistical model for estimating a bitrate of the encoded one or more frames; determine a gradient of the estimated bitrate using the encoder agent; and train the encoder agent to predict the estimated bitrate based on the actual bitrate, the estimated bitrate, and the gradient.

[0007] In another example, a non-transitory computer-readable storage medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to: receive one or more frames of video data for encoding by a video encoder, the video encoder comprising at least a quantization process; predict an estimated bitrate of the one or more frames after encoding by the video encoder using an encoder agent, wherein the encoder agent comprises a statistical model for estimating the estimated bitrate, and wherein the statistical model is trained based on a gradient of the estimated bitrate; adjust one or more quality parameters based on the predicted estimated bitrate; and encode the one or more frames of video data using the video encoder, the video encoder comprising at least the quantization process.

[0008] In another example, an apparatus is provided. The apparatus comprises means for receiving one or more frames of video data for encoding by a video encoder, the video encoder comprising at least a quantization process; means for predicting an estimated bitrate of the one or more frames after encoding by the video encoder using an encoder agent, wherein the encoder agent comprises a statistical model for estimating the estimated bitrate, and wherein the statistical model is trained based on a gradient of the estimated bitrate; means for adjusting one or more quality parameters based on the predicted estimated bitrate; and means for encoding the one or more frames of video data using the video encoder, the video encoder comprising at least the quantization process.

[0009] In some aspects, the apparatus includes a mobile device (e.g., a mobile phone or so-called “smart phone,” a tablet computer, or other type of mobile device), a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a video server, a television (e.g., a networked television), a vehicle (or a computing device of a vehicle), or other device. In some aspects, the apparatus includes at least one camera for capturing one or more images or video frames. For example, the apparatus can include a camera (e.g., an RGB camera) or multiple cameras for capturing one or more images and / or one or more videos including video frames. In some aspects, the apparatus includes a display for displaying one or more images, videos, notifications, or other displayable data. In some aspects, the apparatus includes a transmitter configured to transmit one or more video frames and / or syntax data to at least one device over a transmission medium. In some aspects, the processor includes a neural processing unit (NPU), a central processing unit (CPU), a graphics processing unit (GPU), or other processing device or component.

[0010] This Summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used to determine the scope of the claimed subject matter. The subject matter should be understood from readin the entire specification of the patent, including any claims, the

[0011] The foregoing, along with other features and embodiments, will become more apparent from the following description, including the claims, when read in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0012] The illustrative embodiments of the present application are described below with reference to the following drawings:

[0013] Figure 1 An example implementation of a system on chip (SOC) is shown in accordance with some examples.

[0014] Figure 2 is a block diagram illustrating an encoding device and a decoding device in accordance with some examples.

[0015] Figure 3 is a diagram illustrating an example of a system including a device operable to perform image and / or video coding (encoding and decoding) using a machine learning coding system in accordance with some examples;

[0016] Figures 4A-4D is a diagram illustrating an example of a neural network in accordance with some examples;

[0017] Figure 5 is a diagram illustrating an example of a deep convolutional network in accordance with some examples;

[0018] Figure 6A and Figure 6B is a block diagram illustrating an example implementation of a video codec neural boost system according to some examples;

[0019] Figure 7 is a diagram illustrating encoder elements of an example video encoder according to some examples;

[0020] Figure 8 is a block diagram illustrating elements of an example differential encoder agent according to some examples.

[0021] Figure 9 is a graph plotting an example adjustment function and derivative of the adjustment function according to some examples;

[0022] Figures 10-16 is a histogram of example rates between frame bitrates according to some examples;

[0023] Figure 17 is a flowchart illustrating techniques for performing bitrate estimation according to aspects of the present disclosure;

[0024] Figure 18 is a flowchart illustrating techniques for performing bitrate estimation according to aspects of the present disclosure;

[0025] Figure 19 is an example computing device that can implement various techniques described herein. DETAILED DESCRIPTION

[0026] Certain aspects and embodiments of the present disclosure are provided below. Some of these aspects and embodiments can be applied independently, and some of them can be applied in combination, as will be apparent to those skilled in the art. In the following description, for the purposes of explanation, specific details are set forth in order to provide a thorough understanding of embodiments of the application. It will be apparent, however, that various embodiments can be practiced without

[0027] The following description is provided for the purpose of illustrating example embodiments, and is not intended to limit the scope, applicability, or configuration of the disclosure. Rather, the following description of example embodiments is provided as an enabling teaching of implementations. It should be understood that various changes can be made in the function and arrangement of elements without departing from the spirit and scope of the application as set forth in the appended claims.

[0028] Digital video data can include large amounts of data, particularly as the demand for high quality video data continues to grow. For example, consumers of video data often expect increasing quality of video with high fidelity, high resolution, high frame rates, etc. However, the large amounts of video data required to meet such demands can place a heavy burden on communicating networks and devices processing and storing the video data.

[0029] Video data can be coded using various techniques. Video coding can be performed according to a particular video coding standard or can use one or more machine learning systems or algorithms. Example video coding standards include Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), Moving Picture Experts Group (MPEG) coding (e.g., MPEG-5 Part 1 Basic Video Coding (EVC) or other MPEG-based coding), AOMedia Video 1 (AV1), etc. Video coding often uses prediction methods (e.g., inter-prediction or intra-prediction) that take advantage of redundancy in video images or sequences. One common goal is to compress video data into a form that uses fewer bits to represent the video data, while avoiding or minimizing distortion of the video data. As the demand for video services grows and new video services become available, there is a need for coding techniques that have better coding efficiency, performance, and rate control.

[0030] Video coding devices implement video compression techniques to efficiently encode and decode video data. Video compression techniques can include applying different prediction modes to reduce or eliminate redundancy inherent in video sequences, including spatial prediction (e.g., intra-prediction or intra- prediction), temporal prediction (e.g., inter-prediction or inter-prediction), inter-layer prediction (across different layers of video data), and / or other prediction techniques. A video encoder can partition each picture of an original video sequence into rectangular regions referred to as video blocks or coding units (described in greater detail below). The video blocks can be encoded using a particular prediction mode.

[0031] Video blocks can be partitioned into one or more groups of smaller blocks in one or more ways. Blocks can include coding tree blocks, prediction blocks, transform blocks, and / or other suitable blocks. Unless otherwise indicated, a reference to a “block” can generally refer to such video blocks (e.g., coding tree blocks, coding blocks, prediction blocks, transform blocks, or other suitable blocks or sub-blocks, as will be appreciated by one of ordinary skill in the art). Furthermore, each of these blocks can also be interchangeably referred to herein as a “unit” (e.g., a coding tree unit (CTU), a coding unit, a prediction unit (PU), a transform unit (TU), etc.). In some cases, a unit can indicate a coding logic unit encoded in a bitstream, while a block can indicate a portion of a video frame buffer to which a process is directed.

[0032] For inter prediction modes, the video encoder can search for a block that is similar to the block being encoded in a frame (or picture) located at another temporal location, referred to as a reference frame or reference picture. The video encoder can limit the search to a certain spatial displacement from the block to be encoded. The best match can be located using a two-dimensional (2D) motion vector that includes a horizontal displacement component and a vertical displacement component. For intra prediction modes, the video encoder can form the prediction block using a spatial prediction technique based on data from neighboring blocks previously encoded within the same picture.

[0033] The video encoder can determine a prediction error. For example, the prediction can be determined as a difference between pixel values in the block being encoded and the prediction block. The prediction error can also be referred to as a residual. The video encoder can also apply a transform to the prediction error using transform coding (e.g., in the form of a discrete cosine transform (DCT), a discrete sine transform (DST), or other suitable transform) to generate transform coefficients. After the transform, the video encoder can quantize the transform coefficients. The quantized transform coefficients and the motion vectors can be represented using syntax elements and, along with control information, form a coded representation of the video sequence. In some cases, the video encoder can entropy code the syntax elements, further reducing the number of bits needed for their representation.

[0034] The video decoder can use the syntax elements and control information discussed above to construct predictive data (e.g., a predictive block) for decoding the current frame. For example, the video decoder can add the prediction block and the compressed prediction error. The video decoder can determine the compressed prediction error by weighting transform basis functions using the quantized coefficients. The difference between the reconstructed frame and the original frame is referred to as the reconstruction error.

[0035] As described in more detail below, systems, apparatuses, methods (also referred to as processes), and computer-readable media (collectively referred to as “systems and techniques”) are described herein for performing bitrate estimation to enhance video coding using machine learning. In some devices, video codecs can be implemented using custom hardware, such as an application-specific integrated circuit (ASIC). Custom hardware facilitates high efficiency in terms of processing speed and power usage, but greatly reduces flexibility, as additional features require the design and deployment of new hardware, which can be both slow and expensive.

[0036] Techniques have been proposed to improve performance by adaptively modifying the video before encoding and after decoding, while taking advantage of the ubiquity of codec hardware. Recent proposals have shown the advantage of using machine learning and neural networks for this purpose, as they can use advanced algorithms and extensive training to better recognize how video compression can be improved. This approach can be referred to as standard video codec neural boosting.

[0037] A fundamental problem with many such techniques is that while neural-based solutions can use derivatives of bitrate and distortion performance measures, their design and optimization can be more effective. However, these derivatives can not be available from complex standard video codecs, and thus they cannot be effectively integrated into end-to-end system design, resulting in suboptimal performance.

[0038] The present disclosure addresses at least this problem using a new type of codec agent. For example, a codec agent in accordance with the systems and techniques described herein can efficiently and accurately estimate codec performance as well as the required partial derivatives (e.g., all partial derivatives) for optimization. The codec agent can enable integration with backpropagation techniques (e.g., gradient backpropagation) used by neural network training tools. Experimental results discussed below demonstrate the accuracy of the estimates using comparisons to bitrate of the HEVC / H.265 video standard.

[0039] Various aspects of the present disclosure will be described with reference to the drawings.

[0040] Figure 1 An example implementation of a system on a chip (SOC) 100 is shown, which can include a central processing unit (CPU) 102 or multi-core CPU configured to perform one or more functions described herein. Parameters or variables (e.g., neural signals and synaptic weights), system parameters associated with a computing device (e.g., a neural network with weights), delays, frequency bin information, task information, and other information can be stored in a memory block associated with a neural processing unit (NPU) 108, in a memory block associated with the CPU 102, in a memory block associated with a graphics processing unit (GPU) 104, in a memory block associated with a digital signal processor (DSP) 106, in a memory block 118, and / or can be distributed among multiple blocks. Instructions executed at the CPU 102 can be loaded from a program memory associated with the CPU 102 or can be loaded from the memory block 118.

[0041] The SOC 100 can also include additional processing blocks tailored for specific functions, such as the GPU 104; the DSP 106; a connectivity block 110, which can include fifth generation (5G) connectivity, fourth generation long term evolution (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, etc.; and a multimedia processor 112, which can, for example, detect and recognize gestures. In one implementation, the NPU is implemented in the CPU 102, the DSP 106, and / or the GPU 104. The SOC 100 can also include a sensor processor 114, an image signal processor (ISP) 116, and / or a navigation module 120, which can include a global positioning system.

[0042] The SOC 100 can be based on an ARM instruction set. In an aspect of the disclosure, instructions loaded into the CPU 102 can include code to search a lookup table (LUT) for a stored multiplication result corresponding to a product of an input value and a filter weight. The instructions loaded into the CPU 102 can also include code to disable a multiplier during a multiplication operation of the product when a lookup table hit for the product is detected. Additionally, the instructions loaded into the CPU 102 can include code to store a computed product of the input value and the filter weight when a lookup table miss for the product is detected.

[0043] The SOC 100 and / or components thereof can be configured to perform video compression and / or decompression (also referred to as video encoding and / or decoding, collectively referred to as video coding) using standards-based video coding and / or using machine learning techniques. With respect to standards-based video coding, the SOC 100 and / or components thereof can be configured to perform video coding in accordance with one or more video coding standards, such as the ITU-T H.261, MPEG-l Visual, ITU-T H.262 or MPEG-2 Visual, ITU-T H.263 or MPEG-4 Visual, ITU-T H.264 or MPEG-4 AVC, including its Scalable Video Coding (SVC) and Multiview Video Coding (MVC) extensions, ITU-T H.265 or MPEG-4 AVC, including its SVC and MVC extensions, ITU-T H.266 or Versatile Video Coding (VVC), and / or other standards. Figure 2 and Figure 3 Examples of standards-based and machine learning-based video coding systems are described.

[0044] Figure 2 is a block diagram illustrating an example of a system 200 including an encoding device 204 and a decoding device 212 that can respectively encode and decode video data in accordance with the examples described herein. In some examples, the encoding device 204 and / or the decoding device 212 can include a SOC 100 of FIG. 1. Figure 1 The encoding device 204 can be a part of a source device, and the decoding device 212 can be a part of a receiving device (sometimes referred to as a client device). In some examples, the source device can also include a decoding device similar to the decoding device 212. In some examples, the receiving device can also include an encoding device similar to the encoding device 204. The source device and / or the receiving device can include an electronic device such as a mobile or fixed telephone handset (e.g., a smart phone, a cellular telephone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video gaming console, an Internet Protocol (IP) camera, a server device in a server system including one or more server devices (e.g., a video streaming server system or other suitable server system), a head-mounted display (HMD), a heads-up display (HUD), smart glasses (e.g., virtual reality (VR) glasses, augmented reality (AR) glasses, or other smart glasses), or any other suitable electronic device.

[0045] The components of System 200 may include electronic circuits or other electronic hardware and / or may be implemented using electronic circuits or other electronic hardware, which may include SOC 100 and / or one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), neural processing units (NPUs), and / or other suitable electronic circuits), and / or may include computer software, firmware, or any combination thereof and / or may be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein.

[0046] Although system 200 is shown to include certain components, those skilled in the art will understand that system 200 may include more than Figure 2 The components shown may include more or fewer components. For example, in some cases, system 200 may also include one or more memory devices other than storage 208 and storage 218 (e.g., one or more random access memory (RAM) components, read-only memory (ROM) components, cache memory components, buffer components, database components, and / or other memory devices), one or more processing devices (e.g., one or more CPUs, GPUs, NPUs, and / or other processing devices) communicating with and / or electrically connected to one or more memory devices, one or more wireless interfaces for performing wireless communication (e.g., one or more transceivers and baseband processors for each wireless interface), one or more wired interfaces for performing communication via one or more hardwired connections (e.g., serial interfaces such as Universal Serial Bus (USB) inputs, lighting connectors, and / or other wired interfaces), and / or Figure 2 Other components not shown.

[0047] The decoding techniques described herein are applicable to video decoding in a variety of multimedia applications, including streaming video transmission (e.g., via the Internet), television broadcasting or transmission, encoding of digital video for storage on data storage media, decoding of digital video stored on data storage media, or other applications. In some examples, system 200 may support one-way or two-way video transmission to support applications such as video conferencing, video streaming, video playback, video broadcasting, gaming, and / or video telephony.

[0048] In some examples, the encoding device 204 (or encoder) can be configured to encode video data using a video coding standard or protocol to generate an encoded video bitstream. Examples of video coding standards include ITU-T H.261, ISO / IEC MPEG-l Visual, ITU-T H.262 or ISO / IEC MPEG-2 Visual, ITU-T H.263, ISO / IEC MPEG-4 Visual, ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC), including its Scalable Video Coding (SVC) and Multiview Video Coding (MVC) extensions, High Efficiency Video Coding (HEVC) or ITU-T H.265, Versatile Video Coding (VVC) or ITU-T H.266, and / or other video coding standards. One or more video coding standards have extensions associated therewith that relate to other aspects of video coding. For example, various extensions of HEVC address multi-layer video coding, including the range and screen content coding extensions, 3D video coding (3D-HEVC) and multi-view extension (MV-HEVC), and scalable extension (SHVC).

[0049] Many of the embodiments described herein can be performed using a video codec such as VVC, HEVC, AVC, and / or extensions thereof. However, the techniques and systems described herein can also be applicable to other coding standards, such as MPEG, JPEG (or other coding standards for still images), VP9, AV1, extensions thereof, or other suitable coding standards that are already available or not yet available or developed, such as machine learning based video coding described below. Thus, although the techniques and systems described herein can be described with reference to a particular video coding standard, one of ordinary skill in the art will understand that the description should not be interpreted as applying only to that particular standard.

[0050] Referring to Figure 2 The video source 202 can provide the video data to the encoding device 204. The video source 202 can be part of the source device or can be part of a device other than the source device. The video source 202 can include a video capture device (e.g., a video camera, a camcorder, a webcam, etc.), a video archive containing stored video, a video server or content provider that provides video data, a video feed interface that receives video from a video server or content provider, a computer graphics system for generating computer graphics video data, a combination of such sources, or any other suitable video source.

[0051] Video data from video source 202 can include one or more input pictures. Pictures can also be referred to as “frames.” A picture or frame is a still image that in some cases is part of a video. In some examples, data from video source 202 can be a still image that is not part of a video. In HEVC, VVC, and other video coding specifications, a video sequence can include a series of pictures. A picture can include three sample arrays (denoted as SL, SCb, and SCr). SL is a two-dimensional array of luma samples, SCb is a two-dimensional array of Cb chrominance samples, and SCr is a two-dimensional array of Cr chrominance samples. Chrominance samples can also be referred to herein as “chroma” samples. In other cases, a picture can be monochrome and can include only an array of luma samples.

[0052] Encoder engine 206 (or encoder) of encoding device 204 encodes the video data to generate an encoded video bitstream. In some examples, the encoded video bitstream (or “video bitstream” or “bitstream”) is a series of one or more coded video sequences. According to HEVC, a coded video sequence (CVS) includes a series of access units (AUs) starting with an AU having a random access point picture in the base layer and having certain properties (e.g., a RASL flag (e.g., NoRaslOutputFlag) equal to 1) up to, but not including, the next AU having a random access point picture in the base layer and having certain properties. An AU includes one or more coded pictures and control information corresponding to coded pictures that share the same output time. Coded slices of a picture are encapsulated at the bitstream level as data units called network abstraction layer (NAL) units. For example, an HEVC video bitstream can include one or more CVSs including NAL units. Each of the NAL units has a NAL unit header. Syntax elements in the NAL unit header take up specified bits and are thus visible to all types of systems and transport layers, such as transport streams, real-time transport (RTP) protocols, file formats, etc.

[0053] There are two types of NAL units in the HEVC standard, including video coding layer (VCL) NAL units and non-VCL NAL units. VCL NAL units include coded picture data that forms a coded video bitstream. For example, a bit sequence that forms a coded video bitstream is present in VCL NAL units. A VCL NAL unit can include one slice or slice segment (described below) of coded picture data, and non-VCL NAL units include control information related to one or more coded pictures. In some cases, a NAL unit can be referred to as a packet. An HEVC AU includes VCL NAL units (which contain coded picture data) and non-VCL NAL units (if any) corresponding to the coded picture data. Among other information, non-VCL NAL units can contain parameter sets, which have high-level information related to the encoded video bitstream. For example, parameter sets can include a video parameter set (VPS), a sequence parameter set (SPS), and a picture parameter set (PPS). In some cases, each slice or other portion of a bitstream can reference a single active PPS, SPS, and / or VPS to allow a decoding device 212 to access information that can be used to decode the slice or other portion of the bitstream.

[0054] A NAL unit can contain a bit sequence (e.g., an encoded video bitstream, a CVS of a bitstream, etc.) that forms a coded representation of video data (e.g., a coded representation of a picture in a video). The encoder engine 206 generates a coded representation of a picture by dividing the picture into multiple slices. A slice is independent of other slices, so information in a slice can be coded without having to rely on data from other slices within the same picture. A slice includes one or more slice segments (including an independent slice segment and one or more dependent slice segments (if any) that depend on a previous slice segment).

[0055] A slice is then divided in HEVC into coding tree blocks (CTBs) of luma samples and chroma samples. A CTB of luma samples and one or more CTBs of chroma samples, along with the syntax of the samples, is referred to as a coding tree unit (CTU). A CTU can also be referred to as a “treeblock” or a “largest coding unit” (LCU). A CTU is the basic processing unit of HEVC encoding. A CTU can be divided into multiple coding units (CUs) of varying sizes. A CU contains an array of luma and chroma sample arrays, which is referred to as a coding block (CB).

[0056] Luma and chroma CBs can be further partitioned into prediction blocks (PBs). A PB is a block of samples of a luma component or a chroma component that is inter-predicted using the same motion parameters or intra-block copy (IBC) prediction (when available or enabled for use). A luma PB and one or more chroma PBs, together with associated syntax, form a prediction unit (PU). For inter-prediction, a set of motion parameters (e.g., one or more motion vectors, reference indices, etc.) are signaled in the bitstream for each PU and used for inter-prediction of the luma PB and the one or more chroma PBs. The motion parameters can also be referred to as motion information. The CBs can also be partitioned into one or more transform blocks (TBs). A TB represents a square block of samples of a color component on which a residual

[0057] The size of a CU corresponds to the size of the coding mode and can be square in shape. For example, the size of a CU can be 8x8 samples, 16x16 samples, 32x32 samples, 64x64 samples, or any other appropriate size up to the size of a corresponding CTU. The phrase “NxN” is used herein to refer to pixel dimensions of a video block in terms of vertical and horizontal dimensions (e.g., 8 pixels by 8 pixels). The pixels in a block can be arranged in rows and columns. In some embodiments, a block can not have the same number of pixels in the horizontal direction as in the vertical direction as a CU. Syntax data associated with a CU can describe, for example, partitioning of the CU into one or more PUs. The partitioning mode can differ between whether the CU is intra- or inter- mode encoded. PUs can be partitioned into non-square shapes. Syntax data associated with a CU can also describe, for example, partitioning of the CU into one or more TUs according to a CTU. The shape of a TU can be square or non-square.

[0058] According to HEVC, a transform unit (TU) can be used to perform a transform. The TU can be different for different CUs. The size of a TU can be determined based on the size of a PU within a given CU. The size of a TU can be the same as or smaller than a PU. In some examples, a quadtree structure referred to as a residual quadtree (RQT) can be used to subdivide residual samples corresponding to a CU into smaller units. Leaf nodes of the RQT can correspond to TUs. Pixel difference values associated with a TU can be transformed to produce transform coefficients. The transform coefficients can be quantized by the encoder engine 206.

[0059] Once the pictures of the video data are partitioned into CUs, the encoder engine 206 predicts each PU using a prediction mode. The prediction units or blocks are subtracted from the original video data to obtain residuals (described below). For each CU, the prediction mode can be signaled inside the bitstream using syntax data. The prediction mode can include intra prediction (or intra-picture prediction) or inter prediction (or inter-picture prediction). Intra prediction exploits the correlation between spatially neighboring samples within a picture. For example, in the case of using intra prediction, each PU is predicted from neighboring image data in the same picture using, e.g., DC prediction to find an average value for the PU, planar prediction to fit a planar surface to the PU, directional prediction to extrapolate from neighboring data, or any other suitable type of prediction. Inter prediction uses the temporal correlation between pictures to derive a motion compensated prediction for a block of image samples. For example, in the case of using inter prediction, each PU is predicted from image data in one or more reference pictures (that precede or follow the current picture in output order) using motion compensated prediction. The decision whether to code a picture region using inter-picture or intra-picture prediction can be made, e.g., on a CU level.

[0060] As described above, in some cases, the encoder engine 206 and the decoder engine 216 (described in more detail below) can be configured to operate according to VVC. According to VVC, a video coder (e.g., the encoder engine 206 and / or the decoder engine 216) partitions a picture into a plurality of coding tree units (CTUs) (where a CTB of luma samples and one or more CTBs of chroma samples, along with syntax for the samples, are referred to as a CTU). The video coder can partition a CTU according to a tree structure, such as a quad-tree binary tree (QTBT) structure or Multi-Type Tree (MTT) structure. The QTBT structure removes the concepts of multiple partition types, such as the separation between CUs, PUs, and TUs in HEVC. The QTBT structure includes two levels, including a first level partitioned according to quad-tree partitioning and a second level partitioned according to binary tree partitioning. A root node of the QTBT structure corresponds to a CTU. Leaf nodes of the binary trees correspond to coding units (CUs).

[0061] In the MTT partitioning structure, quad-tree partitioning, binary tree partitioning, and one or more types of ternary tree partitioning can be used to partition a block. Ternary tree partitioning is a partitioning that splits one block into three sub-blocks. In some examples, a ternary tree partitioning splits one block into three sub-blocks without splitting the original block by a center. The partitioning types (e.g., quad-tree, binary tree, and ternary tree) in the MTT can be symmetric or asymmetric.

[0062] In some examples, a video coder can use a single QTBT or MTT structure to represent each of the luma and chroma components, while in other examples, the video coder can use two or more QTBT or MTT structures, e.g., one QTBT or MTT structure for the luma component and another QTBT or MTT structure for the two chroma components (or two QTBT and / or MTT structures for the respective chroma components).

[0063] A video coder can be configured to use quad-tree partitioning per HEVC, QTBT partition, MTT partition, or other partition structure. For illustrative purposes, the description herein can refer to QTBT partitioning. However, it should be understood that the techniques of this disclosure can also apply to video coders configured to use quad-tree partitioning or other types of partitioning.

[0064] As described above, intra-picture prediction exploits correlation between spatially neighboring samples within a picture. There are multiple intra-prediction modes (also referred to as “intra modes”). In some examples, intra-prediction of a luma block includes 35 modes, including a planar mode, a DC mode, and 33 angular modes (e.g., diagonal intra-prediction modes and angular modes adjacent to the diagonal intra-prediction modes). The 35 modes of intra-prediction are indexed as shown in Table 1 below. In other examples, more intra modes can be defined, including prediction angles that can not yet be represented by the 33 angular modes. In other examples, the prediction angles associated with the angular modes can be different from those used in HEVC.

[0065] Intra prediction mode Association name 0 INTRA_PLANAR 1 INTRA_DC 2..34 INTRA_ANGULAR2..INTRA_ANGULAR34

[0066] Table 1 - Specification of Intra Prediction Modes and Associated Names

[0067] Inter-picture prediction uses temporal correlation between pictures to derive a motion-compensated prediction for a block of image samples. Using a translational motion model, the position of a block in a previously decoded picture (reference picture) is indicated by a motion vector (Ax, Ay), where Ax specifies the horizontal displacement and Ay specifies the vertical displacement of the position of the reference block relative to the current block. In some cases, the motion vector (Ax, Ay) can have integer sample precision (also referred to as integer precision), in which case the motion vector points to an integer pixel grid (or integer pixel sampling grid) of the reference frame. In some cases, the motion vector (Ax, Ay) can have fractional sample precision (also referred to as fractional pixel precision or non-integer precision) to more accurately capture the motion of the underlying object without being restricted to the integer pixel grid of the reference frame. The accuracy of the motion vector can be represented by the quantization level of the motion vector. For example, the quantization level can be integer precision (e.g., 1 pixel) or fractional pixel precision (e.g., ¼ pixel, ½ pixel, or other sub-pixel values). When the corresponding motion vector has fractional sample precision, interpolation is applied to the reference picture to derive the prediction signal. For example, samples available at integer positions can be filtered (e.g., using one or more interpolation filters) to estimate the values at fractional positions. The previously decoded reference picture is indicated by a reference index (refIdx) of a reference picture list. The motion vector and the reference index can be referred to as motion parameters. Two types of inter-picture prediction can be performed, including uni-prediction and bi-prediction.

[0068] In the case of inter prediction using bi-prediction, two sets of motion parameters (Ax0, Ay0, refldx0 and Ax1, Ay1, refldx1) are used to generate two motion-compensated predictions (from the same reference picture or possibly from different reference pictures). For example, in the case of bi-prediction, each prediction block uses two motion-compensated prediction signals, and B prediction units are generated. The two motion-compensated predictions are combined to derive the final motion-compensated prediction. For example, the two motion-compensated predictions can be combined by taking an average. In another example, weighted prediction can be used, in which case different weights can be applied to each motion-compensated prediction. The reference pictures that can be used for bi-prediction are stored in two separate lists, denoted as List 0 and List 1. The motion parameters can be derived at the encoder using a motion estimation process.

[0069] In the case of inter prediction using uni-prediction, one set of motion parameters (Ax0, Ay0, refldx0) is used to generate a motion-compensated prediction from a reference picture. For example, in the case of uni-prediction, each prediction block uses at most one motion-compensated prediction signal, and P prediction units are generated.

[0070] The PU can include data related to the prediction process (e.g., motion parameters or other suitable data). For example, when the PU is encoded using intra prediction, the PU can include data describing an intra prediction mode used for the PU. As another example, when the PU is encoded using inter prediction, the PU can include data defining a motion vector used for the PU. The data defining the motion vector used for the PU can describe, for example, a horizontal component of the motion vector (Ax), a vertical component of the motion vector (Ay), a resolution of the motion vector (e.g., integer precision, quarter-pel precision, or eighth-pel precision), a reference picture to which the motion vector points, a reference index, a reference picture list of the motion vector (e.g., List 0, List 1, or List C), or any combination thereof.

[0071] After performing prediction using intra prediction and / or inter prediction, the encoding device 204 can perform transform and quantization. For example, after prediction, the encoder engine 206 can calculate residual values corresponding to the PU. The residual values can include pixel difference values between the current block (PU) being coded and a prediction block used to predict the current block (e.g., a predicted version of the current block). For example, after generating the prediction block (e.g., using inter prediction or intra prediction), the encoder engine 206 can generate a residual block by subtracting the prediction block produced by the prediction unit from the current block. The residual block includes a set of pixel difference values that quantize the differences between the pixel values of the current block and the pixel values of the prediction block. In some examples, the residual block can be represented in a two-dimensional block format (e.g., a two-dimensional matrix or array of pixel values). In such examples, the residual block is a two-dimensional representation of pixel values.

[0072] Any residual data that can remain after performing prediction is transformed using a block transform, which can be based on a discrete cosine transform (DCT), a discrete sine transform (DST), an integer transform, a wavelet transform, other suitable transform function, or any combination thereof. In some cases, one or more block transforms (e.g., kernels sized 32x32, 16x16, 8x8, 4x4, or other suitable size) can be applied to the residual data in each CU. In some examples, TUs can be used for the transform and quantization processes implemented by the encoder engine 206. A given CU having one or more PUs can also include one or more TUs. As described in further detail below, residual values can be transformed into transform coefficients using a block transform, and the residual values can be quantized and scanned using a TU to generate serialized transform coefficients for entropy coding.

[0073] In some embodiments, after intra-predictive coding or inter-predictive coding using the PUs of a CU, the encoder engine 206 can calculate residual data for the TUs of the CU. The PUs can include pixel data in the spatial domain (or pixel domain). As previously noted, the residual data can correspond to pixel difference values between pixels of the unencoded picture and the prediction values corresponding to the PUs. The encoder engine 206 can form one or more TUs including the residual data for the CU (which includes the PUs) and can transform the TUs to produce transform coefficients for the CU. The TUs can include coefficients in the transform domain after application of a block transform.

[0074] The encoder engine 206 can perform quantization of the transform coefficients. Quantization provides further compression by quantizing the transform coefficients to reduce the amount of data used to represent the coefficients. For example, quantization can reduce the bit depth associated with some or all of the coefficients. In one example, a coefficient having an n-bit value can be rounded down to an m-bit value during quantization, where n is greater than m.

[0075] Once quantization is performed, the coded video bitstream includes the quantized transform coefficients, prediction information (e.g., prediction modes, motion vectors, block vectors, etc.), partitioning information, and any other suitable data (such as other syntax data). The different elements of the coded video bitstream can be entropy encoded by the encoder engine 206. In some examples, the encoder engine 206 can scan the quantized transform coefficients using a predefined scan order to produce a serialized vector that can be entropy encoded. In some examples, the encoder engine 206 can perform an adaptive scan. After scanning the quantized transform coefficients to form a vector (e.g., a one-dimensional vector), the encoder engine 206 can entropy encode the vector. For example, the encoder engine 206 can use context-adaptive variable length coding, context-adaptive binary arithmetic coding, syntax-based context-adaptive binary arithmetic coding, probability interval partitioning entropy coding, or another suitable entropy encoding technique.

[0076] The output 210 of the encoding device 204 can transmit the NAL units that make up the encoded video bitstream data to a decoding device 212 of a receiving device over a communication link 220. The input 214 of the decoding device 212 can receive the NAL units. The communication link 220 can include a channel provided by a wireless network, a wired network, or a combination of wired and wireless networks. The wireless network can include any wireless interface or combination of interfaces, and can include any suitable wireless network (e.g., the Internet or other wide area networks, a packet-based network, WiFi™, radio frequency (RF), UWB, WiFi Direct, cellular, Long Term Evolution (LTE), WiMax™, etc.). The wired network can include any wired interface (e.g., fiber, Ethernet, powerline Ethernet, Ethernet over coaxial cable, digital signal line (DSL), etc.). The wired and / or wireless networks can be implemented using various devices such as base stations, routers, access points, bridges, gateways, switches, etc. The encoded video bitstream data can be modulated according to a communication standard (e.g., a wireless communication protocol) and transmitted to the receiving device.

[0077] In some examples, the encoding device 204 can store the encoded video bitstream data in a storage 208. The output 210 can retrieve the encoded video bitstream data from the encoder engine 206 or from the storage 208. The storage 208 can include any of a variety of distributed or locally accessed data storage media. For example, the storage 208 can include a hard drive, storage disk, flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data. The storage 208 can also include a decoded picture buffer (DPB) for storing reference pictures used in inter-prediction. In another example, the storage 208 can correspond to a file server, or another intermediate storage device that can store the encoded video generated by the source device. In such cases, a receiving device including the decoding device 212 can access the stored video data from the storage device via streaming or download. The file server can be any type of server capable of storing encoded video data and transmitting that encoded video data to the receiving device. Example file servers include web servers (e.g., for a website), FTP servers, network attached storage (NAS) devices, or local disk drives. The receiving device can access the encoded video data through any standard data connection, including an Internet connection. The access can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both that is suitable for accessing encoded video data stored on a file server. The transmission of the encoded video data from the storage 208 can be a streaming transmission, a download transmission, or a combination thereof.

[0078] Input 214 of the decoding device 212 receives encoded video bitstream data and can provide the video bitstream data to a decoder engine 216 or to storage 218 for later use by the decoder engine 216. For example, the storage 218 can include a DPB for storing reference pictures used in inter-prediction. The receiving device including the decoding device 212 can receive encoded video data to be decoded via the storage 208. The encoded video data can be modulated according to a communications standard, such as a wireless communication protocol, and transmitted to the receiving device. The communication medium used to transmit the encoded video data can comprise any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network, such as a local area network, a wide-area network, or a global network such as the Internet. The communication medium can include routers, switches, base stations, or any other equipment that can be useful to facilitate communications from a source device to a destination device.

[0079] The decoder engine 216 can decode the encoded video bitstream data by entropy decoding (e.g., using an entropy decoder) and extracting elements of one or more coded video sequences that make up the encoded video data. The decoder engine 216 can rescale and perform inverse transforms of the encoded video bitstream data. The residual data is passed to a prediction stage of the decoder engine 216. The decoder engine 216 predicts pixel blocks (e.g., PUs). In some examples, the prediction is added to the output of the inverse transform (residual data).

[0080] The video decoding device 212 can output the decoded video to a video destination device 222, which can include a display or other output device for displaying the decoded video data to a consumer of the content. In some aspects, the video destination device 222 can be part of the receiving device that includes the decoding device 212. In some aspects, the video destination device 222 can be part of a separate device from the receiving device.

[0081] In some embodiments, the video encoding device 204 and / or the video decoding device 212 can be integrated with an audio encoding device and an audio decoding device, respectively. The video encoding device 204 and / or the video decoding device 212 can also include other hardware or software that is necessary for implementing the coding techniques described above, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware or any combinations thereof. The video encoding device 204 and the video decoding device 212 can be integrated as part of a combined encoder / decoder (CODEC) in the respective device.

[0082] Figure 2The example system shown is one illustrative example that can be used herein. Techniques for processing video data using the techniques described herein can be performed by any digital video encoding and / or decoding device. While the techniques of this disclosure are generally performed by a video encoding device or a video decoding device, these techniques can also be performed by a combined video encoder-decoder, commonly referred to as a “CODEC.” Also, the techniques of this disclosure can also be performed by a video preprocessor. Source device and receive device are merely examples of such coding devices in which the source device generates coded video data for transmission to the receive device. In some examples, the source device and receive device can operate in a substantially symmetrical manner, such that each of the devices includes video encoding and decoding components. Hence, example systems can support one-way or two-way video transmission between video devices, e.g., for video streaming, video playback, video broadcasting, or video telephony.

[0083] As described above, in some examples, the SOC 100 and / or components thereof can be configured to perform video compression and / or decompression (also referred to as video encoding and / or decoding, collectively referred to as video coding) using machine learning techniques. For example, the encoding device 204 (or encoder) can be configured to encode video data using a machine learning system having a deep learning architecture, e.g., by leveraging the NPU 108 of the SOC 100. Figure 1 In some cases, using a deep learning architecture to perform video compression and / or decompression can improve the efficiency of video compression and / or decompression on a device. For example, the encoding device 204 can use machine learning-based video coding techniques to more efficiently compress video, can transmit the compressed video to the decoding device 212, and the decoding device 212 can use machine learning-based techniques to decompress the compressed video.

[0084] Neural networks are one example of machine learning systems, and a neural network can include an input layer, one or more hidden layers, and an output layer. Data is provided from input nodes of the input layer, processing is performed by hidden nodes of the one or more hidden layers, and an output is produced through output nodes of the output layer. Deep learning networks typically include multiple hidden layers. Each layer of a neural network can include a feature map or activation map, which can include artificial neurons (or nodes). A feature map can include filters, kernels, etc. A node can include one or more weights to indicate the importance of the node in one or more of the layers. In some cases, a deep learning network can have a series of many hidden layers, with early layers determining simple and low-level features of the input, and later layers building a hierarchy of more complex and abstract features.

[0085] Deep learning architectures can learn hierarchies of features. For example, if presented with visual data, a first layer can learn to recognize relatively simple features in the input stream, such as edges. In another example, if presented with auditory data, a first layer can learn to recognize spectral power in certain frequencies. A second layer, taking the output of the first layer as input, can learn to recognize combinations of features, such as simple shapes in visual data or combinations of sounds in auditory data. Higher layers can learn to represent complex shapes in visual data or words in auditory data, for example. Higher layers can also learn to recognize common visual objects or spoken phrases.

[0086] Deep learning architectures can perform particularly well when applied to problems with a natural hierarchical structure. For example, classification of motor vehicles can benefit from first learning to recognize wheels, windshields, and other features. These features can be combined in different ways at higher layers to recognize cars, trucks, and airplanes.

[0087] Neural networks can be designed with a variety of connection patterns. In feedforward networks, information passes from lower to higher layers, with each neuron in a given layer communicating with neurons in higher layers. A hierarchical representation can be built up in successive layers of a feedforward network, as described above. Neural networks can also have recurrent or feedback (also known as top-down) connections. In recurrent connections, the output from a neuron in a given layer can be passed to another neuron in the same layer. Recurrent architectures can be helpful in recognizing patterns that span more than one block of input data that is passed to the neural network in sequence. Connections from a neuron in a given layer to a neuron in a lower layer are called feedback (or top-down) connections. Networks with many feedback connections can be helpful when recognition of high-level concepts can help to distinguish particular low-level features of the input. Connections between layers of a neural network can be fully connected or locally connected. Various examples of neural network architectures are described below with respect to Figures 4A-5 Various examples of neural network architectures are described.

[0088] Figure 3A system 300 including a device 302 configured to perform video encoding using a machine-learned coding system 310 is depicted. The device 302 is coupled to a camera 307 and a storage medium 314 (e.g., a data storage device). In some implementations, the camera 307 is configured to provide image data 308 (e.g., a stream of video data) to the processor 304 for encoding by the machine-learned coding system 310. In some implementations, the device 302 can be coupled to and / or can include multiple cameras (e.g., a dual-camera system, three cameras, or other number of cameras). In some cases, the device 302 can be coupled to a microphone and / or other input devices (e.g., a keyboard, a mouse, a touch input device such as a touchscreen and / or touchpad, and / or other input devices). In some examples, the camera 307, the storage medium 314, the microphone, and / or other input devices can be part of the device 302.

[0089] The device 302 is also coupled to a second device 390 via a transmission medium 318 (e.g., one or more wireless networks, one or more wired networks, or a combination thereof). For example, the transmission medium 318 can include a channel provided by a wireless network, a wired network, or a combination of a wired network and a wireless network. The transmission medium 318 can form part of a packet-based network, such as a local area network, a wide-area network, or a global network such as the Internet. The transmission medium 318 can include routers, switches, base stations, or any other devices that can facilitate communication from a source device to a receiving device. The wireless network can include any wireless interface or combination of wireless interfaces and can include any suitable wireless network (e.g., the Internet or other wide-area networks, packet-based networks, WiFi™, radio frequency (RF), UWB, WiFi Direct, cellular, Long Term Evolution (LTE), WiMax™, etc.). The wired network can include any wired interface (e.g., fiber, Ethernet, power-line Ethernet, Ethernet over coaxial cable, digital signal line (DSL), etc.). Wired and / or wireless networks can be implemented using a variety of devices such as base stations, routers, access points, bridges, gateways, switches, etc. The encoded video bitstream data can be modulated according to a communication standard (e.g., a wireless communication protocol) and transmitted to a receiving device.

[0090] Device 302 includes one or more processors (referred to herein as "processors") 304 coupled to memory 306, a first interface ("I / F 1") 312, and a second interface ("I / F 2") 316. Processors 304 are configured to receive image data 308 from camera 307, from memory 306, and / or from storage medium 314. Processors 304 are coupled to storage medium 314 via first interface 312 (e.g., via a memory bus) and to transmission medium 318 via second interface 316 (e.g., a network interface device, a wireless transceiver and antenna, one or more other network interface devices, or a combination thereof).

[0091] Device 390 is similar to device 302 and includes one or more processors (referred to herein as "processors") 394 coupled to memory 392, a first interface ("I / F 1") 396, and a second interface ("I / F 2") 398. Processors 392 are configured to receive data from transmission medium 318, from memory 306, and / or from storage medium 314 via second interface 396. Processors 394 are coupled to storage medium 399 via first interface 398 (e.g., via a memory bus) and to transmission medium 318 via second interface 396 (e.g., a network interface device, a wireless transceiver and antenna, one or more other network interface devices, or a combination thereof).

[0092] Processors 304 include a machine learning encoding system 310. Machine learning encoding system 310 includes an encoder portion 362. Encoder portion 362 is configured to receive input data 370 and process input data 370 to generate encoding data 374 based at least in part on input data 370. In some cases, machine learning encoding system 310 can include both an encoder portion 362 and a decoder portion (e.g., decoder portion 366), which is shown here as being included in processors 394 of device 390. In some implementations, machine learning encoding system 310 can include one or more autoencoders.

[0093] In some implementations, encoder portion 362 of machine learning encoding system 310 is configured to perform lossy compression of input data 370 to generate encoding data 374 such that encoding data 374 has fewer bits than input data 370.

[0094] As shown, the encoder portion 362 of the machine learning transcoding system 310 can include a machine learning based preprocessor 363 and an encoder 364. As described above, some implementations of the processor 304 include multiple processors, and the elements of the encoder portion 362 can be included on (e.g., executed by) different ones of the multiple processors. For example, the encoder 364 can be included on a custom processor, such as an application specific integrated circuit (ASIC).

[0095] In some cases, while video consumption is growing very rapidly, the technical requirements are becoming more and more demanding as resolutions, frame rates, and dynamic ranges continue to increase. To provide the desired quality of service as well as the required data processing speed, high throughput, and low power consumption, it is preferable to use custom hardware to support current consumer video applications, such as ASICs that implement video compression standards such as H.264 / AVC and H.265 / HEVC.

[0096] In some cases, updating ASICs can be difficult and relatively slow, which can limit opportunities to improve existing ASICs as new technologies and features become available. Machine learning and neural network implementations can be used to help enhance the performance of existing ASICs using techniques referred to herein as video codec neural boosting. The machine learning based preprocessor 363 can be used to perform video codec neural boosting.

[0097] In some implementations, the machine learning based preprocessor 363 can be included on a machine learning accelerator or other processor optimized for performing machine learning / artificial intelligence processing. The machine learning based preprocessor 363 can include one or more convolutional neural networks (CNNs), one or more fully connected neural networks, one or more gated recurrent units (GRUs), one or more long short-term memory (LSTM) networks, one or more ConvRNNs, one or more ConvGRUs, one or more ConvLSTMs, one or more GANs, any combination thereof, and / or other types of neural network architectures. The machine learning based preprocessor 363 can receive, as input, encoder parameters 366 for configuring the encoder 364. The encoder parameters 366 can include parameters for configuring the operation of the encoder 364, such as quality parameters (e.g., quality settings, parameters to enable / disable optimizations, or other parameters for controlling how the encoder 364 operates). In some cases, the machine learning based preprocessor 363 can also receive, as input, codec internal data 368 from the encoder 364. The codec internal data 368 can include information or settings for performing one or more steps of the encoding process, such as scaling transform coefficients, quantizer step sizes, etc. Based on the inputs (e.g., the input data 370, the encoder parameters 366, and / or the codec internal data 368), the machine learning based preprocessor 363 can generate codec boost information 372 (e.g., control side information and / or pixel side information) for improving the performance of the encoder 364. The machine learning based preprocessor 363 can pass the codec boost information 372 along with the encoded data 374 to the decoder portion 366 for decoding. The machine learning based preprocessor 363 can pass the encoder parameters 366 along with the input data 370 to the encoder 364 for encoding. In some cases, the encoder 364 can encode the input data 370 based on the encoder parameters 366 using existing video compression standards such as H.264 / AVC and H.265 / HEVC, and generate the encoded data 374. The encoded data 374 can be transmitted to the device 390 over the transmission medium 318 via the second interface 316.

[0098] The processor 394 of the device 390 includes the machine learning decoding system 350. The machine learning decoding system 350 includes the decoder portion 366. The device 390 can receive the encoded data 374 and the codec boost information 372 via the second interface 396 and pass the encoded data 374 and the codec boost information 372 to the decoder portion 366. In this example, the decoder portion 366 includes the decoder 352 and the machine learning based post-processor 354. The decoder portion 366 is configured to receive the encoded data 376 and process the encoded data 376 to generate a representation 378 based on the input data 370, which can be displayed to a user, e.g., via a display (not shown). In some cases, the decoder 352 can decode the encoded data 376 using an existing video compression standard, such as H.264 / AVC and H.265 / HEVC, and generate decoded data 356. The decoded data 356 can be passed to the machine learning based post-processor 354. In some cases, the machine learning decoding system 350 can include both an encoder portion (e.g., the encoder portion 362) and the decoder portion 366.

[0099] The machine learning based post-processor 354 can receive the decoded data 356 and the codec boost information 372. The machine learning based post-processor 354 can enhance the decoded 356 using one or more trained machine learning algorithms based on the input codec boost information 372 to generate the representation 378. In some examples, the machine learning based post-processor 354 of the decoder portion includes a neural network, which can include one or more CNNs, one or more fully connected neural networks, one or more GRUs, one or more long short-term memory (LSTM) networks, one or more ConvRNNs, one or more ConvGRUs, one or more ConvLSTMs, one or more GANs, any combination thereof, and / or other types of neural network architectures. The processor 394 can be configured to send the representation 378 to the storage medium 399 or output the representation 378 for display on a display (not shown).

[0100] The components of the system 300 can include and / or be implemented using electronic circuitry or other electronic hardware, which can include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits), and / or can include and / or be implemented using computer software, firmware, or any combination thereof, to perform various operations described herein.

[0101] Although system 300 is shown to include certain components, those skilled in the art will understand that system 300 may include more than [other components]. Figure 3 The components shown may include more or fewer components. For example, system 300 may also include or be part of a computing device that includes input devices and output devices (not shown). In some embodiments, system 300 may also include or be part of a computing device that includes one or more memory devices (e.g., one or more random access memory (RAM) components, read-only memory (ROM) components, cache memory components, buffer components, database components, and / or other memory devices), one or more processing devices (e.g., one or more CPUs, GPUs, and / or other processing devices) that communicate with and / or are electrically connected to one or more memory devices, one or more wireless interfaces for performing wireless communication (e.g., one or more transceivers and baseband processors for each wireless interface), one or more wired interfaces for performing communication via one or more hardwired connections (e.g., serial interfaces such as Universal Serial Bus (USB) inputs, lighting connectors, and / or other wired interfaces), and / or Figure 3 Other components not shown.

[0102] In some implementations, system 300 may be implemented locally by and / or included in a computing device. For example, a computing device may include a mobile device, a personal computer, a tablet computer, a virtual reality (VR) device (e.g., a head-mounted display (HMD) or other VR device), an augmented reality (AR) device (e.g., an HMD, AR glasses, or other AR device), a wearable device, a server (e.g., in a Software as a Service (SaaS) system or other server-based system), a television set, and / or any other computing device with the resource capability to perform the techniques described herein.

[0103] As previously noted, some video coding systems utilize neural networks or other machine learning systems to compress video and / or image data. Neural networks can be designed to have a variety of connection patterns. In feed-forward networks, information passes from lower to higher layers, with each neuron in a given layer communicating with neurons in higher layers. A hierarchical representation can be built up in successive layers of a feed-forward network, as described above. Neural networks can also have recurrent or feedback (also known as top-down) connections. In recurrent connections, the output from a neuron in a given layer can be passed to another neuron in the same layer. Recurrent architectures can be helpful in recognizing patterns that span more than one block of input data passed sequentially to a neural network. Connections from a neuron in a given layer to a neuron in a lower layer are referred to as feedback (or top-down) connections. Networks with many feedback connections can be helpful when recognition of high-level concepts can be helpful in distinguishing particular low-level features of an input.

[0104] Connections between layers of a neural network can be fully connected or locally connected. Figure 4A An example of a fully connected neural network 402 is shown. In a fully connected neural network 402, a neuron in a first layer can pass its output to every neuron in a second layer, so that every neuron in the second layer will receive input from every neuron in the first layer. Figure 4B An example of a locally connected neural network 404 is shown. In a locally connected neural network 404, a neuron in a first layer can be connected to a limited number of neurons in a second layer. More generally, locally connected layers of a locally connected neural network 404 can be configured so that each neuron in a layer will have the same or a similar connection pattern, but with connection strengths that can have different values (e.g., 410, 412, 414, and 416). The locally connected connection pattern can result in spatially different receptive fields in higher layers, as higher layer neurons in a given region can receive input tuned through training to attributes of a restricted portion of the total input to the network.

[0105] One example of a locally connected neural network is a convolutional neural network. Figure 4C An example of a convolutional neural network 406 is shown. A convolutional neural network 406 can be configured so that the connection strengths associated with the input to each neuron in a second layer are shared (e.g., 408). Convolutional neural networks can be well suited to problems where the spatial location of input is meaningful. According to aspects of the present disclosure, a convolutional neural network 406 can be used to perform one or more aspects of video compression and / or decompression.

[0106] One type of convolutional neural network is a deep convolutional network (DCN). Figure 4DA detailed example of a DCN 400 is shown, which is designed to recognize visual features from an image 426 input from an image capture device 430 (e.g., an in-vehicle camera). The DCN 400 in this example can be trained to recognize traffic signs and the numbers provided on them. Of course, the DCN 400 can be trained for other tasks, such as recognizing lane markings or traffic lights.

[0107] Supervised learning can be used to train the DCN 400. During training, an image, such as image 426 of a speed limit sign, can be presented to the DCN 400, and the forward pass can be computed to produce output 422. The DCN 400 may include a feature extraction part and a classification part. Upon receiving image 426, convolutional layer 432 may apply a convolutional kernel (not shown) to image 426 to generate a first feature map set 418. As an example, the convolutional kernel of convolutional layer 432 may be a 5x5 kernel that generates a 28x28 feature map. In this example, because four different feature maps are generated in the first feature map set 418, four different convolutional kernels are applied to image 426 at convolutional layer 432. The convolutional kernel may also be referred to as a filter or convolutional filter.

[0108] The first feature map set 418 can be second-sampled by a max-pooling layer (not shown) to generate a second feature map set 420. The max-pooling layer reduces the size of the first feature map set 418. That is, the size of the second feature map set 420 (e.g., 14x14) is smaller than the size of the first feature map set 418 (e.g., 28x28). The reduced size provides similar information to subsequent layers while reducing memory consumption. The second feature map set 420 can be further convolved via one or more subsequent convolutional layers (not shown) to generate one or more subsequent feature map sets (not shown).

[0109] exist Figure 4D In the example, the second feature map set 420 is convolved to generate a first feature vector 424. Furthermore, the first feature vector 424 is further convolved to generate a second feature vector 428. Each feature of the second feature vector 428 may include a number corresponding to a possible feature of the image 426, such as "symbol", "60", and "100". A softmax function (not shown) can convert the numbers in the second feature vector 428 into probabilities. Therefore, the output 422 of the DCN 400 is the probability that the image 426 includes one or more features.

[0110] In this example, the probabilities of "symbol" and "60" in the output 422 are higher than the probabilities of other in the output 422, e.g., "30," "40," "50," "70," "80," "90," and "100." Before training, the output 422 produced by the DCN 400 can be incorrect. Thus, an error between the output 422 and a target output can be computed. The target output is the ground truth of the image 426 (e.g., "symbol" and "60"). The weights of the DCN 400 can then be adjusted so that the output 422 of the DCN 400 more closely aligns with the target output.

[0111] To adjust the weights, a learning algorithm can compute a gradient vector of the weights. The gradient can indicate an amount by which the error will increase or decrease when the weights are adjusted. At the top layer, the gradient can directly correspond to the values of the weights connecting the activated neurons in the penultimate layer and the neurons in the output layer. In lower layers, the gradient can depend on the values of the weights and the computed error gradient of the higher layers. The weights can then be adjusted to reduce the error. This way of adjusting the weights can be referred to as "backpropagation" because it involves a "backward pass" through the neural network.

[0112] In practice, the error gradient of the weights can be computed over a small number of examples so that the computed gradient is close to the true error gradient. This approximation method can be referred to as stochastic gradient descent. Stochastic gradient descent can be repeated until the achievable error rate of the entire system has stopped decreasing or until the error rate has reached a target level. After learning, new images can be presented to the DCN and the forward pass through the network can produce an output 422 that can be considered an inference or prediction of the DCN.

[0113] A deep belief network (DBN) is a probabilistic model that includes multiple layers of hidden nodes. A DBN can be used to extract a hierarchical representation of a training dataset. A DBN can be obtained by stacking layers of restricted Boltzmann machines (RBMs). An RBM is an artificial neural network that can learn a probability distribution over a set of inputs. Because an RBM can learn a probability distribution without information about the class to which each input belongs, RBMs are often used for unsupervised learning. Using a hybrid unsupervised and supervised paradigm, the bottom RBMs of a DBN can be trained in an unsupervised manner and can be used as feature extractors, while the top RBMs can be trained in a supervised manner (on the joint distribution of inputs from the previous layer and target classes) and can be used as classifiers.

[0114] A deep convolutional network (DCN) is a network of convolutional networks that are configured with additional pooling and normalization layers. DCNs have achieved state-of-the-art performance on many tasks. A DCN can be trained using supervised learning, in which both the input and the output target are known for many examples, and used to modify the weights of the network by using a gradient descent method.

[0115] A DCN can be a feedforward network. Further, as described above, connections from a neuron in a first layer of a DCN to a group of neurons in a next higher layer are shared among the neurons in the first layer. The feedforward and shared connections of a DCN can be used for fast processing. For example, the computational burden of a DCN can be much less than that of a similar sized neural network that includes recurrent or feedback connections.

[0116] The processing of each layer of a convolutional network can be thought of as a spatially invariant template or basis projection. If the input is first decomposed into multiple channels (e.g., the red, green, and blue channels of a color image), then a convolutional network trained on that input can be thought of as three-dimensional, with two spatial dimensions along the axes of the image and a third dimension that captures color information. The output of a convolutional connection can be thought of as forming a feature map in a subsequent layer, with each element of the feature map (e.g., 420) receiving input from a series of neurons in the previous layer (e.g., feature map 418) and from each of the multiple channels. The values in the feature map can be further processed with a nonlinearity (e.g., rect, max(0,x)). Values from adjacent neurons can be further combined, which corresponds to downsampling, and can provide additional local invariance and dimensionality reduction.

[0117] Figure 5 is a block diagram illustrating an example of a deep convolutional network 550. The deep convolutional network 550 can include multiple different types of layers based on connections and weight sharing. As shown, Figure 5 The deep convolutional network 550 includes convolutional blocks 554A, 554B. Each of the convolutional blocks 554A, 554B can be configured with a convolutional layer (CONV) 556, a normalization layer (LNorm) 558, and a max pooling layer (MAX POOL) 560.

[0118] The convolutional layers 556 can include one or more convolutional filters that can be applied to the input data 552 to generate feature maps. Although only two convolutional blocks 554A, 554B are shown, the present disclosure is not so limited, but rather any number of convolutional blocks (e.g., blocks 554A, 554B) can be included in the deep convolutional network 550 according to design preference. The normalization layers 558 can normalize the output of the convolutional filters. For example, the normalization layers 558 can provide whitening or lateral inhibition. The max-pooling layers 560 can provide spatially down-sampling aggregation for local invariance and dimension reduction.

[0119] For example, the parallel filter banks of the deep convolutional network can be loaded on the CPU 102 or GPU 104 of the SOC 100 to achieve high performance and low power consumption. In alternative embodiments, the parallel filter banks can be loaded on the DSP 106 or ISP 116 of the SOC 100. Further, the deep convolutional network 550 can access other processing blocks that can be present on the SOC 100, such as the sensor processor 114 and the navigation module 120 that are specialized for sensors and navigation, respectively.

[0120] The deep convolutional network 550 can also include one or more fully connected layers, such as layer 562A (labeled “FC1”) and layer 562B (labeled “FC2”). The deep convolutional network 550 can also include a logistic regression (LR) layer 564. Between each layer 556, 558, 560, 562A, 562B, 564 of the deep convolutional network 550 are weights (not shown) to be updated. The output of each of the layers (e.g., 556, 558, 560, 562A, 562B, 564) can be used as input to a subsequent one of the layers (e.g., 556, 558, 560, 562A, 562B, 564) in the deep convolutional network 550 to learn a hierarchical feature representation from the input data 552 (e.g., images, audio, video, sensor data, and / or other input data) provided at the first one of the convolutional blocks 554A. The output of the deep convolutional network 550 is a classification score 566 of the input data 552. The classification score 566 can be a set of probabilities, where each probability is a probability that the input data includes a feature from a set of features.

[0121] In some cases, traditional video compression techniques are designed empirically using human knowledge and intuition, while neural networks and other machine learning tools rely on large amounts of training data and efficient learning algorithms. Generally speaking, there are two main categories of learning algorithms: (1) techniques based on discrete trial-and-error testing with on-the-fly generation of optimization strategies, and (2) techniques that compute the gradient of an objective (or loss) function and use automatic differentiation in the form of backpropagation to optimize system parameters. Both approaches have been successfully used to solve practical problems, but the availability of the derivatives of the performance metrics generally makes the second approach more efficient and thus more commonly used. For example, the second approach is the one employed in popular development tools such as TensorFlow and PyTorch.

[0122] One problem with video coding systems is that several operations implemented by standard video codecs are not differentiable. Moreover, video coding systems sometimes include several stages, such as the adaptive deblocking filter, the adaptive loop filter, etc., that can be linear and differentiable, but are quite complex and difficult to integrate into a training procedure.

[0123] In some cases, when designing a neural boosting system (e.g., in the learning (or training) phase), the compressed data and decoded video of the codec are not directly used, because only the measurement of the performance of the codec and its derivatives are needed for the optimization. Therefore, if a good enough estimate of the measurement and its derivatives can be obtained, the full codec is not needed. Systems designed for this purpose are referred to herein as differentiable codec surrogates.

[0124] As mentioned above, entropy coding is one of the last stages of encoding (compression) (and in some cases the last stage), which defines the values and the number of bits to be added to the compressed data bitstream. Modern standard-based video encoding methods (e.g., VVC, HEVC, AV1, etc.) employ adaptive arithmetic coding to achieve high-quality compression performance. The bitstream generated by adaptive arithmetic coding can only be sequentially encoded and decoded. For example, a data element can only be recovered by first decoding all previous elements, because the decoder needs to reach the same state that the encoder had when coding that element.

[0125] Figure 6Ais a block diagram 600 illustrating an example implementation of a video codec neural boost system in accordance with aspects of the present disclosure. In the example video codec neural boost system, when deploying a codec with neural boosting inside the video codec neural boost system, a standard codec (e.g., VVC, HEVC, AVC, MPEG, AV1, or other similar codec) can be used. Video data 604 can be input to an ML pre-processing engine 606. The ML pre-processing engine 606 can include one or more ML models. The one or more ML models can perform a variety of pre-processing tasks. Examples of pre-processing tasks can include spatial upsampling, temporal upsampling, quality optimization, compression artifact removal, selective denoising, dynamic frame size adjustment, etc. Generally, these pre-processing tasks aim to improve the quality of the video and / or the compression of the video without modifying the video encoder 608, which can be a video encoder (e.g., VVCV, HEVC, AVC, MPEG, AV1, or other similar encoder) that complies with existing standards. For example, one pre-processing task can be to apply an ML model to perform selective denoising to eliminate types of noise that the encoder can not be able to handle well, thereby improving the compression of the video. As another example, the ML model can dynamically adjust the encoder parameters 614 based on the input video data 604 to improve the compression or quality of the video. The compressed video 610 can be sent 612 to a playback device (which can be any device, including an encoding device) for decoding by a video decoder 616. In some cases, the video decoder 616 can be a standard decoder (e.g., VVCV, HEVC, AVC, MPEG, AV1, or other similar decoder). The video data 604 output by the video decoder 616 can be input to an ML post-processing engine 618. The ML post-processing engine 618 can perform various post-processing tasks. For example, the ML post-processing engine 618 can apply one or more ML models to dynamically add noise to the output video that is similar to the noise removed by the ML pre-processing engine 606.

[0126] Figure 6Bis a block diagram 650 illustrating the training of a video codec neural boost system according to aspects of the present disclosure. In some cases, during a network learning (training) phase, the codec is replaced by a differentiable codec proxy 652 that enables gradient backpropagation 654. Video data 654 can be input to an ML pre-processing engine 656. In some cases, the ML pre-processing engine 656 can be the ML pre-processing engine 606 that is being trained. The output from the ML pre-processing engine 656 is input to the differentiable codec proxy 652. The differentiable codec proxy 652 can estimate the bit rate 670 of the compressed video that a video encoder (e.g., video encoder 608) can output. The differentiable codec proxy 652 can pass the estimated bit rate 670 to a loss measure engine 668. The differentiable codec proxy 652 can also pass the video data 654 to an ML post-processing engine 668. In some cases, the ML post-processing engine 668 can be the ML post-processing engine 618 that is being trained. The output from the ML post-processing engine 668 can be passed to the loss measure engine 668. The loss measure engine 668 can compare the output from the ML post-processing engine 668 and the estimated bit rate 670 to ground truth references to compute a loss function and gradients 672. The gradients 672 can be backpropagated to the ML pre-processing engine 656, the differentiable codec proxy 652, and the ML post-processing engine 668 for training.

[0127] In some cases, the efficacy of using the differentiable codec proxy 652 depends on the accuracy of the proxy estimates, which are defined by two factors, a selected quality parameter, referred to here as QP, controls. The bit rate R (which is a function of QP) corresponds to the number of bits used by the encoder to compress a video frame or a block within a frame, divided by the number of pixels used to normalize to bits per pixel. The bit rate decreases as QP increases. The distortion D (which is also a function of QP) is a measure of the difference between the original video pixels and the decoded video pixels, also normalized per pixel. For example, the distortion can correspond to a mean squared error or a more complex measure that approximates human subjective preference. The distortion increases as QP increases.

[0128] The bit rate and the distortion can be combined into a single loss function, such as the loss function defined in Equation 1 below:

[0129] L(QP) = D(QP) + λ(QP)R(QP),

[0130] Equation (1)

[0131] where λ(QP) is a Lagrange multiplier defined in terms of the definition of QP in the video standard along with the quantization step size.

[0132] One observation is that distortion measures are strongly dependent on the quantization step size predefined by the encoder parameter QP, and are thus more predictable and easier to estimate. On the other hand, bitrates are more difficult to estimate as they are subject to complex statistical dependencies between many pixels.

[0133] Some video encoders, such as some HEVC / H.265 standard video encoders, implement a so-called hybrid coding configuration that combines predictive coding and transform coding. The components that perform prediction and transform are linear and thus differentiable and easier to directly integrate into a training process.

[0134] In some cases, some encoder elements are non-linear and non-differentiable, such as those defined by the quantization and entropy coding processes. As used herein, non-differentiable can refer to a process that cannot be usefully differentiated with respect to gradient backpropagation and / or loss measurement. For example, a function that rounds a decimal number can be a basic form of quantization, where an input decimal number (e.g., 3.14 and 2.56) is rounded to 3. The derivative of the output of the function with respect to the input range (e.g., a set of decimal numbers from 2.51 to 3.50 (assuming two decimal places)) would simply be zero from 2.51 to 3.49, and then infinite for 3.50. Such a differential output can be useless for gradient backpropagation because there is no real gradient.

[0135] Figure 7 is a block diagram illustrating encoder elements of a video encoder 700 (e.g., an HEVC or H.265 encoder) in accordance with aspects of the present disclosure. In some cases, certain encoder elements can be non-linear and non-differentiable, such as the quantization 702 and entropy coding 704 processes. Assuming a block has size M x N pixels, a linear transform (typically a discrete cosine transform) is applied to the residual (e.g., an array of differences between predicted pixel values and actual pixel values). The resulting array of transform coefficients d_(m,n) is then scaled (divided by a positive quantizer step size s 706), m = 0, 1,..., M - 1, n = 0, 1,..., N - 1, and quantized according to Equation 2 below:

[0136] q m,n = sign(c m,n )[|c m,n | + ξ],

[0137] Equation (2)

[0138] where ξ is an offset that defines a dead-zone quantization type.

[0139] The quantized array of transform coefficients q m,nentropy coded, defines the number of bits used to encode the block. Since quantization is not differentiable, any proxy estimate needs to use values from the scaled transform coefficient array c m,n , instead of q m,n .

[0140] In some cases, an image coding differentiable proxy can estimate the number of bits as a function, for example, Equation 3 below:

[0141]

[0142] where μ is a normalization constant.

[0143] However, current video coding standards employ complex binarization adaptive arithmetic coding, and they exploit statistical dependencies among elements of the array c m,n by processing the array in multiple passes, using coding contexts. Therefore, any estimator like the above equation that does not take into account the inter-coefficient dependencies will have very limited accuracy.

[0144] As described above, in accordance with aspects of the present disclosure, systems and techniques are described herein that at least address this problem by defining a joint statistical model for all values in the array c m,n , computing maximum likelihood estimates of the model parameters, and then using the model to estimate the number of bits. Such a process is differentiable.

[0145] Figure 8 is a block diagram illustrating elements of a differentiable encoder proxy 800, in accordance with aspects of the present disclosure. In the differentiable encoder proxy 800, the non-differentiable quantization 702 and entropy coding 704 elements of an encoder (e.g., video encoder 700) can be replaced by a differentiable rate estimation process of the encoder proxy 800. The differentiable rate estimation process includes a noise and adjustment process 802, a parameter estimation process 804, and a block mode 806. The proxy implementation assumes that it approximates the main encoder decisions, such as selecting the correct prediction 808, block size (M, N), and transform type 810.

[0146] Regularization and coefficient adjustment will now be described. For example, in some cases, quantization can alter the scaled transform coefficients in a process similar to adding random quantization noise η with roughly uniform distribution in the interval [-0.5, 0.5]. For this purpose, two arrays of random numbers with uniform distribution can be defined as In some cases, for the two arrays of random values, one can be chosen to be non-zero, while the other can be chosen to be zero.

[0147] In some cases, this technique of adding random numbers is used when training the end-to-end neural codec to account for quantization of latent variables. However, in this example, it is included because it was empirically observed to improve estimation accuracy, and because it can also be used as a form of regularization, as it avoids numerical instability during estimation when all transform coefficients have very small or zero magnitude.

[0148] Another factor is that, although not required by the standard, video codecs often use a dead-zone quantizer (e.g., Equation 2), which increases the probability of quantization to zero values. To approximate this feature, an “adjustment” function is defined as Equation 4 below:

[0149]

[0150] where x is the argument of the function, and A and an integer K are positive constants (parameters). In some cases, the adjustment function can be updated, for example, during a training process of the encoder agent. For example, the values of A and K or x can be adjusted based on a loss function.

[0151] An example derivative of Equation 4 is illustrated in Equation 5 below:

[0152]

[0153]

[0154] Figure 9 is a plot of an example adjustment function and derivative of the adjustment function 900, in accordance with aspects of the present disclosure. In this example, Figure 9 A first line 902 is shown corresponding to Equation 4, given According to these definitions, a random variable can be defined in a statistical model for bit rate estimation, for example, as shown in Equation 6 below:

[0155]

[0156] Figure 9 A second line 904 is shown corresponding to the derivative of Equation 4 (e.g., Equation 5).

[0157] Definitions of the statistical model and parameter estimation will now be discussed. For example, in some cases, the statistical model can be defined with a vector x containing three parameters of the model, and can be based on one or more assumptions. A first assumption for which the statistical model can be defined is that all elements of the array t m,n are independent and have a Laplace distribution (which has a zero mean and a scale s m,n (x)). For example, the probability distribution function of the elements of the array t m,n may be described by Equation 7 below:

[0158]

[0159] For simplicity of notation, the Laplace distribution can be parameterized by scaling s instead of the standard deviation

[0160] A second assumption that can be defined for the statistical model is that the array of Laplace distribution scales is defined by equation 8 as follows:

[0161]

[0162] In some cases, the probability distribution of the transform coefficients can be well approximated by a Laplace distribution. The fast decay of the scale with frequency (defined by the indices m and n) is also well known. In some cases, the choice of the exponential function to model the decay greatly simplifies the derivative formula and its computation. In some cases, the values of the indices m and n are updated, for example, during the training process of the encoder agent. For example, the indices m and n can be adjusted during training based on a loss function.

[0163] Model parameter estimation will now be described. For example, given the array of modified transform coefficients t m,n in a block, the parameter vector x can be estimated using a maximum likelihood (ML) method. As explained above, by replacing the array t ′ m,n with the array t m,n defined in equation 6, the prediction accuracy can be improved and numerical instability can be avoided. For simplicity of notation, equations 9 and 10 can be defined as follows:

[0164]

[0165] From equations 6, 8, 9, and 10, the negative of the log-likelihood (e.g., a function to be minimized to obtain the optimal parameter vector) can be shown, for example, as shown in equation 11 below:

[0166]

[0167] The optimal solution corresponding to the maximum likelihood parameters is defined by setting the gradient of equation 11 to zero, which can be shown to correspond to the first set of equations as follows:

[0168] The first set of equations does not have a known closed-form solution, but the solution can be efficiently computed using a Newton method iteration, defined by equation 12 below:

[0169]

[0170] ​where 0 < ξ < 1 is a multiplicative factor added to ensure convergence.

[0171] In Equation 12, the gradient is:

[0172]

[0173] and the corresponding Hessian matrix is:

[0174]

[0175] In some cases, the output gradient of Equation 12 can be determined and used, e.g., by a loss measurement system (e.g., loss measurement engine 668 of FIG. 6) for backpropagation as part of a loss function for training a preprocessor (e.g., ML preprocessor engine 656) or a differentiable codec agent (e.g., differentiable codec agent 652).

[0176] Some observations about practical applications of the Newton method are as follows:

[0177] (1) As mentioned previously, the use of the exponential function simplifies the derivative formula and its computation because several factors appear repeatedly in the above equations. For example, the array with factors {w m,n ,mw m,n ,nw m,n ,m 2 w m,n ,n 2 w m,n ,mn w m,n} can be computed once and reused in each iteration.

[0178] (2) Divergence can be avoided by using an adaptive step-size correction method, e.g., for each iteration, first try ξ = 1, and if L(x) does not decrease, halve its value.

[0179] (3) Since the Hessian matrix is symmetric, it is more efficient to compute the Newton method step size in Equation 12 using Cholesky matrix decomposition instead of inverse matrix decomposition.

[0180] Experimental results show that when a reasonably good initial solution is used, quadratic convergence starts after only 2 or 3 iterations. For example, in experiments using 8 x 8 DCT coefficients, the following initial solution can be used:

[0181]

[0182] Bit rate estimation will now be described. As mentioned previously, Figure 7As shown, entropy coding 704 in video compression is applied to the integers obtained by quantization, where a probability value is assigned to each quantization interval. In some cases, to approximate this stage with a differentiable method, a relaxation technique developed for training end-to-end neural image codecs can be employed, where the fixed set of intervals for quantization and probability computation is replaced by intervals around the coefficient values.

[0183] With this approach, the probability value corresponding to the value to be used for entropy coding can be defined by Equation 13 below:

[0184]

[0185] In Equation 13, the cumulative distribution of the Laplace distribution can be defined as defined in Equation 14 below:

[0186]

[0187] With these definitions, an estimate of the number of bits used to code a block can be defined by Equation 15 below:

[0188]

[0189] where a is a normalization constant.

[0190] In some cases, the estimated number of bits can be generated, for example, by a differentiable codec agent (e.g., differentiable codec agent 652 of FIG. 6). For example, a loss measurement system (e.g., loss measurement engine 668) can use the estimated number of bits for backpropagation.

[0191] Now, derivative computation will be described. To train the neural network, the gradient (e.g., an array of partial derivatives) of the estimated number of bits can be used according to Equation 16 below:

[0192]

[0193] In some cases, the array of partial derivatives can be computed using automatic differentiation tools available in software tools such as PyTorch, but it is noted that to enable backpropagation, Equation 12 can need to be replaced with l = 0, 1, 2,..., creating an aggregated sequence of solutions and allowing the software to create the correct links for automatic differentiation. Such techniques can increase computational complexity as it adds multiple stages of gradient computation. If automatic differentiation is removed from the maximum likelihood optimization and direct computation of those partial derivatives using the optimal solution is used, a more efficient implementation can be performed.

[0194] To compute those derivatives, one can define And

[0195] Or partial derivative calculations, optimality conditions Can correspond to the following set of MN vector equations:

[0196]

[0197] Where,

[0198]

[0199] Using such results, equation 16 can be expanded as:

[0200]

[0201] Into equations defining gradient calculations:

[0202]

[0203] Where,

[0204] Then, from equations 5 and 6, equation 17 can be derived as follows:

[0205]

[0206] In some aspects, equation (17) can be simplified by collecting sums of terms as a function of indices m, n. For example, the following functions can be defined:

[0207]

[0208] Ψ m,n (b) = (t m,n +b) ψ m,n (b), equation (20)

[0209] And vector z such that:

[0210]

[0211] Thus, equation (17) can correspond to:

[0212]

[0213] An illustrative example of experimental verification will now be described. For example, the system and techniques described herein can be tested by modifying a reference implementation of the HEVC / H.265 video compression standard, known as HM version 20.0, to output the DCT coefficients used by the encoder in a low-latency configuration, with the luma transformation dimension limited to 8×8 blocks, QP values ​​of 22, 27, 32, and 37, and the number of bits used in each case.

[0214] Compression was applied to 201 frames of 10 HD test videos, and only the luminance coefficients between frames (e.g., excluding intra-frames) were preserved. For each frame, ∈ in equation (5) was used. (c) =0.5,∈ (t) =0, A=1, K=2 in Equation 5 to estimate the bit rate, and α=1 was initially used in Equation 15. After comparing the estimated bit rate with the actual bit rate, the parameter α was calibrated to α=5 / 3, which is the value used in all the results presented thereafter.

[0215] In some cases, the choice of parameter α is not important because the bit rate estimate is multiplied by a constant during neural network training, as shown in Equation 1, and several constant values ​​need to be tested to obtain optimal results. Similarly, the estimator calibration form HEVC can be used with AVC by testing new scaling factors.

[0216] The results obtained using the system and techniques described herein can be evaluated by measuring the ratio of the estimated frame bit rate to the corresponding value from the HM software. Figures 10-13 Histograms of these ratios calculated for 2,000 frames of the test video are shown, with each plot displaying results for QP = 22, 27, 32, and 37. As expected, the results peak at low QP values, corresponding to the high-rate case where the per-coefficient factor dominates. As the QP value increases, the histograms become more dispersed, and well-constrained errors are observed even for the maximum QP = 27.

[0217] Figure 10 It is a histogram of the ratio between the frame bit rate estimated using the system and techniques described herein and the bit rate used by the HM software (HEVC / H.265 reference implementation) and QP = 22.

[0218] Figure 11 It is a histogram of the ratio between the frame bit rate estimated using the system and techniques described herein and the bit rate used by the HM software (HEVC / H.265 reference implementation) and QP = 27.

[0219] Figure 12is a histogram of the ratio between the frame bitrates estimated using the systems and techniques described herein and the bitrates used by the HM software (HEVC / H.265 reference implementation) for QP = 22.

[0220] Figure 13 is a histogram of the ratio between the frame bitrates estimated using the systems and techniques described herein and the bitrates used by the HM software (HEVC / H.265 reference implementation) for QP = 37.

[0221] Figure 14 is a histogram of the ratio between the frame bitrates estimated using the techniques for bitrate estimation using machine learning to enhance video coding and the bitrates used by the HM software (HEVC / H.265 reference implementation) for the combined results for QP = 22, 27, 32, and 37.

[0222] Figure 14 The combined results for all QP values are shown, which can be compared to the values estimated using the non-differentiable AGP entropy coding method shown in Figure 15 which uses a much simpler way to model the transform coefficients than the method used by HEVC. Finally, Figure 16 The results obtained with the per-coefficient differentiable estimation of Equation 3 are shown, where the results are widely distributed, indicating much lower precision.

[0223] Figure 15 is a histogram of the ratio between the frame bitrates estimated using the AGP entropy coding method (non-differentiable) and the bitrates used by the HM software (HEVC / H.265 reference implementation) for the combined results for QP = 22, 27, 32, and 37.

[0224] Figure 16 is a histogram of the ratio between the frame bitrates estimated using the per-coefficient differentiable estimation of Equation 3 and the bitrates used by the HM software (HEVC / H.265 reference implementation) for the combined results for QP = 22, 27, 32, and 37.

[0225] For reference, Table 1 shows the average bitrates for different QP values, indicating that they roughly cover one order of magnitude.

[0226] Quality setting QP = 22 QP = 27 QP = 32 QP = 37 Average bitrate 0.045 bits / pixel 0.11 bits / pixel 0.23 bits / pixel 0.48 bits / pixel

[0227] Table 1

[0228] This disclosure addresses various issues that arise when neural networks are used to enhance the performance of standard video codecs in several ways (referred to as standard video codec neural boosting) and when the gradient of performance measurements cannot be backpropagated through the codec, thus hindering overall optimization.

[0229] The aspects of this disclosure solve the problem of reliably estimating the bit rate used by the video encoder while simultaneously estimating the corresponding derivative and allowing gradient backpropagation to be used during end-to-end training.

[0230] This estimate is more accurate and reliable because, similar to entropy decoding in a standard encoder, it incorporates information from many discrete cosine transform (DCT) coefficients used for compression. This is achieved using a new form of statistical model that assumes the DCT coefficients have a Laplace distribution and three parameters determined using the maximum likelihood criterion.

[0231] The model generates scaling parameters for the Laplace distribution, a process similar to that of entropy decoding super-prior neural networks in end-to-end neural video codecs, and this similarity simplifies the integration of the systems and techniques described in this paper into systems used for neural network training.

[0232] Experimental results comparing the frame bitrate obtained using the system and techniques described in this paper with precise values ​​from the HEVC / H.265HM codec demonstrate that the method provides accuracy that is very similar to nonspecific and nondifferentiable entropy decoding methods and is much more accurate than methods that use per-coefficient estimation.

[0233] Because a mathematical statistical model is used, it is possible to derive formulas that define all necessary derivatives, and all necessary derivatives can be computed more efficiently using C++ or CUDA GPU implementations, rather than automatic differentiation.

[0234] Figure 17 This is a flowchart illustrating techniques for performing bitrate estimation 1700 according to various aspects of this disclosure. At operation 1702, technique 1700 may include: encoding one or more frames of video data using a video encoder, the video encoder including at least a quantization process, such as... Figures 6A-6B and Figure 7 As shown. At operation 1704, technique 1700 may include: determining as Figures 6A-6B and Figure 7 The actual bit rate of one or more encoded frames shown.

[0235] At operation 1706, technique 1700 may include: using an encoder agent to predict an estimated bit rate, the encoder agent including a statistical model for estimating the bit rate of one or more encoded frames, such as... Figure 8As shown. In some cases, the encoder agent estimates an output of one or more processes of the video encoder. In some cases, the statistical model is based on a Laplacian distribution of coefficients input to a quantization process of the video encoder. In some cases, the technique 1700 can further include estimating an output of the quantization process based on a maximum likelihood estimate of the Laplacian distribution.

[0236] At operation 1708, the technique 1700 can include determining a gradient of the estimated bitrate using the encoder agent, as Figure 8 As shown. In some cases, the video encoder can include an entropy coding process after the quantization process, and further include estimating an output of the entropy coding process based on a spacing around coefficients input to the quantization process of the video encoder. In some cases, the estimated bitrate is based on the estimated output of the quantization process and the estimated output of the quantization process. In some cases, the gradient is determined based on at least a derivative of the statistical model.

[0237] At operation 1710, the technique can include training the encoder agent to predict the estimated bitrate based on the actual bitrate, the estimated bitrate, and the gradient.

[0238] Figure 18 is a flow diagram illustrating a technique for performing bitrate estimation 1800, in accordance with aspects of the present disclosure. At operation 1802, the technique 1800 can include receiving one or more frames of video data for encoding by a video encoder, the video encoder including at least a quantization process, as Figures 6A-6B and Figure 7 As shown.

[0239] At operation 1804, the technique 1800 can include predicting an estimated bitrate of the one or more frames after encoding by the video encoder using an encoder agent, wherein the encoder agent includes a statistical model for estimating the estimated bitrate, and wherein the statistical model is trained based on a gradient of the estimated bitrate, as Figures 6B-8 As shown. In some cases, the encoder agent estimates an output of one or more processes of the video encoder. In some cases, the statistical model is based on a Laplacian distribution of coefficients input to a quantization process of the video encoder. In some cases, the technique 1800 can further include estimating an output of the quantization process based on a maximum likelihood estimate of the Laplacian distribution. In some cases, the video encoder can include an entropy coding process after the quantization process, and further include estimating an output of the entropy coding process based on a spacing around coefficients input to the quantization process of the video encoder. In some cases, the estimated bitrate is based on the estimated output of the quantization process and the estimated output of the quantization process. In some cases, the gradient is determined based on at least a derivative of the statistical model.

[0240] At operation 1806, the technique 1800 can adjust one or more quality parameters based on the predicted estimated bitrate, as shown in Figure 6A At operation 1808, the technique 1800 can encode one or more frames of the video data using a video encoder that includes at least a quantization process, as shown in Figure 6A and Figure 7 .

[0241] Figure 19 An example computing device architecture 1900 of an example computing device that can implement various techniques described herein is shown. In some examples, the computing device can include a mobile device, a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a video server, a vehicle (or a computing device of a vehicle), or other device. For example, the computing device architecture 1900 can be used as part of the system 200 of Figure 2 and / or the system 300 of Figure 3 . Components of the computing device architecture 1900 are shown in electrical communication with each other using connection 1905 (e.g., a bus). Example computing device architecture 1900 includes a processing unit (CPU or processor) 1910 and computing device connections 1905 that couple various computing device components (including computing device memory 1915, such as read only memory (ROM) 1920 and random access memory (RAM) 1925) to the processor 1910.

[0242] The computing device architecture 1900 can include a cache of high-speed memory directly coupled to, in close proximity to, or integrated as part of the processor 1910. The computing device architecture 1900 can copy data from memory 1915 and / or storage device 1930 to cache 1912 for quicker access by processor 1910. In this way, the cache can provide a performance boost that avoids processor 1910 delays while waiting for data. These and other modules can control or be configured to control the processor 1910 to perform various actions. Other computing device memory 1915 can be available for use as well. The memory 1915 can include multiple different types of memory with different performance characteristics. The processor 1910 can include any general purpose processor and a hardware or software service (e.g., service 1 1932, service 2 1934, and service 3 1936) stored in storage device 1930 and configured to control processor 1910 as well as specialized processors in which software instructions are incorporated into the processor design. The processor 1910 can be a self-contained computing system, with multiple cores or processors, a bus, memory controller, and cache, among other components. Multi-core processors can be symmetric or asymmetric.

[0243] To enable user interaction with the computing device architecture 1900, an input device 1945 can represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, keyboard, mouse, motion input, speech and the like. An output device 1935 can also be one or more of a number of output mechanisms known to those of skill in the art, such as a display, projector, television, speaker device, etc. In some instances, multi-modal computing devices can enable a user to provide multiple types of input to communicate with the computing device architecture 1900. The communications interface 1940 can generally govern and manage the user input and computing device outputs. There is no restriction on operating on any particular hardware arrangement and therefore the basic features here can easily be substituted for improved hardware or firmware arrangements as they are developed.

[0244] Storage device 1930 is a non-transitory memory and can be a hard disk or other types of computer readable media which can store data that are accessible by a computer, such as magnetic cassettes, flash memory cards, solid state memory devices, digital versatile disks, cartridges, random access memories (RAMs) 1925, read only memory (ROM) 1920, and hybrids thereof. The storage device 1930 can include services 1932, 1934, 1936 that, in operation, control the processor 1910. Other hardware or software modules are contemplated. The storage device 1930 can be connected to the computing device connection 1905. In one aspect, a hardware module that performs a particular function can include the software component stored in a computer-readable medium that, in operation, is connected with the necessary hardware components, such as the processor 1910, connection 1905, output device 1935, etc., to carry out that function.

[0245] Aspects of the present disclosure are applicable to any suitable electronic device (e.g., a security system, a smartphone, a tablet, a laptop, a vehicle, a drone, or other device) that includes or is coupled to one or more active depth sensing systems. Although described below with respect to a device having or coupled to one light projector, aspects of the present disclosure are applicable to devices having any number of light projectors and are therefore not limited to a particular device.

[0246] The term "device" is not limited to one or a particular number of physical objects (e.g., a smartphone, a controller, a processing system, etc.). As used herein, a device can be any electronic device having one or more portions that can implement at least some portions of the present disclosure. While the description and examples below use the term "device" to describe various aspects of the present disclosure, the term "device" is not limited to a particular configuration, type, or number of objects. Furthermore, the term "system" is not limited to multiple components or a particular embodiment. For example, a system can be implemented on one or more printed circuit boards or other substrates and can have movable or static components. While the description and examples below use the term "system" to describe various aspects of the present disclosure, the term "system" is not limited to a particular configuration, type, or number of objects.

[0247] In the description above, specific details are set forth in order to provide a thorough understanding of embodiments and examples provided herein. However, persons having ordinary skill in the relevant arts will appreciate that embodiments can be practiced without the specific details. In some instances, well-known structures have not been described in detail in order to avoid obscuring embodiments. Circuits, systems, networks, processes, and other components might be shown as components in block diagram form in order to not obscure embodiments. In other instances, well-known circuits, processes, algorithms, structures, techniques, and components might not be described in detail in order to avoid obscuring embodiments.

[0248] Various embodiments can be described, in the foregoing disclosure, as a process or method being performed in a flow direction. Although the process or method can be described in a sequential manner, some operations can in fact be performed in parallel or in a different order. Additionally, some operations can be omitted, combined, or substituted for other operations. The process or method can be terminated when its operations are completed but can also generate additional steps not included in the figure. A process can correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination can correspond to a return of the function to the calling function or the main function.

[0249] Processes and methods according to the examples described above can be implemented using computer-executable instructions, which are stored or otherwise available from computer-readable media. Such instructions can include, for example, instructions and data which cause or otherwise configure a general purpose computer, special purpose computer, or a processing device to perform a certain function or group of functions. Portions of computer resources used can be accessible via a network. The computer-executable instructions can be, for example, binaries, intermediate format instructions such as assembly language, firmware, source code, or any other

[0250] The term "computer-readable medium" includes, but is not limited to portable or fixed storage devices, optical storage devices, and various other mediums capable of storing, containing or carrying instruction and / or data. The computer-readable medium can include a non-transitory medium in which data can be stored and not a carrier wave or signals per se, although a carrier wave or signals can be used to propagate or carry the data. Examples of a non-transitory medium can include, but are not limited to, a magnetic disk, a magnetic tape, an optical storage medium, a memory or memory device, a compact disc (CD) or digital versatile disc (DVD), a flash memory, a punch card, a paper tape, an optical mark sheet, an

[0251] In some embodiments, computer-readable storage devices, media and memory can include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media expressly excludes media such as energy, carrier signals, electromagnetic waves, and signals per se.

[0252] Devices implementing processes and methods according to these disclosures can include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and can take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments to perform the necessary tasks (e.g., a computer-program product) can be stored in a computer-readable or machine-readable medium. A processor(s) can perform the necessary tasks. Typical examples of form factors include laptops, smart phones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rackmount devices, standalone devices, and so on. Functionality described herein also can be embodied in peripherals or add-on

[0253] Instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functionality described in this disclosure.

[0254] In the foregoing description, various aspects of the present application have been described in reference to particular embodiments that have been described in considerable detail. It will be apparent to those skilled in the art that many variations can be made to the embodiments described without departing from the spirit and scope of the application. Accordingly, it is intended that the application be limited only by the spirit and scope of the appended claims, including the equivalents thereof.

[0255] Ordinary skill will appreciate that the less than (“<”) and greater than (“>”) symbols or terms used herein can be replaced with less than or equal to (“≤”) and greater than or equal to (“≥”) symbols without departing from the scope of this specification.

[0256] Where components are described as being “configured to” perform certain operations, such configuration can be accomplished, for example, by designing the electronic circuitry or other hardware of the component to perform the operation, by programming the component to perform the operation, or any combination thereof. For example, a component can be configured to perform an operation by being designed to perform the operation, by being programmed to perform the operation, or any combination thereof.

[0257] The phrase "coupled to" means any direct or indirect physical connection between elements, and / or any connection between elements that is capable of communicating signals between them (e.g., a wired or wireless connection, and / or other suitable connections that are known in the art).

[0258] Claim language referring to "at least one of' and / or "one or more of a collection of items refers to one member of the collection or more than one member of the collection, in any combination. For example, claim language referring to "at least one of A and B" or "at least one of A or B" means A or B or both A and B. In another example, claim language referring to "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, A and B, A and C, B and C, or A and B and C. The language "at least one of' and / or "one or more of a collection of items does not restrict the collection to consisting only of members listed. For example, claim language stating "at least one of A and B" or "at least one of A or B" can be satisfied by A alone, B alone, or A and B.

[0259] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0260] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of various devices such as a general purpose computer, a wireless communication device handset, or an integrated circuit device having a multitude of uses including use in wireless communication device handsets and other devices. Any features described as modules or components can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques can be realized at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, performs one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which can include packaging material. The computer-readable medium can comprise memory or data storage media, such as random access memory (RAM) such as synchronous dynamic random access memory (SDRAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), FLASH memory, magnetic or optical data storage media, and the like. Additionally or in the alternative, the techniques described can be realized at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.

[0261] The program code can be executed by a processor, which can include one or more of a set of processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such a processor can be configured to perform any of the techniques described in this disclosure. A general purpose processor can be a microprocessor; microcontrollers; an ASIC; a FPGA; or any other equivalent integrated or discrete logic circuitry. In this description, the term "processor" is used to generally refer to any of the foregoing logic circuitry, alone or in combination with other logic circuitry. Other elements of a practical implementation can include one or more memory / storage devices which can be implemented as any combination of volatile and non-volatile storage, such as RAM (random access memory, such as DRAM (dynamic

[0262] Illustrative aspects of the disclosure include:

[0263] Aspect 1 : A method of processing video data. The method comprises encoding one or more frames of video data using a video encoder, the video encoder comprising at least a quantization process; determining an actual bitrate of the encoded one or more frames; predicting an estimated bitrate using an encoder agent, the encoder agent comprising a statistical model for estimating a bitrate of the encoded one or more frames; determining a gradient of the estimated bitrate using the encoder agent; and training the encoder agent to predict the estimated bitrate based on the actual bitrate, the estimated bitrate, and the gradient.

[0264] Aspect 2. The method of claim 1, wherein the encoder agent estimates an output of one or more processes of the video encoder.

[0265] Aspect 3. The method of any of claims 1-2, wherein the statistical model is based on a Laplace distribution of coefficients input to the quantization process of the video encoder.

[0266] Aspect 4. The method of claim 3, further comprising estimating an output of the quantization process based on a maximum likelihood estimate of the Laplace distribution.

[0267] Aspect 5. The method of any of claims 1-4, wherein the video encoder further comprises an entropy coding process after the quantization process, and further comprising estimating an output of the entropy coding process based on a spacing around coefficients input to the quantization process of the video encoder.

[0268] Aspect 6. The method of claim 5, wherein the estimated bitrate is based on the estimated output of the quantization process and the estimated output of the quantization process.

[0269] Aspect 7. The method of claim 1, wherein the gradient is determined based on at least a derivative of the statistical model.

[0270] Aspect 8. An apparatus for processing video data, the apparatus comprising at least one memory; and at least one processor coupled to the at least one memory, wherein the at least one processor is configured to: encode one or more frames of video data using a video encoder, the video encoder comprising at least a quantization process; determine an actual bitrate of the encoded one or more frames; predict an estimated bitrate using an encoder agent, the encoder agent comprising a statistical model for estimating a bitrate of the encoded one or more frames; determine a gradient of the estimated bitrate using the encoder agent; and train the encoder agent to predict the estimated bitrate based on the actual bitrate, the estimated bitrate, and the gradient.

[0271] Aspect 9. The apparatus of claim 8, wherein the encoder proxy estimates an output of one or more processes of the video encoder.

[0272] Aspect 10. The apparatus of any of claims 8-9, wherein a statistical model is based on a Laplacian distribution of coefficients input to the quantization process of the video encoder.

[0273] Aspect 11. The apparatus of claim 10, wherein the at least one processor is further configured to estimate an output of the quantization process based on a maximum likelihood estimate of the Laplacian distribution.

[0274] Aspect 12. The apparatus of any of claims 8-11, wherein the video encoder further comprises an entropy coding process after the quantization process, and wherein the processor is further configured to estimate an output of the entropy coding process based on a spacing around coefficients input to the quantization process of the video encoder.

[0275] Aspect 13. The apparatus of claim 12, wherein the estimated bitrate is based on the estimated output of the quantization process and the estimated output of the quantization process.

[0276] Aspect 14. The apparatus of claim 8, wherein the gradient is determined based on at least a derivative of the statistical model.

[0277] Aspect 15. A method for processing video data, the method comprising: receiving one or more frames of video data for encoding by a video encoder, the video encoder comprising at least a quantization process; predicting an estimated bitrate of the one or more frames after encoding by the video encoder using an encoder proxy, wherein the encoder proxy comprises a statistical model for estimating the estimated bitrate, and wherein the statistical model is trained based on a gradient of the estimated bitrate; adjusting one or more quality parameters based on the predicted estimated bitrate; and encoding the one or more frames of video data using the video encoder.

[0278] Aspect 16. The method of claim 15, wherein the encoder proxy estimates an output of one or more processes of the video encoder.

[0279] Aspect 17. The method of any of claims 15-16, wherein a statistical model is based on a Laplacian distribution of coefficients input to the quantization process of the video encoder.

[0280] Aspect 18. The method of claim 17, further comprising estimating an output of the quantization process based on a maximum likelihood estimate of the Laplacian distribution.

[0281] Aspect 19. The method of any of claims 15-18, wherein the video encoder further comprises an entropy coding process after the quantization process, and wherein the method further comprises estimating an output of the entropy coding process based on a spacing around coefficients input to the quantization process of the video encoder.

[0282] Aspect 20. The method of claim 19, wherein the estimated bitrate is based on the estimated output of the quantization process and the estimated output of the quantization process.

[0283] Aspect 21. The method of claim 15, wherein the gradient is determined based on at least a derivative of the statistical model.

[0284] Aspect 22. An apparatus for processing video data, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor configured to: receive one or more frames of video data for encoding by a video encoder, the video encoder comprising at least a quantization process; predict an estimated bitrate of the one or more frames after encoding by the video encoder using an encoder proxy, wherein the encoder proxy comprises a statistical model for estimating the estimated bitrate, and wherein the statistical model is trained based on a gradient of the estimated bitrate; adjust one or more quality parameters based on the predicted estimated bitrate; and encode the one or more frames of video data using the video encoder.

[0285] Aspect 23. The apparatus of claim 22, wherein the encoder proxy estimates an output of one or more processes of the video encoder.

[0286] Aspect 24. The apparatus of any of claims 22-23, wherein the statistical model is based on a Laplacian distribution of coefficients input to the quantization process of the video encoder.

[0287] Aspect 25. The apparatus of claim 24, wherein the at least one processor is further configured to estimate an output of the quantization process based on a maximum likelihood estimate of the Laplacian distribution.

[0288] Aspect 26, the apparatus of claims 22-25, wherein the video encoder further comprises an entropy coding process after the quantization process, and wherein the processor is further configured to estimate an output of the entropy coding process based on a spacing around coefficients input to the quantization process of the video encoder.

[0289] Aspect 27, the apparatus of claim 26, wherein the estimated bitrate is based on an estimated output of the quantization process and an estimated output of the quantization process.

[0290] Aspect 28, the apparatus of claim 22, wherein the gradient is determined based on at least a derivative of the statistical model.

[0291] Aspect 29, a method for processing video data, the method comprising: receiving one or more frames of video data for encoding by a video encoder; predicting an estimated bitrate of the one or more frames after encoding by the video encoder using an encoder proxy, wherein the encoder proxy comprises a statistical model for estimating the estimated bitrate; determining a gradient of the estimated bitrate using the encoder proxy; encoding the one or more frames of video data using the video encoder, the video encoder comprising at least a quantization process; obtaining an actual bitrate of the encoded one or more frames; and updating the encoder proxy based on a comparison between the estimated bitrate and actual bitrate.

[0292] Aspect 30, the method of claim 29, wherein the encoder proxy estimates an output of one or more processes of the video encoder.

[0293] Aspect 31, the method of any of claims 29-30, wherein the statistical model is based on a Laplacian distribution of coefficients input to the quantization process of the video encoder.

[0294] Aspect 32, the method of claim 31, further comprising estimating an output of the quantization process based on a maximum likelihood estimate of the Laplacian distribution.

[0295] Aspect 33, the method of any of claims 29-32, wherein the video encoder further comprises an entropy coding process after the quantization process, and wherein the method further comprises estimating an output of the entropy coding process based on a spacing around coefficients input to the quantization process of the video encoder.

[0296] Aspect 34, the method of claim 33, wherein the estimated bitrate is based on an estimated output of the quantization process and an estimated output of the quantization process.

[0297] Aspect 35. The method of claim 29, wherein the gradient is determined based at least on a derivative of the statistical model.

[0298] Aspect 36. An apparatus for processing video data, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor configured to: receive one or more frames of video data for encoding by a video encoder; predict an estimated bitrate of the one or more frames after encoding by the video encoder using an encoder agent, wherein the encoder agent comprises a statistical model for estimating the estimated bitrate; determine a gradient of the estimated bitrate using the encoder agent; encode the one or more frames of video data using the video encoder, the video encoder comprising at least a quantization process; obtain an actual bitrate of the encoded one or more frames; and update the encoder agent based on a comparison between the estimated and actual bitrates.

[0299] Aspect 37. The apparatus of claim 36, wherein the encoder agent estimates an output of one or more processes of the video encoder.

[0300] Aspect 38. The apparatus of any of claims 36-37, wherein the statistical model is based on a Laplacian distribution of coefficients input to the quantization process of the video encoder.

[0301] Aspect 39. The apparatus of claim 38, wherein the at least one processor is further configured to estimate an output of the quantization process based on a maximum likelihood estimate of the Laplacian distribution.

[0302] Aspect 40. The apparatus of any of claims 36-39, wherein the video encoder further comprises an entropy coding process after the quantization process, and wherein the at least one processor is further configured to estimate an output of the entropy coding process based on a spacing around coefficients input to the quantization process of the video encoder.

[0303] Aspect 41. The apparatus of claim 39, wherein the estimated bitrate is based on an estimated output of the quantization process and an estimated output of the quantization process.

[0304] Aspect 42. The apparatus of claim 36, wherein the gradient is determined based at least on a derivative of the statistical model.

[0305] Aspect 43. The apparatus of any of aspects 8-14, aspects 22-28, and aspects 37-42, wherein the apparatus comprises an encoder.

[0306] Aspect 44. The apparatus of any of aspects 8-14, 22-28, and 37-43, further comprising a display configured to display one or more output pictures.

[0307] Aspect 45. The apparatus of any of aspects 8-14, 22-28, and 37-44, further comprising a camera configured to capture one or more pictures.

[0308] Aspect 46. The apparatus of any of aspects 8-14, 22-28, and 37-45, wherein the apparatus is a mobile device.

[0309] Aspect 47. An apparatus for processing video data, comprising means for performing one or more of the operations in accordance with any of aspects 8-14, 22-28, and 37-46.

[0310] Aspect 48. A non-transitory computer-readable storage medium having instructions stored thereon, the instructions, when executed by one or more processors, cause the one or more processors to perform operations in accordance with any one or more of aspects 1-7, 15-21, and / or 29-35.

[0311] Aspect 49. An apparatus comprising one or more elements for performing operations in accordance with any one or more of aspects 1-7, 15-21, and / or 29-35.

Claims

1. A method for processing video data, the method comprising: encoding one or more frames of video data using a video encoder, the video encoder comprising at least a quantization process; determining an actual bitrate of the encoded one or more frames; predicting an estimated bitrate using an encoder agent, the encoder agent comprising a statistical model for estimating a bitrate of the encoded one or more frames; determining a gradient of the estimated bitrate using the encoder agent; and training the encoder agent to predict the estimated bitrate based on the actual bitrate, the estimated bitrate, and the gradient. The encoder agent estimates an output of one or more processes of the video encoder.

2. The method of claim 1, wherein, The statistical model is based on a Laplacian distribution of coefficients input to the quantization process of the video encoder.

3. The method of claim 1, wherein, 4. The method of claim 3, further comprising estimating an output of the quantization process based on a maximum likelihood estimate of the Laplacian distribution. The video encoder further comprises an entropy coding process after the quantization process, and further comprising estimating an output of the entropy coding process based on a spacing around coefficients input to the quantization process of the video encoder.

5. The method of claim 1, wherein, The estimated bitrate is based on an estimated output of the quantization process.

6. The method of claim 5, wherein, The gradient is determined based on at least a derivative of the statistical model.

7. The method of claim 1, wherein, 8. An apparatus for processing video data, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory, wherein the at least one processor is configured to: encode one or more frames of video data using a video encoder, the video encoder comprising at least a quantization process; determine an actual bitrate of the encoded one or more frames; predict an estimated bitrate using an encoder agent, the encoder agent comprising a statistical model for estimating a bitrate of the encoded one or more frames; determine a gradient of the estimated bitrate using the encoder agent; and train the encoder agent to predict the estimated bitrate based on the actual bitrate, the estimated bitrate, and the gradient. The encoder agent estimates an output of one or more processes of the video encoder.

9. The apparatus of claim 8, wherein, The statistical model is based on a Laplacian distribution of coefficients input to the quantization process of the video encoder.

10. The apparatus of claim 8, wherein, The at least one processor is further configured to estimate an output of the quantization process based on a maximum likelihood estimate of the Laplacian distribution.

11. The apparatus of claim 10, wherein, The video encoder further comprises an entropy coding process after the quantization process, and wherein the at least one processor is further configured to estimate an output of the entropy coding process based on a spacing around coefficients input to the quantization process of the video encoder.

12. The apparatus of claim 8, wherein, The estimated bitrate is based on an estimated output of the quantization process.

13. The apparatus of claim 12, wherein, The gradient is determined based on at least a derivative of the statistical model.

14. The apparatus of claim 8, wherein, 15. A method for processing video data, the method comprising: receiving one or more frames of video data for encoding by a video encoder, the video encoder comprising at least a quantization process; ​ predict an estimated bitrate of the one or more frames after encoding by the video encoder using an encoder agent, wherein the encoder agent estimates the estimated bitrate using a statistical model, and wherein the statistical model is trained based on a gradient of the estimated bitrate; adjust one or more quality parameters of the video encoder based on the predicted estimated bitrate; and encode the one or more frames of video data using the video encoder.

16. The method of claim 15, wherein, the encoder agent estimates an output of one or more processes of the video encoder.

17. The method of claim 15, wherein, the statistical model is based on a Laplacian distribution of coefficients input to the quantization process of the video encoder.

18. The method of claim 17, further comprising estimating an output of the quantization process based on a maximum likelihood estimate of the Laplacian distribution.

19. The method of claim 15, wherein, the video encoder further comprises an entropy coding process after the quantization process, and wherein the method further comprises estimating an output of the entropy coding process based on a spacing around coefficients input to the quantization process of the video encoder.

20. The method of claim 19, wherein, the estimated bitrate is based on an estimated output of the quantization process.

21. The method of claim 15, wherein, the gradient is determined based on at least a derivative of the statistical model.

22. An apparatus for processing video data, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory, wherein the at least one processor is configured to: receive one or more frames of video data for encoding by a video encoder, the video encoder comprising at least a quantization process; predict an estimated bitrate of the one or more frames after encoding by the video encoder using an encoder agent, wherein the encoder agent estimates the estimated bitrate using a statistical model, and wherein the statistical model is trained based on a gradient of the estimated bitrate; adjust one or more quality parameters of the video encoder based on the predicted estimated bitrate; and encode the one or more frames of video data using the video encoder.

23. The apparatus of claim 22, wherein, the encoder agent estimates an output of one or more processes of the video encoder.

24. The apparatus of claim 22, wherein, the statistical model is based on a Laplacian distribution of coefficients input to the quantization process of the video encoder.

25. The apparatus of claim 24, wherein, the at least one processor is further configured to estimate an output of the quantization process based on a maximum likelihood estimate of the Laplacian distribution.

26. The apparatus of claim 22, wherein, the video encoder further comprises an entropy coding process after the quantization process, and wherein the at least one processor is configured to estimate an output of the entropy coding process based on a spacing around coefficients input to the quantization process of the video encoder.

27. The apparatus of claim 26, wherein, the estimated bitrate is based on an estimated output of the quantization process.

28. The apparatus of claim 22, wherein, the gradient is determined based on at least a derivative of the statistical model.

29. The apparatus of claim 22, wherein, the apparatus is the video encoder.

30. The apparatus of claim 22, wherein, the apparatus comprises the video encoder.

Citation Information

Patent Citations

  • Method and Apparatus for Video Codec Quantization

    US20080080615A1

  • Video Characterization For Smart Encoding Based On Perceptual Quality Optimization

    US20190289296A1