Joint Termination of Bidirectional Data Blocks for Parallel Coding

By employing joint termination of bidirectional data blocks in parallel entropy coding, the systems and techniques address the inefficiencies in existing video coding, achieving improved compression and reduced overhead for high-quality video data processing.

JP7815230B2Active Publication Date: 2026-02-17QUALCOMM INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023521759
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-04-07
Filing Date
2021-09-20
Publication Date
2026-02-17
Estimated Expiration
2041-09-20

AI Technical Summary

Technical Problem

The increasing demand for high-quality video data strains communication networks and devices due to the large amount of data required, and existing video coding techniques face challenges in achieving efficient compression without degrading video quality, particularly with sequential arithmetic coding leading to throughput bottlenecks.

Method used

The implementation of systems and techniques for joint termination of bidirectional data blocks using parallel entropy coding, which optimizes the termination of paired arithmetic-coded bitstreams through bidirectional data packing, reducing overhead and compression loss.

Benefits of technology

This approach significantly enhances coding efficiency by minimizing overhead and maintaining video quality, overcoming throughput bottlenecks in hardware limitations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007815230000026
    Figure 0007815230000026
  • Figure 0007815230000027
    Figure 0007815230000027
  • Figure 0007815230000028
    Figure 0007815230000028
Patent Text Reader

Abstract

Techniques for processing video data are described herein. For example, a process may include obtaining encoded video data. The process may include determining a value intersection between a value for a first termination byte of a first parcel of the encoded video data and a value for a second termination byte of a second parcel of the encoded video data. The process may further include determining a co-termination byte for the first termination byte of the first parcel and the second termination byte of the second parcel. The value for the co-termination byte is based on the value intersection. The process may include generating entropy-coded data including the co-termination byte for the first parcel and the second parcel. The entropy-coded data may be generated using arithmetic coding or binary coding.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to systems and techniques for coding (eg, encoding and / or decoding) image and / or video content. [Background technology]

[0002]

[0001] Many devices and systems enable video data to be processed and output for consumption. Digital video data comprises a large amount of data to meet the demands of consumers and video providers. For example, video data consumers desire high-quality video, including high fidelity, resolution, frame rates, etc. As a result, the large amount of video data required to meet these demands places a strain on communication networks and devices that process and store the video data.

[0003]

[0002] Video coding techniques can be used to compress video data. The goal of video coding is to compress video data into a format that uses a lower bit rate while avoiding or minimizing degradation to video quality. As ever-evolving video services become available, encoding techniques with better coding efficiency are needed. Summary of the Invention

[0004] Systems and techniques for coding (e.g., encoding and / or decoding) image and / or video content are described. In one illustrative example, a method for processing video data is provided. The method includes obtaining encoded video data, determining a value intersection between a value for a first termination byte of a first parcel of the encoded video data and a value for a second termination byte of a second parcel of the encoded video data, determining a joint termination byte for the first termination byte of the first parcel and the second termination byte of the second parcel, and generating entropy-coded data including the joint termination byte for the first parcel and the second parcel, wherein the value for the joint termination byte is based on the value intersection (the intersection value).

[0005] In another example, an apparatus for processing video data is provided, the apparatus including: a memory configured to store video data; and a processor (e.g., implemented in a circuit) coupled to the memory. In some examples, two or more processors may be coupled to the memory and used to perform one or more of the operations. The one or more processors are configured to: obtain encoded video data; determine a value intersection between a value for a first termination byte of a first parcel of the encoded video data and a value for a second termination byte of a second parcel of the encoded video data; determine a co-termination byte for the first termination byte of the first parcel and the second termination byte of the second parcel; and generate entropy-coded data including a co-termination byte for the first parcel and the second parcel, wherein the value for the co-termination byte is based on the value intersection (the value of the intersection).

[0006]

[0005] In another example, a non-transitory computer-readable medium is provided having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to obtain encoded video data; determine a value intersection between a value for a first termination byte of a first parcel of the encoded video data and a value for a second termination byte of a second parcel of the encoded video data; determine a co-termination byte for the first termination byte of the first parcel and the second termination byte of the second parcel; and generate entropy coded data including a co-termination byte for the first parcel and the second parcel, wherein the value for the co-termination byte is based on the value intersection (the intersection value).

[0007] In another example, an apparatus for processing video data is provided, the apparatus including: means for obtaining encoded video data; means for determining a value intersection between a value for a first termination byte of a first parcel of the encoded video data and a value for a second termination byte of a second parcel of the encoded video data; means for determining a co-termination byte for the first termination byte of the first parcel and the second termination byte of the second parcel; and means for generating entropy coded data including a co-termination byte for the first parcel and the second parcel, wherein the value for the co-termination byte is based on the value intersection (the value of the intersection).

[0008] In some aspects, the entropy coded data is generated using arithmetic coding.

[0009]

[0008] In some aspects, the value for the first termination byte includes a first range of termination byte values ​​that are allowed to be decoded, the value for the second termination byte includes a second range of termination byte values ​​that are allowed to be decoded, and the intersection of the values ​​(the intersection value) includes a value that is in the first range and the second range.

[0010] In some aspects, the entropy coded data is generated using binary coding.

[0011] In some aspects, the value for the first termination byte includes a first number of bits, the value for the second termination byte includes a second number of bits, and an intersection of the values ​​(intersection value) includes at least one of a common value in the first number of bits and the second number of bits, a subset of values ​​from the first number of bits, and a subset of values ​​from the second number of bits. In some cases, the order of the first number of bits and the order of the second number of bits are unchanged in the co-termination byte compared to the order of the first number of bits in the first termination byte and the order of the second number of bits in the second termination byte.

[0012] In some aspects, generating the entropy coded data includes performing parallel entropy coding of the first parcel and the second parcel.

[0013] In some aspects, the first parcel is encoded using a first encoder and the second parcel is encoded using a second encoder.

[0014]

[0013] In some aspects, the methods, apparatus, and computer-readable media described above comprise storing a first parcel in a first buffer and storing a second parcel in a second buffer.

[0015] In some aspects, the above-described methods, apparatus, and computer-readable media comprise transmitting a bitstream including entropy coded data.

[0016] In some aspects, the above-described methods, apparatus, and computer-readable media comprise storing a bitstream including entropy coded data.

[0017]

[0016] In some aspects, the methods, apparatus, and computer-readable media described above comprise performing parallel entropy decoding of the first parcel and the second parcel using a joint termination byte for the first parcel and the second parcel.

[0018]

[0017] In some aspects, the methods, apparatus, and computer-readable media described above comprise reading a first parcel in a forward order and reading a second parcel in a reverse order.

[0019] In some aspects, the above-described methods, apparatus, and computer-readable media comprise converting the bytes of the second parcel into reverse order.

[0020] In some aspects, the joint termination byte is the final termination byte of the first parcel and the second parcel for processing.

[0021] In some aspects, the encoded video data comprises one or more syntax elements of a video bitstream. In some aspects, the one or more syntax elements indicate one or more parameters defining a neural network for decoding the encoded video data. In some aspects, the one or more parameters defining the neural network comprise at least one of neural network weights and an activation function of the neural network.

[0022] In another illustrative example, a method for processing video data is provided, the method including obtaining a first parcel of entropy coded data and a second parcel of entropy coded data, and performing parallel entropy decoding of the first parcel and the second parcel using a co-termination byte for the first parcel and the second parcel, the first parcel and the second parcel sharing a co-termination byte, where a value for the co-termination byte is based on an intersection of values ​​between a value for the first termination byte of the first parcel and a value of a second termination byte of the second parcel.

[0023] In another example, an apparatus for processing video data is provided, the apparatus including: a memory configured to store video data; and a processor (e.g., implemented in a circuit) coupled to the memory. In some examples, two or more processors may be coupled to the memory and used to perform one or more of the operations. The one or more processors are configured to: obtain a first parcel of entropy coded data and a second parcel of entropy coded data; and perform parallel entropy decoding of the first parcel and the second parcel using a common termination byte for the first parcel and the second parcel, the first parcel and the second parcel sharing a common termination byte, where a value for the common termination byte is based on an intersection of values ​​between a value for the first termination byte of the first parcel and a value for the second termination byte of the second parcel.

[0024]

[0023] In another example, a non-transitory computer-readable medium is provided that stores instructions that, when executed by one or more processors, cause the one or more processors to obtain a first parcel of entropy coded data and a second parcel of entropy coded data, and perform parallel entropy decoding of the first parcel and the second parcel using a common termination byte for the first parcel and the second parcel, where the first parcel and the second parcel share a common termination byte, and where the value for the common termination byte is based on an intersection of values ​​between the value for the first termination byte of the first parcel and the value of the second termination byte of the second parcel.

[0025] In another example, an apparatus for processing video data is provided, the apparatus including: means for obtaining a first parcel of entropy coded data and a second parcel of entropy coded data; and means for performing parallel entropy decoding of the first parcel and the second parcel using a common termination byte for the first parcel and the second parcel, the first parcel and the second parcel sharing a common termination byte, where a value for the common termination byte is based on an intersection of values ​​between a value for the first termination byte of the first parcel and a value of the second termination byte of the second parcel.

[0026] In some aspects, the entropy coded data is encoded using arithmetic coding.

[0027]

[0026] In some aspects, the value for the first termination byte includes a first range of termination byte values ​​that are allowed to be decoded, the value for the second termination byte includes a second range of termination byte values ​​that are allowed to be decoded, and the intersection of the values ​​(the intersection value) includes a value that is in the first range and the second range.

[0028] In some aspects, the entropy coded data is generated using binary coding.

[0029]

[0028] In some aspects, the value for the first termination byte includes a first number of bits, the value for the second termination byte includes a second number of bits, and the intersection of the values ​​(intersection value) includes at least one of a common value between the first number of bits and the second number of bits, and a subset of values ​​from the first number of bits and a subset of values ​​from the second number of bits.

[0030]

[0029] In some aspects, the order of the first number of bits and the order of the second number of bits are not changed in the co-termination byte compared to the order of the first number of bits in the first termination byte and the order of the second number of bits in the second termination byte.

[0031]

[0030] In some aspects, the methods, apparatus, and computer-readable media described above comprise obtaining a first parcel from a first buffer and obtaining a second parcel from a second buffer.

[0032]

[0031] In some aspects, the methods, apparatus, and computer-readable media described above comprise reading a first parcel in a forward order and reading a second parcel in a reverse order.

[0033] In some aspects, the methods, apparatus, and computer-readable media described above comprise converting the bytes of the second parcel into reverse order.

[0034] In some aspects, the joint termination byte is the final termination byte of the first parcel and the second parcel for processing.

[0035] In some aspects, the above-described methods, apparatus, and computer-readable media comprise receiving a video bitstream, the video bitstream including a first parcel, a second parcel, and one or more syntax elements. In some aspects, the one or more syntax elements indicate one or more parameters defining a neural network for decoding the encoded video data. In some aspects, the one or more parameters defining the neural network comprise at least one of neural network weights and an activation function of the neural network.

[0036] In some aspects, the apparatus comprises a mobile device (e.g., a mobile phone or so-called “smartphone,” a tablet computer, or other type of mobile device), a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a video server, a television (e.g., a network-connected television), a vehicle (or a vehicle's computing device), or other device. In some aspects, the apparatus includes at least one camera for capturing one or more images or video frames. For example, the apparatus may include a camera (e.g., an RGB camera) or multiple cameras for capturing one or more images and / or one or more videos including the video frames. In some aspects, the apparatus includes a display for displaying one or more images, videos, notifications, or other displayable data. In some aspects, the apparatus includes a transmitter configured to transmit one or more video frames and / or syntax data to the at least one device over a transmission medium. In some aspects, the processor includes a neural processing unit (NPU), a central processing unit (CPU), a graphics processing unit (GPU), or other processing device or component.

[0037] This Summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used independently to determine the scope of the claimed subject matter, which subject matter should be understood by reference to the entire specification of this patent, any or all drawings, and appropriate portions of each claim.

[0038]

[0037] The above, together with other features and embodiments, will become more apparent with reference to the following specification, claims, and accompanying drawings.

[0039]

[0038] Exemplary embodiments of the present application are described in detail below with reference to the following figures: [Brief explanation of the drawings]

[0040] [Figure 1]

[0039] FIG. 1 illustrates an exemplary implementation of a system-on-chip (SOC), according to some examples. [Figure 2]

[0040] FIG. 1 is a block diagram illustrating an encoding device and a decoding device, according to some examples. [Figure 3]

[0041] FIG. 1 illustrates an example of a system including a device operable to perform image and / or video coding (encoding and decoding) using a neural network-based system, according to some examples. [Figure 4]

[0042] FIG. 1 illustrates an example of a data organization and coding process for bidirectional byte packing, according to some examples. [Figure 5]

[0043] FIG. 1 illustrates an example of extending bidirectional byte packing to support parallel entropy encoding and decoding, according to some examples. [Figure 6]

[0044] 10A-10C are diagrams that graphically illustrate factors used for correct arithmetic coding termination, according to some examples; [Figure 7A]

[0045] FIG. 10 illustrates an example of byte termination in bidirectional byte packing when bits are written in reverse order within bytes in a reverse stream, in accordance with some examples, where byte concatenation is used. [Figure 7B]

[0046] FIG. 10 illustrates an example of byte termination in bidirectional byte packing when bits are written in reverse order within bytes in a reverse stream, where bits are copied to a shared termination byte if there is no overlap in the bit positions used, according to some examples. [Figure 8A]

[0047] FIG. 10 illustrates an example of byte termination in bidirectional byte packing when bits are written in the same order for both streams, with byte concatenation used, in accordance with some examples. [Figure 8B]

[0048] FIG. 10 illustrates an example of byte termination in bidirectional byte packing when bits are written in the same order for both streams, where bits are copied to a shared termination byte if the first bit values ​​are equivalent, in accordance with some examples. [Figure 9A]

[0049] A diagram showing an example of valid arithmetic coding termination byte value ranges (gray area) and equal values ​​(dashed lines) with non-empty intersections (bold dashed lines), according to some examples. [Figure 9B]

[0050] FIG. 10 illustrates an example of a range of valid arithmetic coding termination byte values ​​(gray area) and equal values ​​(dashed lines) with empty intersections, according to some examples. [Figure 10A]

[0051] FIG. 10 illustrates an example of bidirectional byte packing of an arithmetic coded stream defined by sets of allowable termination bytes with no intersection between the sets, in accordance with some examples. [Figure 10B]

[0052] FIG. 10 illustrates an example of bidirectional byte packing of an arithmetic coded stream defined by a set of allowed termination bytes, with one byte in the intersection used for a shared termination, in accordance with some examples. [Figure 11A]

[0053] FIG. 10 illustrates an example of how the range interval for the termination byte may be split when there is a possibility of an arithmetic coded carry operation in the addition, according to some examples. [Figure 11B] FIG. 10 illustrates an example of how a range interval for a termination byte may be split when there is a possibility of an arithmetic coding carry operation in the addition, according to some examples. [Figure 12A]

[0054] FIG. 9B illustrates an example similar to that of FIG. 9A, but where bits are stored in reverse order in the reverse stream, according to some examples. [Figure 12B] FIG. 9C illustrates an example similar to that of FIG. 9B, but where bits are stored in reverse order in the rewind stream, according to some examples. [Figure 13]

[0055] 1 is a flowchart illustrating an example of a process for processing video data, according to some examples. [Figure 14]

[0056] 10 is a flowchart illustrating another example of a process for processing video data, according to some examples. [Figure 15A]

[0057] FIG. 1 illustrates an example of a fully connected neural network, according to some examples. [Figure 15B]

[0058] FIG. 1 illustrates an example of a locally connected neural network, according to some examples. [Figure 15C]

[0059] FIG. 1 illustrates an example of a convolutional neural network, according to some examples. [Figure 15D]

[0060] A diagram showing a detailed example of a deep convolutional network (DCN) designed to recognize visual features from images, with some examples. [Figure 16]

[0061] Block diagram showing a deep convolutional network (DCN), with some examples. [Figure 17]

[0062] FIG. 1 illustrates an example computing device architecture for an example computing device capable of implementing various techniques described herein. DETAILED DESCRIPTION OF THE INVENTION

[0041]

[0063]

[0023] Several aspects and embodiments of the present disclosure are provided below. As will be apparent to those skilled in the art, some of these aspects and embodiments may be applied independently, and some of them may be applied in combination. In the following description, for purposes of explanation, specific details are set forth to provide a thorough understanding of the embodiments of the present application. However, it will be apparent that various embodiments may be practiced without these specific details. The figures and descriptions are not limiting.

[0042]

[0064] The following description provides exemplary embodiments only and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of exemplary embodiments will provide those skilled in the art with an enabling description for implementing the exemplary embodiments. It should be understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope of the present application, as set forth in the appended claims.

[0043]

[0065] Digital video data can contain large amounts of data, especially as the demand for high-quality video data continues to grow. For example, consumers of video data generally desire increasingly higher quality video with high fidelity, resolution, frame rates, etc. However, the large amounts of video data required to meet such demand can place a significant strain on communication networks as well as devices that process and store the video data.

[0044]

[0066] Various techniques may be used to code video data. Video coding may be performed according to a particular video coding standard or may be performed using one or more machine learning systems or algorithms. Exemplary video coding standards include Generic Video Coding (VVC), High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), Moving Picture Experts Group (MPEG) coding (e.g., MPEG-5 Essential Video Coding (EVC) or other MPEG-based coding), AOMedia Video 1 (AV1), among others. Video coding often uses prediction methods such as inter-prediction or intra-prediction that exploit redundancy present in a video image or sequence. A common goal of video coding techniques is to compress video data into a format that uses a lower bitrate while avoiding or minimizing degradation of video quality. As demand for video services increases and new video services become available, coding techniques with better coding efficiency, performance, and rate control are needed.

[0045]

[0067] Video coding devices implement video compression techniques for efficiently encoding and decoding video data. Video compression techniques may include applying different prediction modes, including spatial prediction (e.g., intra-frame prediction or intra-prediction), temporal prediction (e.g., inter-frame prediction or inter-prediction), inter-layer prediction (across different layers of video data), and / or other prediction techniques, to reduce or remove redundancy inherent in video sequences. A video encoder may partition each picture of an original video sequence into rectangular regions called video blocks or coding units (described in more detail below). These video blocks may be coded using particular prediction modes.

[0046]

[0068] A video block may be divided into one or more groups of smaller blocks in one or more ways. A block may include a coding tree block, a prediction block, a transform block, and / or other suitable block. Generally, references to a "block" may refer to such a video block (e.g., a coding tree block, a coding block, a prediction block, a transform block, or other appropriate block or sub-block, as understood by those skilled in the art) unless otherwise specified. Furthermore, each of these blocks may also be referred to interchangeably herein as a "unit" (e.g., a coding tree unit (CTU), a coding unit, a prediction unit (PU), a transform unit (TU), etc.). In some cases, a unit may refer to a coding logical unit that is encoded in the bitstream, and a block may refer to a portion of a video frame buffer that a process is targeted to.

[0047]

[0069] In inter-prediction modes, a video encoder may search for a block similar to a block encoded in a frame (or picture) in another temporal location, called a reference frame or picture. The video encoder may limit its search to a certain spatial displacement from the block to be encoded. The best match may be identified using a two-dimensional (2D) motion vector that includes a horizontal displacement component and a vertical displacement component. In intra-prediction modes, the video encoder may use spatial prediction techniques to form a predicted block based on data from previously encoded neighboring blocks in the same picture.

[0048]

[0070] The video encoder may determine a prediction error. For example, the prediction may be determined as the difference between pixel values ​​in the block being coded and pixel values ​​in a predicted block. The prediction error is sometimes referred to as a residual. The video encoder may also apply a transform to the prediction error using transform coding (e.g., using a form of discrete cosine transform (DCT), a form of discrete sine transform (DST), or other suitable transform) to generate transform coefficients. After the transform, the video encoder may quantize the transform coefficients. The quantized transform coefficients and motion vectors may be represented using syntax elements and, together with control information, form a coded representation of the video sequence. In some instances, the video encoder may entropy code the syntax elements, thereby further reducing the number of bits required for their representation.

[0049]

[0071] A video decoder may use the syntax elements and control information described above to construct prediction data (e.g., a prediction block) for decoding a current frame. For example, the video decoder may add the predicted block and the compressed prediction error. The video decoder may determine the compressed prediction error by weighting the transform basis functions using the quantized coefficients. The difference between the reconstructed frame and the original frame is called the reconstruction error.

[0050]

[0072] Arithmetic coding is used by many video compression standards, including VVC, HEVC, VP9, ​​and AV1. Such widespread use is at least in part due to the fact that arithmetic coding enables powerful data modeling, resulting in compression very close to the theoretical limit. One problem with using arithmetic coding for video coding / compression is that as video resolutions and frame rates continue to increase, sequential coding can create a throughput bottleneck, which can increase costs to the point of reaching the limits of current hardware. One solution to offset such problems is to employ parallelism, for example, by dividing the bitstream into independent data blocks that can be processed simultaneously. Furthermore, new coding (encoding and decoding) methods based on machine learning techniques (e.g., using neural networks or other machine learning tools) are being developed to meet coding requirements. Such machine learning techniques can also employ arithmetic coding and / or parallelism. While dividing the bitstream into data blocks enables simultaneous processing of data, it can severely degrade compression. For example, when arithmetic coding is terminated, extra bits are required to ensure correct decoding and, in some cases, for padding to the next byte boundary.

[0051]

[0073] Described herein are systems (collectively referred to as “systems and techniques”), apparatuses, methods (also referred to as processes), and computer-readable media for performing arithmetic coding (e.g., encoding, decoding, or both encoding and decoding), as described in more detail below. The systems and techniques described herein can significantly reduce the overhead and resulting compression loss discussed above. In some cases, aspects of the systems and techniques described herein are based on the discovery that it is possible to jointly optimize the termination of a pair of arithmetic-coded bitstreams using a form of bidirectional data packing that is more efficient for split bitstreams, leading to significantly more efficient results than conventional termination. Theoretical analysis and simulations of realistic coding conditions are provided below to confirm the effectiveness of the proposed systems and techniques.

[0052]

[0074] Various aspects of the present disclosure are described with reference to the figures. Figure 1 shows an example implementation of a system-on-chip (SOC) 100 that may include a central processing unit (CPU) 102 or multi-core CPU configured to perform one or more of the functions described herein. Parameters or variables (e.g., neural signals and synaptic weights), system parameters associated with a computing device (e.g., neural network with weights), delays, frequency bin information, task information, among other information, may be stored in memory blocks associated with a neural processing unit (NPU) 108, in memory blocks associated with the CPU 102, in memory blocks associated with a graphics processing unit (GPU) 104, in memory blocks associated with a digital signal processor (DSP) 106, in memory blocks 118, and / or distributed across multiple blocks. Instructions executed in the CPU 102 may be loaded from a program memory associated with the CPU 102 or from memory blocks 118.

[0053]

[0075] SOC 100 may also include additional processing blocks adapted for specific functions, such as GPU 104, DSP 106, connectivity block 110, which may include fifth generation (5G) connectivity, fourth generation long-term evolution (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, etc., and multimedia processor 112, which may, for example, detect and recognize gestures. In one implementation, the NPU is implemented in CPU 102, DSP 106, and / or GPU 104. SOC 100 may also include sensor processor 114, image signal processor (ISP) 116, and / or navigation module 120, which may include a global positioning system.

[0054]

[0076] The SOC 100 may be based on the ARM instruction set. In one aspect of the present disclosure, instructions loaded into the CPU 102 may include code for searching a lookup table (LUT) for a stored multiplication result corresponding to a multiplication product of an input value and a filter weight. The instructions loaded into the CPU 102 may also include code for disabling a multiplier during a multiplication operation of the multiplication product when a lookup table hit for the multiplication product is detected. Furthermore, the instructions loaded into the CPU 102 may include code for storing the calculated multiplication product of the input value and the filter weight when a lookup table miss for the multiplication product is detected.

[0055]

[0077] SOC 100 and / or its components may be configured to perform video compression and / or decompression (also called video encoding and / or decoding, collectively referred to as video coding) using standards-based video coding and / or using machine learning techniques. Examples of standards-based and machine learning-based video coding systems are described with respect to FIGS. 2 and 3.

[0056]

[0078] 2 is a block diagram illustrating an example of a system 200 including an encoding device 204 and a decoding device 212, which can encode and decode video data, respectively, according to examples described herein. In some examples, the encoding device 204 and / or the decoding device 212 can include the SOC 100 of FIG. 1. The encoding device 204 can be part of a source device, and the decoding device 212 can be part of a receiving device (also referred to as a client device). In some examples, the source device can also include a decoding device similar to the decoding device 212. In some examples, the receiving device can also include an encoding device similar to the encoding device 204. The source device and / or receiving device may include an electronic device such as a mobile or landline telephone handset (e.g., a smartphone, a cellular telephone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video game console, an Internet Protocol (IP) camera, a server device in a server system including one or more server devices (e.g., a video streaming server system or other suitable server system), a head-mounted display (HMD), a head-up display (HUD), smart glasses (e.g., virtual reality (VR) glasses, augmented reality (AR) glasses, or other smart glasses), or any other suitable electronic device.

[0057]

[0079] The system components 200 may include and / or be implemented using electronic circuitry or other electronic hardware, which may include SOC 100 and / or one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), a neural processing unit (NPU), and / or other suitable electronic circuitry), and / or may include and / or be implemented using computer software, firmware, or any combination thereof, to perform various operations described herein.

[0058]

[0080] While system 200 is shown to include several components, those skilled in the art will appreciate that system 200 can include more or fewer components than those shown in Figure 2. For example, system 200, in some examples, can also include one or more memory devices other than storage 208 and storage 218 (e.g., one or more random access memory (RAM) components, read-only memory (ROM) components, cache memory components, buffer components, database components, and / or other memory devices), one or more processing devices (e.g., one or more CPUs, GPUs, NPUs, and / or other processing devices) in communication with and / or electrically connected to the one or more memory devices, one or more wireless interfaces for implementing wireless communications (e.g., including one or more transceivers and a baseband processor for each wireless interface), one or more wired interfaces for implementing communications via one or more wired connections (e.g., a serial interface such as a universal serial bus (USB) input, a lightning connector, and / or other wired interfaces), and / or other components not shown in Figure 2.

[0059]

[0081] The coding techniques described herein are applicable to video coding in various multimedia applications, including streaming video transmission (e.g., over the Internet), television broadcasting or transmission, encoding digital video for storage on a data storage medium, decoding digital video stored on a data storage medium, or other applications. In some examples, system 200 may support one-way or two-way video transmission to support applications such as video conferencing, video streaming, video playback, video broadcasting, gaming, and / or video telephony.

[0060]

[0082] In some examples, encoding device 204 (or encoder) may be used to encode video data using a video coding standard or protocol to generate an encoded video bitstream. Examples of video coding standards include ITU-T H.261, ISO / IEC MPEG-1 Visual, ITU-T H.262 or ISO / IEC MPEG-2 Visual, ITU-T H.263, ISO / IEC MPEG-4 Visual, ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC) including its Scalable Video Coding (SVC) extension and Multiview Video Coding (MVC) extension, High Efficiency Video Coding (HEVC) or ITU-T H.265, Versatile Video Coding (VVC) or ITU-T H.266, and / or other video coding standards. One or more of the video coding standards have extensions related to other aspects of video coding. For example, various extensions to HEVC address multi-layer video coding, including range and screen content coding extensions, 3D video coding (3D-HEVC) and multiview extensions (MV-HEVC) and scalable extensions (SHVC).

[0061]

[0083] Many embodiments described herein may be implemented using video codecs such as VVC, HEVC, AVC, and / or extensions thereof. However, the techniques and systems described herein may also be applicable to other coding standards, such as MPEG, JPEG (or other coding standards for still images), VP9, ​​AV1, extensions thereof, or other suitable coding standards, whether already available or not yet available or developed, such as machine learning-based video coding described below. Thus, while the techniques and systems described herein may be described with reference to a particular video coding standard, those skilled in the art will appreciate that the description should not be construed as applying only to that particular standard.

[0062]

[0084] 2, a video source 202 may provide video data to an encoding device 204. The video source 202 may be part of a source device or part of a device other than the source device. The video source 202 may include a video capture device (e.g., a video camera, a camera phone, a video phone, etc.), a video archive containing stored video, a video server or content provider providing video data, a video feed interface receiving video from a video server or content provider, a computer graphics system for generating computer graphics video data, a combination of such sources, or any other suitable video source.

[0063]

[0085] The video data from the video source 202 may include one or more input pictures. A picture is sometimes called a "frame." A picture or frame, in some cases, is a still image that is part of a video. In some examples, the data from the video source 202 may be a still image that is not part of a video. In HEVC, VVC, and other video coding specifications, a video sequence includes a series of pictures. A picture is a sequence of pictures. L , SCb , and S Cr The sample array may include three sample arrays, denoted as S L is a 2D array of luma samples, and S Cb is a two-dimensional array of Cb chrominance samples, and S Cr where Cr is a two-dimensional array of chrominance samples. Chrominance samples are sometimes referred to herein as "chroma" samples. In other cases, a picture may be monochrome and include only an array of luma samples.

[0064]

[0086] The encoder engine 206 (or encoder) of the encoding device 204 encodes video data to generate an encoded video bitstream. In some examples, an encoded video bitstream (or "video bitstream" or "bitstream") is a series of one or more coded video sequences. According to HEVC, a coded video sequence (CVS) includes a series of AUs starting with an access unit (AU) in a base layer that has a random access point picture with some property (e.g., a RASL flag (e.g., NoRaslOutputFlag) equal to 1) up to, but not including, the next AU in the base layer that has a random access point picture with some property. An AU includes one or more coded pictures and control information corresponding to coded pictures that share the same output time. At the bitstream level, coded slices of a picture are encapsulated in data units called network abstraction layer (NAL) units. For example, an HEVC video bitstream may include one or more CVSs that include NAL units. Each NAL unit has a NAL unit header. Syntax elements in the NAL unit header take designated bits and are therefore visible to all kinds of systems and transport layers, such as transport streams, real-time transport (RTP) protocols, file formats, among others.

[0065]

[0087] Two classes of NAL units exist in the HEVC standard, including video coding layer (VCL) NAL units and non-VCL NAL units. VCL NAL units contain coded picture data that form a coded video bitstream. For example, a sequence of bits that form a coded video bitstream resides in a VCL NAL unit. A VCL NAL unit may contain one slice or slice segment (described below) of coded picture data, while a non-VCL NAL unit contains control information related to one or more coded pictures. In some cases, NAL units may be referred to as packets. An HEVC AU includes VCL NAL units that contain coded picture data and non-VCL NAL units that correspond to the coded picture data (if any). Non-VCL NAL units may contain, in addition to other information, parameter sets with high-level information related to the coded video bitstream. For example, parameter sets may include a video parameter set (VPS), a sequence parameter set (SPS), and a picture parameter set (PPS). In some cases, each slice or other portion of the bitstream may reference a single active PPS, SPS, and / or VPS to enable the decoding device 212 to access information that can be used to decode the slice or other portion of the bitstream.

[0066]

[0088] An NAL unit may contain a sequence of bits (e.g., an encoded video bitstream, a CVS of a bitstream, etc.) that form a coded representation of video data, such as a coded representation of a picture in a video. The encoder engine 206 generates coded representations of pictures by partitioning each picture into multiple slices. Slices are independent of other slices such that information in a slice is coded without dependency on data from other slices within the same picture. A slice includes one or more slice segments, including independent slice segments, and one or more dependent slice segments, if present, that depend on previous slice segments.

[0067]

[0089] In HEVC, a slice is partitioned into coding tree blocks (CTBs) of luma samples and chroma samples. A CTB of luma samples and one or more CTBs of chroma samples, together with the syntax for the samples, are called a coding tree unit (CTU). A CTU is sometimes called a "treeblock" or "largest coding unit" (LCU). A CTU is the basic processing unit for HEVC encoding. A CTU can be split into multiple coding units (CUs) of various sizes. A CU contains luma and chroma sample arrays called coding blocks (CBs).

[0068]

[0090] The luma and chroma CBs may be further split into prediction blocks (PBs). A PB is a block of luma or chroma component samples that uses the same motion parameters for inter prediction or intra block copy (IBC) prediction (when available or enabled for use). A luma PB and one or more chroma PBs, together with associated syntax, form a prediction unit (PU). For inter prediction, a set of motion parameters (e.g., one or more motion vectors, reference indexes, etc.) is signaled in the bitstream for each PU and used for inter prediction of the luma PB and one or more chroma PBs. The motion parameters are sometimes referred to as motion information. A CB may also be partitioned into one or more transform blocks (TBs). A TB represents a square block of color component samples to which a residual transform (e.g., the same two-dimensional transform in some cases) is applied to code the prediction residual signal. A transform unit (TU) represents a TB of luma and chroma samples and corresponding syntax elements. Transform coding is described in more detail below.

[0069]

[0091] The size of a CU corresponds to the size of a coding mode and may be square in shape. For example, the size of a CU may be 8x8 samples, 16x16 samples, 32x32 samples, 64x64 samples, or any other suitable size up to the size of the corresponding CTU. The phrase "NxN" is used herein to refer to the pixel dimensions of a video block in vertical and horizontal dimensions (e.g., 8 pixels x 8 pixels). The pixels in a block may be arranged in rows and columns. In some embodiments, a block may not have the same number of pixels in the horizontal direction as in the vertical direction. Syntax data associated with a CU may, for example, represent the partitioning of the CU into one or more PUs. The partitioning mode may differ between whether the CU is coded in intra-prediction mode or inter-prediction mode. The PU may be partitioned to be non-square in shape. Syntax data associated with a CU may also, for example, represent the partitioning of a CU into one or more TUs according to the CTU. The TUs may be square or non-square in shape.

[0070]

[0092] According to HEVC, transforms may be performed using transform units (TUs). TUs may be different for different CUs. TUs may be sized based on the size of the PUs within a given CU. TUs may be the same size as or smaller than the PUs. In some examples, residual samples corresponding to a CU may be subdivided into smaller units using a quad tree structure known as a residual quad tree (RQT). Leaf nodes of the RQT may correspond to TUs. Pixel difference values ​​associated with the TUs may be transformed to produce transform coefficients. The transform coefficients may be quantized by the encoder engine 206.

[0071]

[0093] Once a picture of video data is partitioned into CUs, the encoder engine 206 predicts each PU using a prediction mode. The prediction unit or prediction block is subtracted from the original video data to obtain a residual (described below). For each CU, a prediction mode may be signaled in the bitstream using syntax data. The prediction mode may include intra-prediction (or intra-picture prediction) or inter-prediction (or inter-picture prediction). Intra-prediction exploits the correlation between spatially adjacent samples within a picture. For example, using intra-prediction, each PU is predicted from neighboring image data in the same picture using, for example, DC prediction to find the average value for the PU, planar prediction to fit a flat surface to the PU, directional prediction to extrapolate from neighboring data, or any other suitable type of prediction. Inter-prediction uses temporal correlation between pictures to derive motion-compensated predictions for blocks of image samples. For example, using inter-prediction, each PU is predicted using motion-compensated prediction from image data in one or more reference pictures (before or after the current picture in output order). The decision of whether to code a picture area using inter-picture prediction or intra-picture prediction may be made, for example, at the CU level.

[0072]

[0094] As mentioned above, in some cases, the encoder engine 206 and the decoder engine 216 (described in more detail below) may be configured to operate according to VVC. According to VVC, a video coder (such as the encoder engine 206 and / or the decoder engine 216) partitions a picture into multiple coding tree units (CTUs) (where a CTB for luma samples and one or more CTBs for chroma samples, together with syntax for the samples, are referred to as a CTU). The video coder may partition the CTUs according to a tree structure, such as a quad-tree binary tree (QTBT) structure or a multi-type tree (MTT) structure. The QTBT structure eliminates the concept of multiple partition types, such as the separation between CUs, PUs, and TUs in HEVC. The QTBT structure includes two levels, including a first level partitioned according to quad-tree partitioning and a second level partitioned according to binary tree partitioning. The root node of the QTBT structure corresponds to a CTU. The leaf nodes of the binary tree correspond to coding units (CUs).

[0073]

[0095] In the MTT partitioning structure, blocks may be partitioned using quadtree partitioning, binary tree partitioning, and one or more types of tripletree partitioning. Tripletree partitioning is a partition in which a block is split into three sub-blocks. In some examples, tripletree partitioning splits a block into three sub-blocks without splitting the original block through the center. Partition types in MTT (e.g., quadtree, binary tree, and tripletree) can be symmetric or asymmetric.

[0074]

[0096] In some examples, the video coder may use a single QTBT or MTT structure to represent each of the luminance and chrominance components, and in other examples, the video coder may use two or more QTBT or MTT structures, such as one QTBT or MTT structure for the luminance component and another QTBT or MTT structure for both chrominance components (or two QTBT and / or MTT structures for each chrominance component).

[0075]

[0097] The video coder may be configured to use quadtree partitioning according to HEVC, QTBT partitioning, MTT partitioning, or other partition structures. For purposes of explanation, the description herein may refer to QTBT partitioning. However, it should be understood that the techniques of this disclosure may also be applied to video coders configured to use quadtree partitioning, or other types of partitioning as well.

[0076]

[0098] As mentioned above, intra-picture prediction exploits the correlation between spatially adjacent samples within a picture. There are multiple intra-prediction modes (also referred to as "intra modes"). In some examples, intra-prediction of luma blocks includes 35 modes, including planar mode, DC mode, and 33 angular modes (e.g., diagonal intra-prediction mode and angular modes adjacent to the diagonal intra-prediction mode). The 35 modes of intra-prediction are indexed as shown below in Table 1. In other examples, more intra-modes may be defined, including prediction angles not already represented by the 33 angular modes. In other examples, the prediction angles associated with the angular modes may differ from those used in HEVC.

[0077] [Table 1]

[0078]

[0099] Inter-picture prediction uses temporal correlation between pictures to derive motion-compensated predictions for blocks of image samples. Using a translational motion model, the position of a block in a previously decoded picture (reference picture) is indicated by a motion vector (Δx, Δy), where Δx specifies the horizontal displacement of the reference block relative to the position of the current block and Δy specifies its vertical displacement. In some cases, the motion vector (Δx, Δy) may be of integer sample precision (also called integer precision), in which case the motion vector points to the integer pel grid (or integer pixel sampling grid) of the reference frame. In some cases, the motion vector (Δx, Δy) may be of fractional sample precision (also called fractional pel precision or non-integer precision) to more accurately capture the movement of underlying objects without being restricted to the integer pel grid of the reference frame. The precision of the motion vector may be represented by the quantization level of the motion vector. For example, the quantization level may be of integer precision (e.g., 1 pixel) or fractional pel precision (e.g., 1 / 4 pixel, 1 / 2 pixel, or other sub-pixel value). When the corresponding motion vector has fractional sample precision, interpolation is applied to the reference picture to derive the prediction signal. For example, samples available at integer positions may be filtered (e.g., using one or more interpolation filters) to estimate values ​​at fractional positions. A previously decoded reference picture is indicated by a reference index (refIdx) into a reference picture list. The motion vector and the reference index may be referred to as motion parameters. Two types of inter-picture prediction may be implemented, including uni-prediction and bi-prediction.

[0079]

[0100] In the case of inter prediction using bi-prediction, two sets of motion parameters (Δx0, y0, refIdx0 and Δx1, y1, refIdx1) are used to generate two motion-compensated predictions (from the same reference picture or possibly from different reference pictures). For example, in the case of bi-prediction, each prediction block uses two motion-compensated prediction signals to generate a B prediction unit. The two motion-compensated predictions are combined to obtain a final motion-compensated prediction. For example, the two motion-compensated predictions may be combined by averaging. In another example, weighted prediction may be used, in which case different weights may be applied to each motion-compensated prediction. Reference pictures that may be used in bi-prediction are stored in two separate lists, denoted as List 0 and List 1. The motion parameters may be derived in the encoder using a motion estimation process.

[0080]

[0101] In the case of inter prediction using uni-prediction, one set of motion parameters (Δx0, y0, refIdx0) is used to generate a motion-compensated prediction from a reference picture. For example, in the case of uni-prediction, each prediction block uses at most one motion-compensated prediction signal to generate a P prediction unit.

[0081]

[0102] A PU may include data related to the prediction process (e.g., motion parameters or other suitable data). For example, when a PU is encoded using intra prediction, the PU may include data representing an intra prediction mode for the PU. As another example, when a PU is encoded using inter prediction, the PU may include data defining a motion vector for the PU. The data defining a motion vector for the PU may represent, for example, a horizontal component (Δx) of the motion vector, a vertical component (Δy) of the motion vector, a resolution of the motion vector (e.g., integer precision, ¼-pixel precision, or ⅛-pixel precision), a reference picture to which the motion vector points, a reference index, a reference picture list for the motion vector (e.g., List 0, List 1, or List C), or any combination thereof.

[0082]

[0103] After performing prediction using intra prediction and / or inter prediction, the encoding device 204 may perform transform and quantization. For example, after prediction, the encoder engine 206 may calculate a residual value corresponding to the PU. The residual value may comprise pixel difference values ​​between the current block of pixels being coded (PU) and a predictive block (e.g., a predicted version of the current block) used to predict the current block. For example, after generating a predictive block (e.g., using inter prediction or intra prediction), the encoder engine 206 may generate a residual block by subtracting the predictive block produced by the prediction unit from the current block. The residual block includes a set of pixel difference values ​​that quantify differences between pixel values ​​of the current block and pixel values ​​of the predictive block. In some examples, the residual block may be represented in a two-dimensional block format (e.g., a two-dimensional matrix or array of pixel values). In such examples, the residual block is a two-dimensional representation of pixel values.

[0083]

[0104] Any residual data that may remain after prediction is performed is transformed using a block transform, which may be based on a discrete cosine transform (DCT), a discrete sine transform (DST), an integer transform, a wavelet transform, other suitable transform functions, or any combination thereof. In some cases, one or more block transforms (e.g., kernels of size 32x32, 16x16, 8x8, 4x4, or other suitable sizes) may be applied to the residual data in each CU. In some examples, TUs may be used for the transform and quantization process implemented by the encoder engine 206. A given CU having one or more PUs may also include one or more TUs. As described in more detail below, residual values ​​may be transformed into transform coefficients using a block transform, and quantized and scanned using TUs to produce serialized transform coefficients for entropy coding.

[0084]

[0105] In some embodiments, after intra-predictive coding or inter-predictive coding using a PU of a CU, the encoder engine 206 may calculate residual data for the TUs of the CU. The PU may comprise pixel data in the spatial domain (or pixel domain). As mentioned above, the residual data may correspond to pixel difference values ​​between pixels of the uncoded picture and predicted values ​​corresponding to the PU. The encoder engine 206 may form one or more TUs including the residual data for the CU (including the PU) and transform the TUs to produce transform coefficients for the CU. The TUs may comprise coefficients in the transform domain after application of a block transform.

[0085]

[0106] The encoder engine 206 may perform quantization of the transform coefficients. Quantization provides further compression by quantizing the transform coefficients to reduce the amount of data used to represent the coefficients. For example, quantization may reduce the bit depth associated with some or all of the coefficients. In one example, a coefficient with an n-bit value may be truncated to an m-bit value during quantization, where n is greater than m.

[0086]

[0107] Once quantization is performed, the coded video bitstream includes the quantized transform coefficients, prediction information (e.g., prediction modes, motion vectors, block vectors, etc.), partition information, and any other suitable data, such as other syntax data. Different elements of the coded video bitstream may be entropy coded by the encoder engine 206. In some examples, the encoder engine 206 may utilize a predefined scan order to scan the quantized transform coefficients to produce serialized vectors that can be entropy coded. In some examples, the encoder engine 206 may perform adaptive scanning. After scanning the quantized transform coefficients to form vectors (e.g., one-dimensional vectors), the encoder engine 206 may entropy code the vectors. For example, the encoder engine 206 may use context-adaptive variable length coding, context-adaptive binary arithmetic coding, syntax-based context-adaptive binary arithmetic coding, probability interval partition entropy coding, or another suitable entropy coding technique.

[0087]

[0108] The output 210 of the encoding device 204 may send the NAL units constituting the encoded video bitstream data to a decoding device 212 of a receiving device via a communication link 220. An input 214 of the decoding device 212 may receive the NAL units. The communication link 220 may include channels provided by a wireless network, a wired network, or a combination of wired and wireless networks. The wireless network may include any wireless interface or combination of wireless interfaces, including any suitable wireless network (e.g., the Internet or other wide area network, packet-based network, WiFi, radio frequency (RF), UWB, WiFi-Direct, cellular, Long Term Evolution (LTE), WiMax, etc.). The wired network may include any wired interface (e.g., fiber, Ethernet, powerline Ethernet, Ethernet over coaxial cable, digital signal line (DSL), etc.). Wired and / or wireless networks may be implemented using a variety of equipment, such as base stations, routers, access points, bridges, gateways, switches, etc. The encoded video bitstream data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to a receiving device.

[0088]

[0109] In some examples, encoding device 204 may store the encoded video bitstream data in storage 208. Output unit 210 may retrieve the encoded video bitstream data from encoder engine 206 or from storage 208. Storage 208 may include any of a variety of distributed or locally accessed data storage media. For example, storage 208 may include a hard drive, a storage disk, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. Storage 208 may also include a decoded picture buffer (DPB) for storing reference pictures for use in inter-prediction. In further examples, storage 208 may correspond to a file server or another intermediate storage device that may store encoded video generated by a source device. In such cases, a receiving device, including decoding device 212, can access the stored video data from the storage device via streaming or download. The file server may be any type of server capable of storing encoded video data and transmitting the encoded video data to a receiving device. Exemplary file servers include web servers (e.g., for websites), FTP servers, network-attached storage (NAS) devices, or local disk drives. Receiving devices may access the encoded video data through any standard data connection, including an Internet connection. Access may include wireless channels (e.g., Wi-Fi connections), wired connections (e.g., DSL, cable modems, etc.), or a combination of both, that are suitable for accessing the encoded video data stored on the file server. Transmission of the encoded video data from storage 208 may be a streaming transmission, a download transmission, or a combination thereof.

[0089]

[0110] The input 214 of the decoding device 212 may receive encoded video bitstream data and provide the video bitstream data to the decoder engine 216 or to the storage 218 for later use by the decoder engine 216. For example, the storage 218 may include a DPB for storing reference pictures for use in inter-prediction. A receiving device including the decoding device 212 may receive encoded video data to be decoded via the storage 208. The encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to the receiving device. The communication medium for the transmitted encoded video data may comprise any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful for enabling communication from a source device to a receiving device.

[0090]

[0111] The decoder engine 216 may decode the encoded video bitstream data by entropy decoding (e.g., using an entropy decoder) to extract elements of one or more coded video sequences that make up the encoded video data. The decoder engine 216 may rescale the encoded video bitstream data and perform an inverse transform on the encoded video bitstream data. Residual data is passed to a prediction stage of the decoder engine 216. The decoder engine 216 predicts blocks of pixels (e.g., PUs). In some examples, the prediction is added to the output of the inverse transform (the residual data).

[0091]

[0112] Video decoding device 212 may output the decoded video to video destination device 222, which may include a display or other output device for displaying the decoded video data to a content consumer. In some aspects, video destination device 222 may be part of a receiving device that includes decoding device 212. In some aspects, video destination device 222 may be part of a separate device other than the receiving device.

[0092]

[0113] In some embodiments, the video encoding device 204 and / or the video decoding device 212 may be integrated with an audio encoding device and an audio decoding device, respectively. The video encoding device 204 and / or the video decoding device 212 may also include other hardware or software necessary to implement the coding techniques described above, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. The video encoding device 204 and the video decoding device 212 may be integrated as part of a combined encoder / decoder (CODEC) in their respective devices.

[0093]

[0114] The exemplary system shown in FIG. 2 is one illustrative example that may be used herein. Techniques for processing video data using the techniques described herein may be implemented by any digital video encoding and / or decoding device. Generally, the techniques of this disclosure are implemented by a video encoding device or a video decoding device, although the techniques may also be implemented by a composite video encoder-decoder, commonly referred to as a “codec.” Additionally, the techniques of this disclosure may also be implemented by a video preprocessor. The source device and the receiving device are merely examples of coding devices, such that the source device generates coded video data for transmission to the receiving device. In some examples, the source device and the receiving device may operate substantially symmetrically, such that each device includes a video encoding component and a video decoding component. Thus, the exemplary system may support one-way or two-way video transmission between video devices, for example, for video streaming, video playback, video broadcasting, or video telephony.

[0094]

[0115] As mentioned above, in some examples, SOC 100 and / or components thereof may be configured to perform video compression and / or decompression (also referred to as video encoding and / or decoding, collectively referred to as video coding) using machine learning techniques. For example, encoding device 204 (or encoder) may be used to encode video data using a machine learning system with a deep learning architecture (e.g., by utilizing NPU 108 of SOC 100 of FIG. 1). In some cases, using a deep learning architecture to perform video compression and / or decompression can increase the efficiency of video compression and / or decompression on the device. For example, encoding device 204 may use machine learning-based video coding techniques to more efficiently compress the video and send the compressed video to decoding device 212, which may decompress the compressed video using machine learning-based techniques.

[0095]

[0116] A neural network is an example of a machine learning system and may include an input layer, one or more hidden layers, and an output layer. Data is provided from input nodes in the input layer, processing is performed by hidden nodes in one or more hidden layers, and output is produced through output nodes in the output layer. Deep learning networks generally include multiple hidden layers. Each layer of a neural network may include a feature map or activation map, which may include artificial neurons (or nodes). The feature map may include filters, kernels, etc. The nodes may include one or more weights used to indicate the importance of one or more nodes in the layer. In some cases, deep learning networks may have a series of many hidden layers, with early layers used to determine simple, low-level characteristics of the input and later layers accumulating a hierarchy of more complex and abstract characteristics.

[0096]

[0117] Deep learning architectures may learn a hierarchy of features. For example, when presented with visual data, a first layer may learn to recognize relatively simple features in the input stream, such as edges. In another example, when presented with auditory data, the first layer may learn to recognize spectral power at specific frequencies. A second layer, taking the output of the first layer as input, may learn to recognize combinations of features, such as simple shapes in the case of visual data or combinations of sounds in the case of auditory data. For example, higher layers may learn to represent complex shapes in visual data or words in auditory data. Even higher layers may learn to recognize common visual objects or spoken phrases.

[0097]

[0118] Deep learning architectures can work particularly well when applied to problems that have a natural hierarchical structure. For example, motor vehicle classification can benefit from first learning to recognize wheels, windshields, and other features. These features can be combined in different ways in higher layers to recognize cars, trucks, and airplanes.

[0098]

[0119] Neural networks can be designed with various connectivity patterns. In feedforward networks, information is passed from lower layers to higher layers, with each neuron in a given layer communicating with neurons in higher layers. As described above, hierarchical representations can be accumulated in successive layers of a feedforward network. Neural networks can also have recurrent or (also called top-down) feedback connections. In recurrent connections, the output from a neuron in a given layer can be communicated to another neuron in the same layer. Recurrent architectures can be useful for recognizing patterns across two or more of the input data chunks delivered sequentially to the neural network. Connections from neurons in a given layer to neurons in a lower layer are called feedback (or top-down) connections. Networks with many feedback connections can be useful when recognizing high-level concepts can help distinguish specific low-level features of the input. Connections between layers of a neural network can be fully or locally connected. Various examples of neural network architectures are described below with reference to Figures 15A-16.

[0099]

[0120] 3 shows a system 300 including a device 302 configured to perform video encoding and decoding using a machine learning coding system 310. The device 302 is coupled to a camera 307 and a storage medium 314 (e.g., a data storage device). In some implementations, the camera 307 is configured to provide image data 308 (e.g., a video data stream) to a processor 304 for encoding by the machine learning coding system 310. In some implementations, the device 302 may be coupled to and / or include multiple cameras (e.g., a dual camera system, three cameras, or other number of cameras). In some cases, the device 302 may be coupled to a microphone and / or other input devices (e.g., a keyboard, a mouse, a touch input device such as a touchscreen and / or touchpad, and / or other input devices). In some examples, the camera 307, the storage medium 314, the microphone, and / or other input devices may be part of the device 302.

[0100]

[0121] The device 302 is also coupled to the second device 390 via a transmission medium 318, such as one or more wireless networks, one or more wired networks, or a combination thereof. For example, the transmission medium 318 may include a channel provided by a wireless network, a wired network, or a combination of a wired network and a wireless network. The transmission medium 318 may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The transmission medium 318 may include routers, switches, base stations, or any other equipment that may be useful for enabling communication from a source device to a receiving device. The wireless network may include any wireless interface or combination of wireless interfaces, including any suitable wireless network (e.g., the Internet or other wide area network, a packet-based network, WiFi, radio frequency (RF), UWB, WiFi-Direct, cellular, Long Term Evolution (LTE), WiMax, etc.). The wired network may include any wired interface (e.g., fiber, Ethernet, powerline Ethernet, Ethernet over coaxial cable, digital signal line (DSL), etc.). Wired and / or wireless networks may be implemented using a variety of equipment, such as base stations, routers, access points, bridges, gateways, switches, etc. The encoded video bitstream data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to a receiving device.

[0101]

[0122] Device 302 includes one or more processors 304 (referred to herein as “processors”) coupled to a memory 306, a first interface (“I / F 1”) 312, and a second interface (“I / F 2”) 316. Processor 304 is configured to receive image data 308 from a camera 307, from memory 306, and / or from a storage medium 314. Processor 304 is coupled to storage medium 314 via first interface 312 (e.g., via a memory bus) and to transmission medium 318 via second interface 316 (e.g., a network interface device, a wireless transceiver and antenna, one or more other network interface devices, or a combination thereof).

[0102]

[0123] The processor 304 includes a machine learning coding system 310. The machine learning coding system 310 includes an encoder portion 362 and a decoder portion 366. In some implementations, the machine learning coding system 310 may include one or more autoencoders. The encoder portion 362 is configured to receive input data 370 and process the input data 370 to generate output data 374 based at least in part on the input data 370.

[0103]

[0124] In some implementations, the encoder portion 362 of the machine learning coding system 310 is configured to perform lossy compression of the input data 370 to generate the output data 374, such that the output data 374 has fewer bits than the input data 370. The encoder portion 362 may be trained to compress the input data 370 (e.g., an image or video frame) without using motion compensation based on any prior representation (e.g., one or more previously reconstructed frames). For example, the encoder portion 362 may compress a video frame using only video data from that video frame and without using any data from previously reconstructed frames. Herein, the video frames processed by the encoder portion 362 may be referred to as intra-predicted frames (I-frames). In some examples, the I-frames may be generated using traditional video coding techniques (e.g., according to HEVC, VVC, MPEG-4, or other video coding standards). In such examples, the processor 304 may include or be coupled to a video coding device (e.g., an encoding device) configured to perform block-based intra prediction, such as that described above with respect to the HEVC standard. In such examples, the machine learning coding system 310 may be excluded from the processor 304.

[0104]

[0125] In some implementations, the encoder portion 362 of the machine learning coding system 310 may be trained to compress input data 370 (e.g., video frames) using motion compensation based on a previous representation (e.g., one or more previously reconstructed frames). For example, the encoder portion 362 may compress a video frame using video data from that video frame and using data from a previously reconstructed frame. Herein, the video frames processed by the encoder portion 362 may be referred to as intra-predicted frames (P-frames). Motion compensation may be used to determine data for a current frame by describing how pixels from a previously reconstructed frame, along with residual information, have moved to new positions in the current frame.

[0105]

[0126] As shown, the encoder portion 362 of the machine learning coding system 310 may include a neural network 363 and a quantizer 364. The neural network 363 may include one or more convolutional neural networks (CNNs), one or more fully connected neural networks, one or more gated recurrent units (GRUs), one or more long short-term memory (LSTM) networks, one or more ConvRNNs, one or more ConvGRUs, one or more ConvLSTMs, one or more GANs, any combination thereof, and / or other types of neural network architectures that generate intermediate data 372. The intermediate data 372 is input to the quantizer 364. The quantizer 364 may be implemented using a machine learning system (e.g., using a neural network system) or may be implemented using standard-based quantization and / or entropy coding techniques (e.g., arithmetic coding). For example, in some cases, the encoder portion 362 may compress the input data 370 using neural network techniques described herein and output intermediate data 372 to a quantizer 364 for performing standards-based quantization and / or entropy coding (e.g., arithmetic coding).

[0106]

[0127] The quantizer 364 is configured to perform quantization, and in some cases, entropy coding, of the intermediate data 372 to produce output data 374. The output data 374 may include quantized (and in some cases, entropy coded) data. The quantization operation performed by the quantizer 364 may result in the generation of quantized codes (or data representing the quantized codes generated by the machine learning coding system 310) from the intermediate data 372. The quantized codes (or data representing the quantized codes) may also be referred to as latent codes or latents (denoted as z). Herein, the entropy model applied to the latents may be referred to as a “prior.” In some examples, the quantization and / or entropy coding operations may be performed using existing quantization and entropy coding operations performed when encoding and / or decoding video data according to existing video coding standards. In some examples, the quantization and / or entropy coding operations may be performed by the machine learning coding system 310. In one illustrative example, the machine learning coding system 310 may be trained using supervised training, during which residual data is used as input and quantized codes and entropy codes are used as known outputs (labels).

[0107]

[0128] The decoder portion 366 of the machine learning coding system 310 is configured to receive output data 374 (e.g., directly from the quantizer 364 and / or from the storage medium 314). The decoder portion 366 may process the output data 374 to generate a representation 376 of the input data 370 based at least in part on the output data 374. In some examples, the decoder portion 366 of the machine learning coding system 310 includes a neural network 368, which may include one or more CNNs, one or more fully connected neural networks, one or more GRUs, one or more long short-term memory (LSTM) networks, one or more ConvRNNs, one or more ConvGRUs, one or more ConvLSTMs, one or more GANs, any combination thereof, and / or other types of neural network architectures.

[0108]

[0129] The processor 304 is configured to send the output data 374 to at least one of the transmission medium 318 or the storage medium 314. For example, the output data 374 may be stored in the storage medium 314 for later retrieval and decoding (or decompression) by the decoder portion 366 to generate a representation 376 of the input data 370 as reconstructed data. The reconstructed data may be used for various purposes, such as for playback of the video data encoded / compressed to generate the output data 374. In some implementations, the output data 374 may be decoded in another decoder device corresponding to the decoder portion 366 (e.g., in the device 302, in a second device 390, or in another device) to generate the representation 376 of the input data 370 as reconstructed data. For example, the second device 390 may include a decoder corresponding to (or substantially corresponding to) the decoder portion 366, and the output data 374 may be transmitted to the second device 390 via the transmission medium 318. The second device 390 can process the output data 374 to generate a representation 376 of the input data 370 as reconstructed data.

[0109]

[0130] The system components 300 may include and / or be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include and / or be implemented using computer software, firmware, or any combination thereof, to perform various operations described herein.

[0110]

[0131] While system 300 is shown to include several components, those skilled in the art will appreciate that system 300 can include more or fewer components than those shown in FIG. 3. For example, system 300 can also include, or be part of, a computing device including input and output devices (not shown). In some implementations, system 300 can also include, or be part of, a computing device that includes one or more memory devices (e.g., one or more random access memory (RAM) components, read-only memory (ROM) components, cache memory components, buffer components, database components, and / or other memory devices), one or more processing devices (e.g., one or more CPUs, GPUs, and / or other processing devices) in communication with and / or electrically connected to the one or more memory devices, one or more wireless interfaces for implementing wireless communications (e.g., including one or more transceivers and a baseband processor for each wireless interface), one or more wired interfaces for implementing communications via one or more wired connections (e.g., serial interfaces such as universal serial bus (USB) inputs, lightning connectors, and / or other wired interfaces), and / or other components not shown in FIG. 3.

[0111]

[0132] In some implementations, system 300 may be implemented locally by and / or included within a computing device. For example, the computing device may include a mobile device, a personal computer, a tablet computer, a virtual reality (VR) device (e.g., a head-mounted display (HMD) or other VR device), an augmented reality (AR) device (e.g., an HMD, AR glasses, or other AR device), a wearable device, a server (e.g., in a Software as a Service (SaaS) system or other server-based system), a television, and / or any other computing device with the resource capabilities to perform the techniques described herein.

[0112]

[0133] In one example, the machine learning coding system 310 may be incorporated into a portable electronic device including a memory 306 coupled to the processor 304 and configured to store instructions executable by the processor 304, and a wireless transceiver coupled to an antenna and the processor 304 and operable to transmit output data 374 to a remote device.

[0113]

[0134] As explained above, entropy coding is one of the final stages (and in some cases the final stage) of encoding (compression) and defines the value and number of bits to be added to the compressed data bitstream. Modern standards-based video encoding methods (e.g., VVC, HEVC, AV1, etc.) employ adaptive arithmetic coding to enable high-quality compression performance. The bitstream generated by adaptive arithmetic coding can only be coded and decoded sequentially. For example, a data element can only be restored by first decoding all previous elements, because the decoder needs to reach the same state the encoder had when it coded that element.

[0114]

[0135] Parallel entropy coding may be implemented, in which the throughput requirements are divided so that they can be processed by less complex circuitry. To enable parallel entropy coding, the compressed data is separated into independently coded bitstream segments that can be coded and decoded simultaneously. Such independently coded bitstream segments are referred to herein as data parcels or parcels. It is assumed that a parcel can be decoded without information from other parcels (this requirement is for the decoding process; interpretation of data in a parcel can depend on information from another parcel) and that each parcel extends to an integer number of bytes.

[0115]

[0136] Whenever there are two or more bitstreams to be decoded simultaneously, in addition to using compressed data, information indicating an entry point (e.g., a byte location where decoding can begin) may be provided for decoding. The data structure containing the entry point is referred to herein as a parcel index. In some cases, the parcel index containing the entry point may be included in a header of the video data, in a parameter set (e.g., a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), etc.), and / or in any other message or signaling related to the video data.

[0116]

[0137] Bidirectional byte packing is one technique that can be used to process parcels. For example, bidirectional byte packing can be used to reduce the number of entry points that need to be specified in the parcel index. Bidirectional byte packing is based on the fact that the locations of the entry points must be known before decoding begins, but the end locations do not need to be stored.

[0117]

[0138] 4 illustrates an example of a data organization and coding process for bidirectional byte packing. As shown, data from two encoded bitstreams (including bitstream 402 and bitstream 404) is first stored in memory buffers (including memory buffer 406 for bitstream 402 and memory buffer 408 for bitstream 404). Once encoding is complete, the final number of bytes required for each stream (byte l for bitstream 402) is calculated. a and byte l for bitstream 404 b ) is known, l a +l b A compressed data byte array 410 is created with the bytes. The forward stream (denoted as "forward" in FIG. 4) is copied in conventional increasing byte order (from left to right in FIG. 4) starting from the beginning of the compressed data byte array 410, and the reverse stream (denoted as "reverse" in FIG. 4) is copied in reverse byte position order (from right to left in FIG. 4) starting from the end of the compressed data byte array 410. The decoder determines the number of compressed data bytes (l a +l b ), the decoder can start decoding from the beginning and the end of the compressed data byte array 410 simultaneously.

[0118]

[0139] 5 illustrates an example of extending the bidirectional byte packing technique to support parallel entropy encoding and decoding. Similar to the example of FIG. 4, an encoded (or compressed) bitstream may be split into several independent parcels (e.g., parcels l1, l2, l3, . . . l), for example, by multiple encoders (e.g., encoder 502, encoder 504, etc.) or by a single encoder. N) Independent parcels may be stored in separate buffers (e.g., parcel l1 is stored in buffer 506). Pairs of parcels may be combined using bidirectional byte packing. For example, bidirectional byte packing may be used to combine parcel l1 with parcel l2, parcel l3 with parcel l4, etc. The compressed data arrangement of compressed data byte array 510 in FIG. 5 enables parallel entropy coding. The compressed data is organized into multiple bidirectional byte arrays, including bidirectional byte array 512, bidirectional byte array 514, and bidirectional byte array 516.

[0119]

[0140] Multiple decoders (e.g., decoder 518, decoder 520, etc.) are also shown in FIG. 5. As shown, a forward decoder (FDEC) and a reverse decoder (BDEC) may be used for each bidirectional byte array. For example, decoder 518 (which is FDEC) and decoder 520 (which is BDEC) may be used to decode bidirectional byte array 512. The FDEC and BDEC decoders may differ only in the order in which they read the compressed data bytes. Compressed data byte array 510 is appended with a parcel index 522. Parcel index 522 is a data structure that contains an indication of the number of bytes in each parcel (and thus the decoding entry point, denoted as the "entry point location"). One technique (called the "NBC" method) that may be used to efficiently code the entry point information is to encode a sequence with the number of bytes in each parcel.

[0120]

number

[0121]

[0141] but,

[0122]

number

[0123] This is based on the idea that a parcel can be coded with an average number of bits per parcel approximately equal to .

[0124]

[0142] When using two-way byte packing, the following sequence needs to be encoded instead (for simplicity, we assume N is even):

[0125]

number

[0126]

[0143] The corresponding average number of bits is

[0127]

number

[0128] is.

[0129]

[0144] This means that bidirectional byte packing can reduce the overhead resulting from encoding the parcel index 522 by approximately half.

[0130]

[0145] Arithmetic coding stream termination is an aspect of arithmetic coding. Arithmetic coding differs from other forms of entropy coding for various reasons. For example, arithmetic coding maintains a "range" state (semi-closed interval) during encoding and decoding that corresponds to a fractional number of pending bits in the information-theoretic sense. Also, when the number of fractional bits reaches a certain threshold, an integer number of bits are written to the output stream, and the state is updated to preserve information corresponding to a fraction of the remaining bits. When encoding is finished, it may be necessary to "flush" the pending bits and add bits to convert the pending information to an integer number of output stream bits. To increase efficiency and avoid checking for an end-of-stream condition every time a byte is read, the decoder can preload some bytes, which ultimately includes some bytes from beyond the end of the current stream. The goal of stream termination is to ensure correct decoding regardless of the values ​​of the bits beyond the end of the stream.

[0131]

[0146] Figure 6 is a diagram that graphically illustrates the factors used for correct arithmetic coding termination. In Figure 6, the interval [u,v) represents the final state of the arithmetic encoder. For example, the interval [u,v) defined by u and v represents a fractional number of bytes (e.g., a 34-bit data set contains 4.2 bytes, and u and v represent 0.2 bytes). Any bit value within the interval [u,v) is valid for the termination byte. For simplicity of presentation, it is assumed that an implementation with byte output (byte-based renormalization) is used and that u and v are

[0132]

number

[0133] It can be assumed that σ is a real number such that σ is a real number.

[0134]

[0147] The minimum number of extra bits required to ensure correct decoding is

[0135]

number

[0136] is equal to the smallest value of exponent s that satisfies

[0137]

[0148] The s most significant bits of the termination byte are

[0138]

number

[0139]

[0149] where the values ​​of the remaining 8-s bits are arbitrary (referred to herein as "do not care bits").

[0140]

[0150] However, when it is possible to select all bits in the termination byte,

[0141]

number

[0142] For any termination byte value b that satisfies

[0143]

[0151] There are some special cases, such as those involving carry operations and the need for an extra byte when s>8.

[0144]

[0152] Is less than or equal to

[0145]

number

[0146]

[0153] , and note that since vu<1, the condition s≧1 always holds (there is always at least one additional bit that needs to be added to the output stream).

[0147]

[0154] The overhead incurred is different for each endpoint and is equal to s+log2(vu). The average value can be calculated from coding simulations. For a high-precision implementation, the average value is

[0148]

number

[0149] It can be measured as being equal to

[0150]

[0155] As mentioned above, systems and techniques are described herein that can reduce such overhead and resulting compression loss. For example, the systems and techniques address the bit overhead that occurs when encoding a parcel is completed. Furthermore, file and data stream formats require bits to be grouped into 8-bit bytes, where there is additional overhead defined by byte boundaries in addition to the overhead associated with the parcel index and coding termination. With unidirectional byte packing, the average overhead (e.g., based on unused bits) per data parcel (where additional bits required to properly terminate arithmetic coding are not taken into account; such additional bits must be added to this value to obtain the total average overhead per parcel) is:

[0151]

number

[0152] is.

[0153]

[0156] There are two choices related to bidirectional byte packing. The first option is to concatenate bytes and have the same average overhead per parcel. FIG. 7A illustrates an example of concatenating bytes. For example, as shown in FIG. 7A, the first terminal byte 702 contains a value of 11011, and the second terminal byte 704 contains a value of 0010. The bits identified with an "x" are called "do not care" bits. In one illustrative example, video data may be coded with 37 bits. A byte contains 8 bits. In the illustrative example of video coded with 37 bits, there would be 4 bytes (containing 32 total bits) and the remaining 5 bits. The remaining five bits and three "don't care" bits may be included in a byte (e.g., bit value 11011 and the "xxx" bits included in first termination byte 702 of FIG. 7A). An encoder and / or decoder can (and in some cases must) write and / or read all five 8-bit bytes, including the last byte containing the "don't care" bit, but the decoder will only parse bits up to the "don't care" bit; for this reason, the encoder can set the "don't care" bit to any value. As shown by bitstream 706, first termination byte 702 is concatenated with second termination byte 704.

[0154]

[0157] A second option is to encode the bits in the reverse stream in reverse order within each byte so that a single shared termination byte (also called a joint termination byte) can be used when the patterns of bits used do not overlap. This is in contrast to the example of FIG. 7A, where two bytes are used in bitstream 706. FIG. 7B illustrates an example of using a shared termination byte 716. For example, as shown in FIG. 7B, shared termination byte 716 includes bits 0 and 1 from first termination byte 712, bits 1, 1, and 1 from second termination byte 714, and three bits of no interest. Because there are a total of five bits (less than eight bits) from first termination byte 712 and second termination byte 714, the bits can be combined into shared termination byte 716.

[0155]

[0158] When the second option is used (e.g., as shown in Figure 7B), the fraction of the trailing bytes that are shared is

[0156]

number

[0157]

[0159] and the average overhead per parcel is

[0158]

number

[0159] is reduced to

[0160] The systems and techniques described herein offer novel solutions to the above-mentioned problems, and in some cases, consider applications for massive parallelization (when termination overhead is important), and consider applications for at least two cases. In the first case, writing bits in reverse order within a byte may be relatively easy when encoding "raw" bits, but it is not always convenient. For example, in arithmetic coding, the bit order comes from the arithmetic operation, and therefore, for a reverse stream, each byte must have its bits reversed before writing and after reading. In the second case, as explained below, the correct termination for arithmetic coding is not unique. In unidirectional byte packing, there is no advantage to choosing between many correct options. However, as explained below, in bidirectional byte packing, a subset of the choices allows for sharing of termination bytes, thus reducing average overhead.

[0161] In some cases, the systems and techniques described herein may be applied to joint binary coding termination. For example, the concept of a shared termination byte may be used when the same convention for filling bits within a byte is used in the forward and reverse streams. FIG. 8A is a diagram illustrating an example using concatenation, and FIG. 8B is a diagram illustrating an example using a shared termination byte using the techniques described herein. For example, a byte value may be shared if the bit positions used for both streams (e.g., for both termination bytes of both streams) have the same bit values ​​and the resulting shared termination byte does not change the bit order of the termination byte of the forward stream and the termination byte of the reverse stream. In some cases, the encoder may determine that a shared termination byte may be used based on determining that the bit positions used in the first termination byte of a first parcel (for the forward stream) and the second termination byte of a second parcel (for the reverse stream) have the same (or common) bit values ​​and the resulting shared termination byte does not change the bit order of the first termination byte and the second termination byte. In some examples, the encoder may determine that a bit position in the first termination byte has a common value with a bit position in the second termination byte by determining an intersection of the value of the first termination byte and the value of the second termination byte. The encoder may include the common bit value (e.g., found to intersect) in the shared termination byte, followed by the other value of the first termination byte and the other value of the second termination byte.

[0162]

[0162] In the example of Figure 8A, the three bit positions containing the value of the first three bit positions (value "101") are used in both the first termination byte 802 of the forward stream and the second termination byte 804 of the reverse stream. However, if a shared termination byte were used, the resulting bits in the shared termination byte would be "101100", which would change the order of the bits in the second termination byte 804 of the reverse stream (which has value "10100"). The order of the bits in each termination byte must be maintained so that correct decoding can be performed by the decoder.

[0163] In Figure 8B, three bit positions containing the value of the first three bit positions (value "100") are used in both the first termination byte 812 of the forward stream and the second termination byte 814 of the reverse stream. Because both the first termination byte 812 and the second termination byte 814 have first three bit values ​​equal to "100," the termination bytes can be shared (in a shared termination byte 816) if the bit values ​​in all occupied positions are copied and the order of the bits in the termination bytes 812 and 814 is maintained (or not changed). Such a condition exists in the example of Figure 8B. The first group 818 of bits (with value "100") of the shared termination byte 816 contains three common bits (the first three bits) from the first termination byte 812 and the second termination byte 814. The next group 820 of bits (with value "1011") includes bits from the second termination byte 814 that are unique (not common) to the first termination byte 812. The final bit of the shared termination byte 816 is a "don't care" bit (represented by an "x"). As shown, the order of the bits in termination byte 812 from the forward stream ("100") and in termination byte 814 from the reverse stream ("1001011") is maintained in the shared termination byte 816.

[0164]

[0164] The decoder is aware of where the bits for termination byte 812 (of the forward stream) are in the shared termination byte 816 and where the bits for termination byte 814 (of the reverse stream) are in the shared termination byte 816. For example, the bits are processed by the decoder according to the data being read and its interpretation. In one illustrative example, the decoder may be defined in such a way that it is to read exactly 1024 values, each value being entropy coded with a different number of bits according to its value (e.g., value 0 coded using 2 bits, values ​​+1, -1 coded with 4 bits, etc.). Because the decoder must know how many bits are used for each value, the decoder is aware that after it has decoded 1024 numbered values, it has reached the end of a particular stream (e.g., the end of the forward stream, and therefore the end of termination byte 812, the end of the reverse stream, and therefore the end of termination byte 814, etc.). Continuing with such an example, there may be two arrays of 1024 values ​​that are coded independently in a similar fashion and packed together using bidirectional packing. The decoder for each stream can decode its own set of 1024 values ​​and can process the bits sequentially. Each decoder will terminate when it has decoded 1024 values. In some cases, the decoder does not know a priori how many bits to read, but the number of data elements the decoder must decode is well-defined, in which case the decoder can continue reading bits from the input stream up to a point where the decoder knows it must stop.

[0165]

[0165] The technique shown in Figure 8B relies on the event that groups of random bit values ​​become equal. Such coincidences are most common for cases where the number of "uninteresting" bits (or "wasted bits") is greatest. For example, if there is only one additional bit in each stream, the end byte can be shared in 50% of cases. This can mean that instead of always having 14 unused bits for those cases, there can be 7 unused bits in 1 / 2 of the cases, and an average of 10.5 unused bits. The exact proportion of end bytes that can be shared can be shown as follows:

[0166]

number

[0167]

[0166] Also, the average overhead per parcel is

[0168]

number

[0169] It can be calculated as:

[0170]

[0167] In some cases, the systems and techniques described herein may be applied to joint arithmetic coding termination (with the same bit order). For example, if the arithmetic coding termination is performed using equation (7) from above, it is straightforward to use equations (11)-(13) and the techniques described with respect to Figures 7A-7B and / or 8A-8B and use a shared termination byte based on the concept of "don't care" bits.

[0171] However, as explained above with respect to equations (5)-(10) and FIG. 6, the termination byte may be freely selected according to equation (8). For one-way byte packing, there may be no advantage to selecting a particular value. For two-way byte packing, the systems and techniques described herein can take advantage of the freedom to select the termination byte by selecting a value that can reduce the average overhead. FIGS. 9A and 9B provide two examples illustrating such a concept.

[0172] In FIG. 9A, the gray area is (b min and b max represents two sets or ranges, F and B, for the forward and reverse streams, with ranges of termination byte values ​​allowed for correct decoding, as defined in equation (8) (grey area 920 represents set F, and grey area 922 represents set B). Sets F and B may be defined as follows:

[0173]

number

[0174]

[0170] The dashed line 925 in Figures 9A and 9B represents the use of the same termination byte value in both the forward stream (represented on the x-axis of Figures 9A and 9B) and the reverse stream (represented on the y-axis). In the example of Figure 9A, there is an intersection between set F and set B (the intersection of set F and set B is non-empty), and all byte values ​​in that intersecting set (shown as bold solid line segment 926 along the dashed line) can be used for the shared termination byte. In the example of Figure 9B, the intersection between set F and set B is empty (because dashed line 925 does not pass through both set F and set B simultaneously), meaning that a shared termination byte cannot be shared.

[0175]

[0171] The range intersection calculation may be determined as follows:

[0176]

number

[0177]

[0172] where ∩ indicates an intersection operation. Figures 10A and 10B illustrate a termination byte selection process based on the concepts shown in Figures 9A and 9B. Figure 10A shows an example where no intersection exists between set F and set B. In Figure 10B, the encoder can determine that a shared termination byte 1016 can be used based on determining that an intersection exists between the bit value of a first termination byte 1012 of a first parcel (forward stream in Figure 10B) and the bit value of a second termination byte 1014 of a second parcel (reverse stream in Figure 10B). For example, a value for the shared termination byte 1016 can be calculated by determining the intersection of the values ​​of the first termination byte 1012 and the second termination byte 1014. The shared terminating byte 1016 is represented as byte c in FIG. 10B and includes all values ​​within the intersecting set or range F′ and set or range B′ (represented by segment 926 in FIG. 9A), and is shown in FIG. 10B as follows:

[0178]

number

[0179]

[0173] Thus, the values ​​that can be contained in the shared termination byte 1016 include the intersecting values ​​from the first termination byte 1012 and the second termination byte 1014.

[0180] By maximizing the use of shared termination bytes (e.g., using the techniques described with respect to Figures 8A-10B), the systems and techniques described herein can reduce the overhead of the bitstream and compression loss. For example, the encoded bitstream will contain fewer bytes, resulting in reduced overhead in the bitstream. Such a solution can result in significant overhead reduction in parallel entropy coding where many forward and reverse streams are included in the bitstream (e.g., such as that shown in Figure 5, which includes bidirectional byte array 512, bidirectional byte array 514, -bidirectional byte array 516).

[0181]

[0175] Note that in the examples shown in Figures 7A, 7B, 8A, and 8B, the only factor considered was the bit value. In the examples of Figures 10A and 10B, which relate to arithmetic coding, there is a more general notion of a set of valid termination bytes. For example, as shown in Figures 10A and 10B, there is the notation a∈F', b∈B' because the range defined in equation (16) is greater than 256.

[0182]

number

[0183] and,

[0184]

number

[0185] This occurs because of the carries that can occur in arithmetic coding addition, including during the terminal. Therefore, the set

[0186]

number

[0187] is defined as:

[0188] Also, during termination, it may be necessary to identify the termination byte value that corresponds to a carry. Figures 11A and 11B show an example where such a case occurs. It can be observed that in Figures 11A and 11B, the determination of the intersection is slightly more complicated (e.g., compared to the illustrations of Figures 9A and 9B), because it is necessary to take into account four possible conditions (i.e., for the forward and backward streams, to carry or not), which means that equation (17) must be modified for each particular case.

[0189]

[0177] Simulation results show that such carry cases may need to be considered for each stream in only 24.6% of the terminations, with carries occurring in 5.5% of the terminations.

[0190] In some cases, the systems and techniques described herein can be applied to joint arithmetic coding termination (with different bit ordering). For example, if the bytes produced by arithmetic coding in the reverse stream are converted to the reverse bit order using the function R(n), the set of valid termination bytes shown in Figures 10A and 10B becomes

[0191]

number

[0192] is defined by

[0193] 12A and 12B show a new set based on reverse bit ordering, using the same termination range as the example in FIGS. 9A and 9B. In fact, it can be observed that reversing the bits in the bytes of the reverse stream "spreads" the values ​​of valid termination bytes in the interval [0, 255], increasing the occurrence of valid intersections. This is confirmed in the simulation results described below. In some cases, the same results can be obtained when the function R(n) is a randomly generated permutation of byte values.

[0194] In some examples, the systems and techniques described herein for determining a shared (or joint) termination byte may be used to encode syntax elements for video coding (e.g., syntax elements in an encoded video bitstream) using traditional standards-based coding and / or using machine learning-based coding. For example, a joint termination byte may be determined to signal a bit of the syntax element in the bitstream.

[0195] Next, simulation results are described. The bidirectional byte packing and termination techniques described herein were tested on an implementation of arithmetic coding with 32-bit precision, using random data from an alphabet of 23 symbols and entropy equal to 2.61 bits / symbol. The average was 2 25 Measured after coding (33 million) parcels.

[0196] Table I below shows simulation results, from which it can be observed that using the systems and techniques described herein for implementing joint termination of arithmetic coding (AC) in the forward and reverse bitstreams significantly increases the proportion of termination bytes that can be shared, thus reducing the average overhead per parcel.

[0197] [Table 2]

[0198] Table II below shows the overall benefits of bidirectional byte packing, co-termination, and the efficient NBC method (described with respect to equations (1)-(4)) for coding the parcel index. The relative overhead is the average number of overhead bits divided by the number of data bits in the parcel.

[0199] [Table 3]

[0200]

[0184] From the simulation results in Tables I and II, it can be observed that in conventional approaches using unidirectional byte packing with universal codes (e.g., Exponential-Golomb), the number of bits defining the overhead of parcel index coding + parcel termination is quite large for small parcel sizes, growing as twice the logarithm of the average number of bytes per parcel.

[0201]

[0185] In comparison, using the systems and techniques described herein, the overhead for parcels with small byte counts is reduced by a significant factor, with the number of overhead bits growing as half the logarithm of the average number of bytes per parcel. In practice, the above simulation results show that using the systems and techniques described herein, the scope of parallelization can be extended to be employed for entropy coding, reducing the cost of achieving high data throughput without compromising compression efficiency.

[0202]

[0186] Figure 13 is a flowchart illustrating an example of a process 1300 for processing video using the co-termination techniques described herein. At block 1302, process 1300 includes obtaining encoded video data. In some cases, the encoded video data may include video data encoded according to a particular video coding standard (e.g., HEVC, VVC, AVC, EVC, etc.) using a machine learning-based video coding system (e.g., in which video frames are coded using one or more neural networks) and / or other video coding techniques. As obtained at block 1302, the encoded video data is not entropy coded. At block 1304, process 1300 includes determining an intersection of values ​​between a value for a first termination byte of a first parcel of the encoded video data and a value for a second termination byte of a second parcel of the encoded video data. Referring to Figure 9A as an illustrative example, a bold solid line segment 926 along the dashed line indicates the intersection of values ​​between the termination byte of the reverse stream and the termination byte of the forward stream.

[0203] At block 1306, the process 1300 includes determining a joint (or shared) termination byte for the first termination byte of the first parcel and the second termination byte of the second parcel, where the value for the joint termination byte is based on an intersection of the values ​​(the value of the intersection). The joint termination byte may be the final termination byte of the first parcel and the second parcel for processing. For example, referring to FIG. 10B as an illustrative example, the shared (or shared) termination byte 1016 is the final termination byte of the forward stream and the reverse stream.

[0204] At block 1308, process 1300 includes generating entropy-coded data including a co-terminating byte for the first parcel and the second parcel. In some examples, the entropy-coded data is generated using arithmetic coding. For example, in some cases, the value for the first terminating byte includes a first range of terminating byte values ​​allowed for decoding, and the value for the second terminating byte includes a second range of terminating byte values ​​allowed for decoding. In such cases, the intersection of the values ​​determined in block 1304 (the intersection value) includes a value that is in the first range and the second range. For example, referring to FIG. 10B as an illustrative example, the intersection may be determined as the intersection between range F′ and range or set B′, in which case the co-terminating byte may be determined as byte c∈F′∩B′, as described above.

[0205] In some examples, the entropy-coded data is generated using binary coding. In some cases, the value for the first termination byte includes a first number of bits, and the value for the second termination byte includes a second number of bits. In such cases, the intersection of the values ​​determined in block 1304 includes at least one of a common value in the first number of bits and the second number of bits, and a subset of values ​​from the first number of bits and a subset of values ​​from the second number of bits. In some examples, the order of the first number of bits and the order of the second number of bits are unchanged in the joint termination byte compared to the order of the first number of bits in the first termination byte and the order of the second number of bits in the second termination byte. For example, referring to FIG. 8B as an illustrative example, the value "100" is common to both the first termination byte 812 and the second termination byte 814. The common bit "100" is included in the first group 818 of bits (having the value "100") of the joint or shared termination byte 816. The next group 820 of bits in the shared termination byte 816 (with value "1011") includes bits from the second termination byte 814 that are unique (not common) to the first termination byte 812. The final bit of the shared termination byte 816 is a "don't care" bit (represented by an "x"). As shown, the order of the bits in termination byte 812 from the forward stream ("100") and the bits in termination byte 814 from the reverse stream ("1001011") is unchanged in the shared termination byte 816.

[0206] In some examples, process 1300 may generate entropy-coded data by performing parallel entropy encoding of a first parcel and a second parcel. For example, the first parcel may be encoded using a first encoder, and the second parcel may be encoded using a second encoder. Referring to FIG. 5 as an illustrative example, the first parcel may be encoded by encoder 502 (e.g., to generate parcel l of entropy-coded data), and the second parcel may be encoded by encoder 504 (e.g., to generate parcel l of entropy-coded data).

[0207]

[0191] In some examples, process 1300 includes storing a first parcel in a first buffer and storing a second parcel in a second buffer. Referring to Figure 5 as an illustrative example, a first parcel of entropy coded data (e.g., parcel l1 encoded by encoder 502) may be stored in buffer 506.

[0208] In some examples, the process 1300 may include transmitting a bitstream including the entropy coded data (e.g., using a transmitter of the encoding device 204 of FIG. 2). In some examples, the process 1300 may include storing the bitstream including the entropy coded data (e.g., stored in the storage 208 of the encoding device 204 of FIG. 2).

[0209] In some examples, process 1300 may include performing parallel entropy decoding of the first parcel and the second parcel using a joint termination byte for the first parcel and the second parcel. The parallel entropy decoding may be performed using the techniques described above. For example, process 1300 may include reading the first parcel in forward order and reading the second parcel in reverse order. In some cases, process 1300 may include converting the bytes of the second parcel to the reverse order. In some cases, once the data is entropy decoded, other decoding operations may be performed (e.g., by performing intra prediction, inter prediction, etc.).

[0210] In some cases, the encoded video data comprises one or more syntax elements of a video bitstream. In some aspects, the one or more syntax elements indicate one or more parameters that define a neural network for decoding the encoded video data. For example, the one or more parameters that define the neural network may include weights of the neural network (e.g., trained using backpropagation, as described below), one or more activation functions of the neural network, and / or other parameters of the neural network.

[0211]

[0195] Figure 14 is a flowchart illustrating another example of a process 1400 for processing video using the co-termination techniques described herein. At block 1402, the process 1400 includes obtaining a first parcel of entropy-coded data and a second parcel of entropy-coded data. In some examples, the process 1400 includes obtaining the first parcel from a first buffer and obtaining the second parcel from a second buffer. Referring to Figure 5 as an illustrative example, the first parcel of entropy-coded data (e.g., parcel l1 encoded by encoder 502) may be obtained from buffer 506. In some examples, the entropy-coded data includes video data. In some examples, the entropy-coded data includes image data.

[0212]

[0196] The first parcel and the second parcel share a joint (or shared) termination byte. The value for the joint termination byte is based on the intersection of the values ​​(intersection value) between the value for the first termination byte of the first parcel and the value of the second termination byte of the second parcel. Referring to Figure 9A as an illustrative example, the bold solid line segment 926 along the dashed line indicates the intersection of the values ​​between the termination byte of the reverse stream and the termination byte of the forward stream.

[0213] In some cases, the process 1400 includes receiving an encoded video bitstream including a first parcel and a second parcel. In some examples, the encoded video bitstream includes one or more syntax elements. In some aspects, the one or more syntax elements indicate one or more parameters that define a neural network for decoding the encoded video data. For example, the one or more parameters that define the neural network may include weights of the neural network (e.g., trained using backpropagation, as described below), one or more activation functions of the neural network, and / or other parameters of the neural network.

[0214]

[0198] In some examples, the entropy-coded data may be generated using arithmetic coding. For example, in some cases, the value for the first termination byte includes a first range of termination byte values ​​that are permitted for decoding, and the value for the second termination byte includes a second range of termination byte values ​​that are permitted for decoding. In such cases, the intersection of the values ​​(the intersection value) includes values ​​that are in the first range and the second range. For example, referring to Figure 10B as an illustrative example, the intersection may be determined as the intersection between range F' and range or set B', in which case the co-termination byte may be determined as byte c∈F'∩B', as described above.

[0215] In some examples, the entropy-coded data may be generated using binary coding. In some cases, the value for the first termination byte includes a first number of bits, and the value for the second termination byte includes a second number of bits. In such cases, the intersection of the values ​​(the intersection value) includes at least one of a common value in the first and second number of bits, a subset of values ​​from the first and second number of bits, and a subset of values ​​from the second number of bits. In some examples, the order of the first and second number of bits is unchanged in the joint termination byte compared to the order of the first and second number of bits in the first and second termination byte. For example, referring to FIG. 8B as an illustrative example, the value "100" is common to both the first termination byte 812 and the second termination byte 814 and is included in the first group 818 of bits (having the value "100") of the joint or shared termination byte 816. The next group 820 of bits in the shared termination byte 816 (with value "1011") includes bits from the second termination byte 814 that are unique (not common) to the first termination byte 812. The final bit of the shared termination byte 816 is a "don't care" bit (represented by an "x"). As shown, the order of the bits in termination byte 812 from the forward stream ("100") and the bits in termination byte 814 from the reverse stream ("1001011") is unchanged in the shared termination byte 816.

[0216] In some examples, a first parcel of entropy coded data may be decoded using a first decoder and a second parcel of entropy coded data may be decoded using a second decoder. Referring to Figure 5 as an illustrative example, the first parcel may be decoded by decoder 518 and the second parcel may be encoded by decoder 520.

[0217] At block 1404, process 1400 includes performing parallel entropy decoding of the first parcel and the second parcel using the joint termination byte for the first parcel and the second parcel. The parallel entropy decoding may be performed using techniques described herein. For example, process 1400 may include reading the first parcel in forward order and reading the second parcel in reverse order. In some cases, process 1400 may include converting the bytes of the second parcel to the reverse order. Once the data is entropy decoded, other decoding operations may be performed (e.g., by performing intra prediction, inter prediction, etc.).

[0218] In some examples, the processes described herein (e.g., process 1300, process 1400, and / or other processes described herein) may be performed by a computing device or apparatus, such as a computing device having computing device architecture 1700 shown in FIG. 17. In some examples, the computing device may include a mobile device (e.g., a mobile phone, a tablet computing device, etc.), a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a video server, a television, a vehicle (or a computing device in a vehicle), a robotic device, and / or any other computing device with the resource capabilities to perform the processes described herein, including process 1300 and / or process 1400.

[0219] In some cases, a computing device or apparatus may include various components such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more transmitters, receivers or combined transmitter-receivers (e.g., referred to as transceivers), one or more cameras, one or more sensors, and / or other component(s) configured to perform the steps of the processes described herein. In some examples, a computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other component(s). The network interface may be configured to communicate and / or receive Internet Protocol (IP)-based data or other types of data.

[0220] Components of a computing device may be implemented in circuitry. For example, components may include and / or be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), a neural processing unit (NPU), and / or other suitable electronic circuitry), and / or may include and / or be implemented using computer software, firmware, or any combination thereof, to perform various operations described herein.

[0221] Process 1300 and process 1400 are illustrated as logical flow diagrams, whose operations represent sequences of operations that may be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform particular functions or implement particular data types. The order in which the operations are described should not be construed as a limitation, and any number of the described operations may be combined in any order and / or in parallel to implement a process.

[0222]

[0206] Furthermore, the processes described herein (including process 1300, process 1400, and / or other processes described herein) may be performed under the control of one or more computer systems configured of executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that collectively execute on one or more processors, by hardware, or a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0223]

[0207] As mentioned above, some video coding systems utilize neural networks or other machine learning systems to compress video and / or image data. Neural networks can be designed with a variety of connectivity patterns. In feedforward networks, information is passed from lower layers to higher layers, with each neuron in a given layer communicating with neurons in higher layers. As explained above, hierarchical representations can be accumulated in successive layers of a feedforward network. Neural networks can also have recurrent or (also called top-down) feedback connections. In recurrent connections, the output from a neuron in a given layer can be communicated to another neuron in the same layer. Recurrent architectures can be useful for recognizing patterns across two or more of the input data chunks delivered sequentially to the neural network. Connections from neurons in a given layer to neurons in lower layers are called feedback (or top-down) connections. Networks with many feedback connections can be useful when recognizing high-level concepts can help discriminate certain low-level features of the input.

[0224]

[0208] Connections between layers of a neural network can be fully connected or locally connected. FIG. 15A shows an example of a fully connected neural network 1502. In a fully connected neural network 1502, a neuron in a first layer can communicate its output to every neuron in a second layer, such that each neuron in the second layer receives input from every neuron in the first layer. FIG. 15B shows an example of a locally connected neural network 1504. In a locally connected neural network 1504, a neuron in a first layer can be connected to a limited number of neurons in the second layer. More generally, the locally connected layers of a locally connected neural network 1504 can be configured so that each neuron in the layer has the same or similar connectivity pattern, but with connection strengths that can have different values ​​(e.g., 1510, 1512, 1514, and 1516). The connectivity patterns of local connections can give rise to spatially distinct receptive fields in the upper layers, as upper layer neurons in a given region can receive inputs that are conditioned through training to the properties of a restricted subset of the total inputs to the network.

[0225]

[0209] An example of a locally connected neural network is a convolutional neural network. FIG. 15C shows an example of a convolutional neural network 1506. The convolutional neural network 1506 may be configured such that the connection strengths associated with the inputs for each neuron in the second layer are shared (e.g., 1508). Convolutional neural networks may be suitable for problems in which the spatial location of the inputs is meaningful. The convolutional neural network 1506 may be used to implement one or more aspects of video compression and / or decompression according to aspects of the present disclosure.

[0226]

[0210] One type of convolutional neural network is the deep convolutional network (DCN). Figure 15D shows a detailed example of a DCN 1500 designed to recognize visual features from images 1526 input from an image capture device 1530, such as an on-board camera. The DCN 1500 in this example may be trained to identify traffic signs and numbers provided on the traffic signs. Of course, the DCN 1500 may be trained for other tasks, such as identifying lane markings or identifying traffic signals.

[0227]

[0211] The DCN 1500 may be trained using supervised learning. During training, the DCN 1500 may be presented with an image, such as an image 1526 of a speed limit sign, and then a forward pass may be computed to produce the output 1522. The DCN 1500 may include a feature extraction section and a classification section. Upon receiving the image 1526, the convolution layer 1532 may apply a convolution kernel (not shown) to the image 1526 to generate a first set of feature maps 1518. As an example, the convolution kernel for the convolution layer 1532 may be a 5x5 kernel that generates a 28x28 feature map. In this example, four different feature maps are generated in the first set of feature maps 1518, so four different convolution kernels were applied to the image 1526 in the convolution layer 1532. A convolution kernel may also be referred to as a filter or a convolution filter.

[0228] The first set of feature maps 1518 may be subsampled by a max pooling layer (not shown) to generate a second set of feature maps 1520. The max pooling layer reduces the size of the first set of feature maps 1518. That is, the size of the second set of feature maps 1520, such as 14×14, is smaller than the size of the first set of feature maps 1518, such as 28×28. The reduced size provides similar information to subsequent layers while reducing memory consumption. The second set of feature maps 1520 may be further convolved via one or more subsequent convolutional layers (not shown) to generate one or more subsequent sets of feature maps (not shown).

[0229] In the example of FIG. 15D , the second set of feature maps 1520 are convolved to generate a first feature vector 1524. Furthermore, the first feature vector 1524 is further convolved to generate a second feature vector 1528. Each feature in the second feature vector 1528 may include a number corresponding to a possible feature of the image 1526, such as “sign,” “60,” and “100.” A softmax function (not shown) may convert the numbers in the second feature vector 1528 into probabilities. Thus, the output 1522 of the DCN 1500 is the probability that the image 1526 contains one or more features.

[0230] In this example, the probabilities in output 1522 for "sign" and "60" are higher than the probabilities for others of output 1522, such as "30," "40," "50," "70," "80," "90," and "100." Prior to training, the output 1522 produced by DCN 1500 may be inaccurate. Thus, an error may be calculated between output 1522 and a target output. The target output is the ground truth of image 1526 (e.g., "sign" and "60"). The weights of DCN 1500 may then be adjusted so that output 1522 of DCN 1500 is more closely aligned with the target output.

[0231]

[0215] To adjust the weights, the learning algorithm may calculate a gradient vector for the weights. The gradient may indicate the amount by which the error increases or decreases if the weights are adjusted. In the top layer, the gradient may correspond directly to the value of the weights connecting activated neurons in the penultimate layer to neurons in the output layer. In lower layers, the gradient may depend on the value of the weights and the calculated error gradient of the upper layer. The weights may then be adjusted to reduce the error. This manner of adjusting weights is sometimes called "backpropagation" because it involves a "backward pass" through the neural network.

[0232]

[0216] In practice, the error gradient of the weights may be calculated over a small number of examples so that the calculated gradient approximates the true error gradient. This approximation method is sometimes called stochastic gradient descent. Stochastic gradient descent may be repeated until the achievable error rate of the overall system no longer decreases, or until the error rate reaches a target level. After training, the DCN may be presented with new images, and a forward pass through the network may result in an output 1522, which may be considered the DCN's inference or prediction.

[0233]

[0217] A deep belief network (DBN) is a probabilistic model with multiple layers of hidden nodes. DBNs can be used to extract hierarchical representations of a training dataset. DBNs can be obtained by stacking layers of restricted Boltzmann machines (RBMs). RBMs are a type of artificial neural network that can learn probability distributions over a set of inputs. Because RBMs can learn probability distributions in the absence of information about the class into which each input should be categorized, RBMs are often used in unsupervised learning. Using a hybrid unsupervised and supervised paradigm, the lower RBM of the DBN can be trained in an unsupervised manner and act as a feature extractor, and the upper RBM can be trained in a supervised manner (on the joint distribution of inputs and target classes from the previous layer) and act as a classifier.

[0234]

[0218] A deep convolutional network (DCN) is a network of convolutional networks composed of additional pooling and normalization layers. DCNs have achieved state-of-the-art performance for many tasks. DCNs can be trained using supervised learning, where both the input and output targets are known for many examples and are used to modify the network weights using gradient descent methods.

[0235]

[0219] A DCN can be a feedforward network. Furthermore, as explained above, connections from neurons in the first layer of a DCN to groups of neurons in the next higher layer are shared across neurons in the first layer. The feedforward and shared connections of a DCN can be exploited for high-speed processing. The computational burden of a DCN can be much less than that of a similarly sized neural network with, for example, recurrent or feedback connections.

[0236]

[0220] The processing of each layer of a convolutional network can be thought of as a spatially invariant template or basis projection. If the input is initially decomposed into multiple channels, such as the red, green, and blue channels of a color image, a convolutional network trained on that input can be thought of as three-dimensional, with two spatial dimensions along the axes of the image and a third dimension capturing color information. The outputs of the convolutional connections can be thought of as forming feature maps in subsequent layers, where each element of a feature map (e.g., 1520) can receive input from various neurons in the previous layer (e.g., feature map 1518) and from each of multiple channels. Values ​​in the feature map can be further processed using nonlinearities, such as rectification, max(0,x), etc. Values ​​from neighboring neurons can be further pooled, which corresponds to downsampling and can provide further local invariance and dimensionality reduction.

[0237]

[0221] Figure 16 is a block diagram illustrating an example of a deep convolutional network 1650. The deep convolutional network 1650 may include multiple different types of layers based on connectivity and weight sharing. As shown in Figure 16, the deep convolutional network 1650 includes convolutional blocks 1654A, 1654B. Each of the convolutional blocks 1654A, 1654B may be composed of a convolutional layer (CONV) 1656, a normalization layer (LNorm) 1658, and a max pooling layer (MAX POOL) 1660.

[0238] The convolution layer 1656 may include one or more convolution filters, which may be applied to the input data 1652 to generate feature maps. While only two convolution blocks 1654A, 1654B are shown, the disclosure is not so limited; instead, any number of convolution blocks (e.g., blocks 1654A, 1654B) may be included in the deep convolutional network 1650 according to design preference. The normalization layer 1658 may normalize the outputs of the convolution filters. For example, the normalization layer 1658 may perform whitening or lateral suppression. The max-pooling layer 1660 may perform downsampling aggregation across space for local invariance and dimensionality reduction.

[0239] For example, the parallel filter bank of the deep convolutional network may be loaded onto the CPU 102 or GPU 104 of the SOC 100 to achieve high performance and low power consumption. In alternative embodiments, the parallel filter bank may be loaded onto the DSP 106 or ISP 116 of the SOC 100. Additionally, the deep convolutional network 1650 may access other processing blocks that may be present on the SOC 100, such as the sensor processor 114 and the navigation module 120, which are dedicated to sensors and navigation, respectively.

[0240] The deep convolutional network 1650 may also include one or more fully connected layers, such as layer 1662A (labeled "FC1") and layer 1662B (labeled "FC2"). The deep convolutional network 1650 may further include a logistic regression (LR) layer 1664. Between each layer 1656, 1658, 1660, 1662A, 1662B, 1664 of the deep convolutional network 1650, there are weights (not shown) that must be updated. The output of each of the layers (e.g., 1656, 1658, 1660, 1662A, 1662B, 1664) may serve as an input of a subsequent one of the layers (e.g., 1656, 1658, 1660, 1662A, 1662B, 1664) in the deep convolutional network 1650 to learn a hierarchical feature representation from the input data 1652 (e.g., image, audio, video, sensor data, and / or other input data) provided in the first one of the convolutional blocks 1654A. The output of the deep convolutional network 1650 is a classification score 1666 for the input data 1652. The classification score 1666 may be a set of probabilities, where each probability is the probability that the input data includes a feature from the set of features.

[0241] FIG. 17 illustrates an exemplary computing device architecture 1700 of an exemplary computing device that may implement various techniques described herein. In some examples, the computing device may include a mobile device, a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a video server, a vehicle (or a computing device in a vehicle), or other device. For example, the computing device architecture 1700 may be used as part of the system 200 of FIG. 2 and / or the system 300 of FIG. 3. The components of the computing device architecture 1700 are shown in electrical communication with each other using a connection 1705, such as a bus. The exemplary computing device architecture 1700 includes a processing unit (CPU or processor) 1710 and computing device connections 1705 that couple various computing device components, including computing device memory 1715, such as read-only memory (ROM) 1720 and random access memory (RAM) 1725, to the processor 1710.

[0242] The computing device architecture 1700 may include a cache of high-speed memory directly connected to the processor 1710, in close proximity to the processor 1710, or integrated as part of the processor 1710. The computing device architecture 1700 may copy data from the memory 1715 and / or the storage device 1730 to the cache 1712 for quick access by the processor 1710. In this manner, the cache may provide performance improvements that avoid delays to the processor 1710 while waiting for data. These and other modules may control or be configured to control the processor 1710 to perform various actions. Other computing device memory 1715 may also be available for use. The memory 1715 may include multiple different types of memory with different performance characteristics. Processor 1710 can include any general-purpose processor, as well as hardware or software services, such as service 1 1732, service 2 1734, and service 3 1736 stored in storage device 1730, configured to control processor 1710, and special-purpose processors in which software instructions are incorporated into the processor design. Processor 1710 can be a self-contained system including multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors can be symmetric or asymmetric.

[0243] To enable user interaction with computing device architecture 1700, input device(s) 1745 can represent any number of input mechanisms, such as a microphone for audio, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice, etc. Output device(s) 1735 can also be one or more of several output mechanisms known to those skilled in the art, such as a display, projector, television, speaker device, etc. In some cases, a multimodal computing device can allow a user to provide multiple types of input to communicate with computing device architecture 1700. Communications interface 1740 can generally govern and manage user input and computing device output. There is no limitation to operating on any particular hardware configuration, and therefore the basic features herein can be easily substituted with improved hardware or firmware configurations as they are developed.

[0244]

[0228] The storage device 1730 is non-volatile memory and may be a hard disk or other type of computer-readable medium capable of storing data accessible by a computer, such as a magnetic cassette, a flash memory card, a solid-state memory device, a digital versatile disk, a cartridge, random access memory (RAM) 1725, read-only memory (ROM) 1720, and hybrids thereof. The storage device 1730 may include services 1732, 1734, 1736 for controlling the processor 1710. Other hardware or software modules are contemplated. The storage device 1730 may be connected to the computing device connections 1705. In one aspect, a hardware module that performs a specific function may include software components stored on a computer-readable medium in relation to the necessary hardware components, such as the processor 1710, connections 1705, output devices 1735, etc., to perform that function.

[0245] Aspects of the present disclosure are applicable to any suitable electronic device (such as a security system, smartphone, tablet, laptop computer, vehicle, drone, or other device) that includes or is coupled to one or more active depth-sensing systems. Although described below with respect to a device having or coupled to one light projector, aspects of the present disclosure are applicable to devices having any number of light projectors and are therefore not limited to any particular device.

[0246] The term "device" is not limited to one or a specific number of physical objects (e.g., a smartphone, a controller, a processing system, etc.). As used herein, a device may be any electronic device having one or more parts and capable of implementing at least some portions of the present disclosure. While the following description and examples use the term "device" to describe various aspects of the present disclosure, the term "device" is not limited to a specific configuration, type, or number of objects. Furthermore, the term "system" is not limited to multiple components or a specific embodiment. For example, a system may be implemented on one or more printed circuit boards or other substrates and may have movable or static components. While the following description and examples use the term "system" to describe various aspects of the present disclosure, the term "system" is not limited to a specific configuration, type, or number of objects.

[0247]

[0231] Specific details are provided in the above description to provide a thorough understanding of the embodiments and examples provided herein. However, those skilled in the art will understand that the embodiments may be practiced without these specific details. For clarity of explanation, in some instances, the technology may be presented as including individual functional blocks, including devices, device components, steps or routines in a method implemented in software, or functional blocks comprising a combination of hardware and software. Additional components other than those shown in the figures and / or described herein may be used. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form so as not to obscure the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail to avoid obscuring the embodiments.

[0248]

[0232] Individual embodiments may be described above as a process or method that is depicted as a flowchart, flow diagram, data flow diagram, structure diagram, or block diagram. While a flowchart may describe operations as a sequential process, many of the operations may be performed in parallel or simultaneously. Moreover, the order of operations may be rearranged. A process is terminated when the operations of a process are completed, but may have additional steps not included in the diagram. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to a return of the function to the calling function or the main function.

[0249]

[0233] The processes and methods according to the examples described above may be implemented using computer-executable instructions stored on or otherwise available from a computer-readable medium. Such instructions may include, for example, instructions and data that cause or otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform a certain function or group of functions. Portions of the computer resources used may be accessible over a network. The computer-executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, firmware, source code, etc.

[0250] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instruction(s) and / or data. Computer-readable media may include non-transitory media on which data may be stored, which does not include carrier waves and / or transitory electronic signals propagating wirelessly or over wired connections. Examples of non-transitory media may include, but are not limited to, magnetic disks or tapes, optical storage media such as flash memory, memory or memory devices, magnetic or optical disks, flash memory, USB devices with non-volatile memory, network-attached storage devices, compact discs (CDs) or digital versatile discs (DVDs), among others. A computer-readable medium may have code and / or machine-executable instructions stored thereon, which may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.

[0251] In some embodiments, computer-readable storage devices, media, and memories may include cable or wireless signals containing bitstreams, etc. However, when stated, non-transitory computer-readable storage media specifically excludes media such as energy, carrier signals, electromagnetic waves, and the signals themselves.

[0252] Devices implementing processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, program code or code segments (e.g., a computer program product) to perform the necessary tasks may be stored on a computer-readable or machine-readable medium. Processor(s) may perform the necessary tasks. Common examples of form factors include laptops, smartphones, mobile phones, tablet devices or other small form-factor personal computers, personal digital assistants, rackmount devices, standalone devices, etc. The functionality described herein may also be embodied in peripheral devices or add-in cards. Such functionality may also be implemented on a circuit board among different chips or different processes executing in a single device, as further examples.

[0253]

[0237] The instructions, media for carrying such instructions, computing resources for executing them, and other structures for supporting such computing resources are exemplary means for providing the functionality described in this disclosure.

[0254]

[0238] In the foregoing description, aspects of the present application have been described with reference to specific embodiments thereof, but those skilled in the art will recognize that the present application is not limited thereto. Accordingly, while illustrative embodiments of the present application have been described in detail herein, it should be understood that the inventive concepts may, in some cases, be variously embodied and employed, and the appended claims are intended to be construed to include such variations, except as limited by the prior art. Various features and aspects of the applications described above may be used individually or together. Moreover, the embodiments may be utilized in any number of environments and applications other than those described herein without departing from the broader spirit and scope of the present specification. Accordingly, the specification and drawings should be considered illustrative and not limiting. For purposes of explanation, methods have been described in a particular order. It should be appreciated that in alternative embodiments, methods may be performed in an order different from that described.

[0255]

[0239] Those skilled in the art will appreciate that the symbols or terminology used herein for less than ("<") and greater than (">") may be replaced with the symbols less than or equal to ("≦") and greater than or equal to ("≧"), respectively, without departing from the scope of this specification.

[0256]

[0240] When a component is described as being "configured to" perform a certain operation, such configuration may be achieved, for example, by designing electronic circuitry or other hardware to perform the operation, by programming a programmable electronic circuit (e.g., a microprocessor or other suitable electronic circuitry) to perform the operation, or by any combination thereof.

[0257]

[0241] The phrase "coupled to" refers to any component that is physically connected to another component, either directly or indirectly, and / or that is in communication with another component, either directly or indirectly (e.g., connected to the other component via a wired or wireless connection and / or other suitable communication interface).

[0258]

[0242] Claim language or other language reciting "at least one of" a set and / or "one or more" of a set indicates that one member of the set or multiple members of the set (in any combination) satisfies the claim. For example, claim language reciting "at least one of A and B" or "at least one of A or B" means A, B, or A and B. As another example, claim language reciting "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, or A, B, and C. The language "at least one of" a set and / or "one or more" of a set does not limit the set to the items listed in the set. For example, claim language reciting "at least one of A and B" or "at least one of A or B" can mean A, B, or A and B, and can further include items not recited in the set of A and B.

[0259]

[0243] The various illustrative logic blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0260] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as a general-purpose computer, a wireless communication device handset, or an integrated circuit device having multiple uses, including applications in wireless communication device handsets and other devices. Features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be realized, at least in part, by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may comprise a memory or data storage medium, such as a random access memory (RAM) such as synchronous dynamic random access memory (SDRAM), a read-only memory (ROM), a nonvolatile random access memory (NVRAM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic or optical data storage medium, or the like. The techniques may additionally or alternatively be realized at least in part by a computer-readable communications medium, such as a propagated signal or radio waves, that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer.

[0261] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the term “processor,” as used herein, may refer to any of the above structures, any combination of the above structures, or any other structure or apparatus suitable for implementing the techniques described herein.

[0262]

[0246] Exemplary aspects of the present disclosure include the following.

[0263] Aspect 1: A method for processing video data. The method comprises obtaining encoded video data, determining a value intersection between a value for a first termination byte of a first parcel of the encoded video data and a value for a second termination byte of a second parcel of the encoded video data, determining a co-termination byte for the first termination byte of the first parcel and the second termination byte of the second parcel, and generating entropy coded data including a co-termination byte for the first parcel and the second parcel, wherein the value for the co-termination byte is based on the value intersection.

[0264]

[0248] Aspect 2: The method of aspect 1, further comprising generating the entropy coded data using arithmetic coding.

[0265]

[0249] Aspect 3: A method as described in any one of aspects 1 or 2, wherein the value for the first termination byte comprises a first range of termination byte values ​​that are permitted to be decoded, and wherein the value for the second termination byte comprises a second range of termination byte values ​​that are permitted to be decoded, and wherein an intersection of the values ​​comprises values ​​that are in the first range and the second range.

[0266]

[0250] Aspect 4: The method of aspect 1, wherein the entropy-coded data is generated using binary coding.

[0267]

[0251] Aspect 5: A method according to any one of aspects 1 or 4, wherein the value for the first termination byte comprises a first number of bits, and wherein the value for the second termination byte comprises a second number of bits, and wherein an intersection of the values ​​comprises at least one of a common value in the first number of bits and the second number of bits, and a subset of values ​​from the first number of bits and a subset of values ​​from the second number of bits.

[0268]

[0252] Aspect 6: The method described in aspect 5, wherein the order of the first number of bits and the order of the second number of bits are not changed in the co-termination byte compared to the order of the first number of bits in the first termination byte and the order of the second number of bits in the second termination byte.

[0269]

[0253] Aspect 7: The method of any one of aspects 1 to 6, wherein generating the entropy-coded data includes performing parallel entropy coding of the first parcel and the second parcel.

[0270]

[0254] Aspect 8: The method of any one of aspects 1 to 7, wherein a first parcel is encoded using a first encoder, and wherein a second parcel is encoded using a second encoder.

[0271]

[0255] Aspect 9: The method of any one of aspects 1 to 8, further comprising storing the first parcel in a first buffer and storing the second parcel in a second buffer.

[0272]

[0256] Aspect 10: The method of any one of aspects 1 to 9, further comprising transmitting a bitstream including entropy-coded data.

[0273]

[0257] Aspect 11: The method of any one of aspects 1 to 10, further comprising storing a bitstream including the entropy coded data.

[0274]

[0258] Aspect 12: A method according to any one of aspects 1 to 11, further comprising performing parallel entropy decoding of the first parcel and the second parcel using a joint termination byte for the first parcel and the second parcel.

[0275] Aspect 13: The method of any one of aspects 1 to 12, further comprising reading the first parcel in forward order and reading the second parcel in reverse order.

[0276]

[0260] Aspect 14: The method of any one of aspects 1 to 13, further comprising converting the bytes of the second parcel to reverse order.

[0277]

[0261] Aspect 15: The method of any one of aspects 1 to 14, wherein the joint termination byte is the final termination byte of the first parcel and the second parcel for processing.

[0278]

[0262] Aspect 16: The method of any one of aspects 1 to 15, wherein the encoded video data comprises one or more syntax elements of a video bitstream.

[0279]

[0263] Aspect 17: The method of aspect 16, wherein the one or more syntax elements indicate one or more parameters that define a neural network for decoding the encoded video data.

[0280]

[0264] Aspect 18: The method of aspect 17, wherein the one or more parameters defining the neural network comprise at least one of weights of the neural network and an activation function of the neural network.

[0281]

[0265] Aspect 19: An apparatus for processing video data. The apparatus includes a memory configured to store video data and one or more processors coupled to the memory (e.g., implemented in a circuit). The one or more processors are configured to obtain encoded video data, determine a value intersection between a value for a first termination byte of a first parcel of the encoded video data and a value for a second termination byte of a second parcel of the encoded video data, determine a co-termination byte for the first termination byte of the first parcel and the second termination byte of the second parcel, and generate entropy coded data including the co-termination byte for the first parcel and the second parcel, wherein the value for the co-termination byte is based on the value intersection.

[0282]

[0266] Aspect 20: The apparatus of aspect 19, wherein the one or more processors are configured to use arithmetic coding to generate the entropy-coded data.

[0283]

[0267] Aspect 21: The apparatus of any one of aspects 19 or 20, wherein the value for the first termination byte comprises a first range of termination byte values ​​that are permitted to be decoded, and wherein the value for the second termination byte comprises a second range of termination byte values ​​that are permitted to be decoded, and wherein an intersection of the values ​​comprises a value that is in the first range and the second range.

[0284]

[0268] Aspect 22: The apparatus of aspect 19, wherein the one or more processors are configured to use binary coding to generate the entropy-coded data.

[0285]

[0269] Aspect 23: The apparatus of any one of aspects 19 or 22, wherein the value for the first termination byte comprises a first number of bits, and wherein the value for the second termination byte comprises a second number of bits, and wherein an intersection of the values ​​comprises at least one of a common value in the first number of bits and the second number of bits, and a subset of values ​​from the first number of bits and a subset of values ​​from the second number of bits.

[0286]

[0270] Aspect 24: The apparatus described in aspect 23, wherein the order of the first number of bits and the order of the second number of bits are not changed in the co-termination byte compared to the order of the first number of bits in the first termination byte and the order of the second number of bits in the second termination byte.

[0287]

[0271] Aspect 25: An apparatus described in any one of aspects 19 to 24, wherein one or more processors are configured to perform parallel entropy coding of the first parcel and the second parcel to generate entropy coded data.

[0288]

[0272] Aspect 26: The apparatus of any one of aspects 19 to 25, further comprising a first encoder configured to encode the first parcel and a second encoder configured to encode the second parcel.

[0289]

[0273] Aspect 27: The apparatus of any one of aspects 19 to 26, further comprising a first buffer configured to store the first parcel and a second buffer configured to store the second parcel.

[0290]

[0274] Aspect 28: The apparatus of any one of aspects 19 to 27, wherein the one or more processors are configured to transmit a bitstream including entropy-coded data.

[0291]

[0275] Aspect 29: The apparatus of any one of aspects 19 to 28, wherein the one or more processors are configured to store a bitstream including entropy coded data.

[0292]

[0276] Aspect 30: The apparatus of any one of aspects 19 to 29, wherein one or more processors are configured to perform parallel entropy decoding of the first parcel and the second parcel using a joint termination byte for the first parcel and the second parcel.

[0293]

[0277] Aspect 31: The apparatus of any one of aspects 19 to 30, wherein the one or more processors are configured to read a first parcel in a forward order and to read a second parcel in a reverse order.

[0294]

[0278] Aspect 32: The apparatus of any one of aspects 19 to 31, wherein the one or more processors are configured to convert the bytes of the second parcel to reverse order.

[0295]

[0279] Aspect 33: The apparatus of any one of aspects 19 to 32, wherein the joint termination byte is a final termination byte of the first parcel and the second parcel for processing.

[0296]

[0280] Aspect 34: The apparatus of any one of aspects 19 to 33, wherein the encoded video data comprises one or more syntax elements of a video bitstream.

[0297]

[0281] Aspect 35: The apparatus of aspect 34, wherein the one or more syntax elements indicate one or more parameters that define a neural network for decoding the encoded video data.

[0298]

[0282] Aspect 36: The apparatus of aspect 35, wherein the one or more parameters defining the neural network comprise at least one of weights of the neural network and an activation function of the neural network.

[0299]

[0283] Aspect 37: The device described in any one of aspects 19 to 36, wherein the processor includes a neural processing unit (NPU).

[0300]

[0284] Aspect 38: An apparatus described in any one of aspects 19 to 37, wherein the apparatus is a mobile device.

[0301]

[0285] Aspect 39: An apparatus described in any one of aspects 19 to 37, wherein the apparatus is an extended reality device.

[0302]

[0286] Aspect 40: The device of any one of aspects 19 to 37, wherein the device is a television.

[0303]

[0287] Aspect 41: A device described in any one of aspects 19 to 40, further comprising a display.

[0304]

[0288] Aspect 42: An apparatus described in any one of aspects 19 to 41, wherein the apparatus comprises a camera configured to capture one or more video frames.

[0305] Aspect 43: A computer-readable storage medium storing instructions that, when executed, cause one or more processors to perform any of the operations recited in aspects 1-42.

[0306]

[0290] Aspect 44: An apparatus comprising means for performing any of the operations recited in aspects 1 to 42.

[0307]

[0291] Aspect 45: A method for processing video data, the method comprising: obtaining a first parcel of entropy coded data and a second parcel of entropy coded data; and performing parallel entropy decoding of the first parcel and the second parcel using a common termination byte for the first parcel and the second parcel, wherein the first parcel and the second parcel share a common termination byte, and wherein a value for the common termination byte is based on an intersection of values ​​between a value for the first termination byte of the first parcel and a value for the second termination byte of the second parcel.

[0308]

[0292] Aspect 46: The method of aspect 45, wherein the entropy-coded data is encoded using arithmetic coding.

[0309]

[0293] Aspect 47: A method according to any one of aspects 45 or 46, wherein the value for the first termination byte comprises a first range of termination byte values ​​that are permitted to be decoded, and wherein the value for the second termination byte comprises a second range of termination byte values ​​that are permitted to be decoded, and wherein an intersection of the values ​​comprises a value that is in the first range and the second range.

[0310]

[0294] Aspect 48: The method of aspect 45, wherein the entropy-coded data is generated using binary coding.

[0311]

[0295] Aspect 49: The method of any one of aspects 45 or 48, wherein the value for the first terminal byte comprises a first number of bits, and wherein the value for the second terminal byte comprises a second number of bits, and wherein an intersection of the values ​​comprises at least one of a common value in the first number of bits and the second number of bits, and a subset of values ​​from the first number of bits and a subset of values ​​from the second number of bits.

[0312]

[0296] Aspect 50: The method described in aspect 49, wherein the order of the first number of bits and the order of the second number of bits are not changed in the co-termination byte compared to the order of the first number of bits in the first termination byte and the order of the second number of bits in the second termination byte.

[0313]

[0297] Aspect 51: The method of any one of aspects 45 to 50, further comprising obtaining a first parcel from a first buffer and obtaining a second parcel from a second buffer.

[0314]

[0298] Aspect 52: The method of any one of aspects 45 to 51, further comprising reading the first parcel in forward order and reading the second parcel in reverse order.

[0315]

[0299] Aspect 53: The method of any one of aspects 45 to 52, further comprising converting the bytes of the second parcel to reverse order.

[0316]

[0300] Aspect 54: The method of any one of aspects 45 to 53, wherein the joint termination byte is the final termination byte of the first parcel and the second parcel for processing.

[0317]

[0301] Aspect 55: A method described in any one of aspects 45 to 54, further comprising receiving a video bitstream, the video bitstream including a first parcel, a second parcel, and one or more syntax elements.

[0318]

[0302] Aspect 56: The method described in aspect 55, wherein the one or more syntax elements indicate one or more parameters that define a neural network for decoding the encoded video data.

[0319]

[0303] Aspect 57: The method of aspect 56, wherein the one or more parameters defining the neural network comprise at least one of weights of the neural network and an activation function of the neural network.

[0320]

[0304] Aspect 58: An apparatus for processing video data. The apparatus includes a memory configured to store video data and one or more processors coupled to the memory (e.g., implemented in a circuit). The one or more processors are configured to: obtain a first parcel of entropy coded data and a second parcel of entropy coded data; and perform parallel entropy decoding of the first parcel and the second parcel using a co-termination byte for the first parcel and the second parcel, where the first parcel and the second parcel share a co-termination byte, and where a value for the co-termination byte is based on an intersection of values ​​between a value for the first termination byte of the first parcel and a value of the second termination byte of the second parcel.

[0321]

[0305] Aspect 59: The apparatus of aspect 58, wherein the entropy-coded data is encoded using arithmetic coding.

[0322]

[0306] Aspect 60: The apparatus of any one of aspects 58 or 59, wherein the value for the first termination byte comprises a first range of termination byte values ​​that are permitted to be decoded, and wherein the value for the second termination byte comprises a second range of termination byte values ​​that are permitted to be decoded, and wherein an intersection of the values ​​comprises a value that is in the first range and the second range.

[0323]

[0307] Aspect 61: The apparatus of aspect 58, wherein the entropy-coded data is generated using binary coding.

[0324]

[0308] Aspect 62: The apparatus of any one of aspects 58 or 61, wherein the value for the first terminal byte comprises a first number of bits, and wherein the value for the second terminal byte comprises a second number of bits, and wherein an intersection of the values ​​comprises at least one of a common value in the first number of bits and the second number of bits, and a subset of values ​​from the first number of bits and a subset of values ​​from the second number of bits.

[0325]

[0309] Aspect 63: The apparatus described in aspect 62, wherein the order of the first number of bits and the order of the second number of bits are not changed in the co-termination byte compared to the order of the first number of bits in the first termination byte and the order of the second number of bits in the second termination byte.

[0326]

[0310] Aspect 64: The apparatus described in any one of aspects 58 to 63, wherein one or more processors are configured to obtain a first parcel from a first buffer and obtain a second parcel from a second buffer.

[0327]

[0311] Aspect 65: The apparatus described in any one of aspects 58 to 64, wherein the one or more processors are configured to read a first parcel in a forward order and to read a second parcel in a reverse order.

[0328]

[0312] Aspect 66: The apparatus of any one of aspects 58 to 65, wherein the one or more processors are configured to convert the bytes of the second parcel to reverse order.

[0329]

[0313] Aspect 67: The apparatus of any one of aspects 58 to 66, wherein the joint termination byte is a final termination byte of the first parcel and the second parcel for processing.

[0330]

[0314] Aspect 68: An apparatus described in any one of aspects 58 to 67, wherein one or more processors are configured to receive a video bitstream, the video bitstream including a first parcel, a second parcel, and one or more syntax elements.

[0331]

[0315] Aspect 69: The apparatus of aspect 68, wherein the one or more syntax elements indicate one or more parameters that define a neural network for decoding the encoded video data.

[0332]

[0316] Aspect 70: The apparatus described in aspect 69, wherein the one or more parameters defining the neural network comprise at least one of weights of the neural network and an activation function of the neural network.

[0333]

[0317] Aspect 71: An apparatus described in any one of aspects 58 to 70, wherein the processor includes a neural processing unit (NPU).

[0334]

[0318] Aspect 72: An apparatus described in any one of aspects 58 to 71, wherein the apparatus is a mobile device.

[0335]

[0319] Aspect 73: An apparatus described in any one of aspects 58 to 71, wherein the apparatus is an extended reality device.

[0336]

[0320] Aspect 74: The device of any one of aspects 58 to 71, wherein the device is a television.

[0337]

[0321] Embodiment 75: The device described in any one of embodiments 58 to 74, further comprising a display.

[0338]

[0322] Aspect 76: An apparatus described in any one of aspects 58 to 75, wherein the apparatus comprises a camera configured to capture one or more video frames.

[0339] Aspect 77: A computer-readable storage medium storing instructions that, when executed, cause one or more processors to perform any of the operations recited in aspects 58-76.

[0340]

[0324] Aspect 78: An apparatus comprising means for performing any of the operations recited in aspects 58 to 76.

[0341]

[0325] Aspect 79: A method configured to perform any of the operations recited in aspects 1 to 42 and aspects 58 to 76.

[0342]

[0326] Aspect 80: An apparatus for processing video data. The apparatus includes a memory configured to store video data and one or more processors coupled to the memory (e.g., implemented in circuitry). The one or more processors are configured to perform any of the operations recited in aspects 1 through 42 and aspects 58 through 76.

[0343]

[0327] Aspect 81: A computer-readable storage medium storing instructions that, when executed, cause one or more processors to perform any of the operations recited in aspects 1 to 42 and aspects 58 to 76.

[0344]

[0328] Aspect 82: An apparatus comprising means for performing any of the operations recited in aspects 1 to 42 and aspects 58 to 76. The inventions described in the claims of the present application as originally filed are set forth below. [C1] 1. A method for processing video data, comprising: Obtaining encoded video data; determining an intersection of a value between a value for a first termination byte of a first parcel of the encoded video data and a value for a second termination byte of a second parcel of the encoded video data; determining a co-termination byte for the first termination byte of the first parcel and the second termination byte of the second parcel, wherein a value for the co-termination byte is based on the intersection of values; generating entropy coded data including the joint termination byte for the first parcel and the second parcel; A method comprising: [C2] The method of C1, wherein the entropy coded data is generated using arithmetic coding. [C3] The method of claim 2, wherein the value for the first termination byte comprises a first range of termination byte values ​​that are permitted to be decoded, and the value for the second termination byte comprises a second range of termination byte values ​​that are permitted to be decoded, and the intersection of values ​​comprises values ​​that are in the first range and the second range. [C4] The method of C1, wherein the entropy coded data is generated using binary coding. [C5] The method of C4, wherein the value for the first termination byte comprises a first number of bits and the value for the second termination byte comprises a second number of bits, and the intersection of values ​​comprises at least one of a common value between the first number of bits and the second number of bits and a subset of values ​​from the first number of bits and a subset of values ​​from the second number of bits. [C6] The method of C5, wherein the order of the first number of bits and the order of the second number of bits are unchanged in the co-termination byte compared to the order of the first number of bits in the first termination byte and the order of the second number of bits in the second termination byte. [C7] The method of C1, wherein generating the entropy coded data includes performing parallel entropy coding of the first parcel and the second parcel. [C8] The method of C7, wherein the first parcel is encoded using a first encoder and the second parcel is encoded using a second encoder. [C9] performing parallel entropy decoding of the first parcel and the second parcel using the joint termination byte for the first parcel and the second parcel; The method of C1, further comprising: [C10] reading the first parcel in forward order; reading the second parcel in reverse order; The method of C9, further comprising: [C11] converting the bytes of said second parcel into reverse order; The method of C10, further comprising: [C12] The method of C1, wherein the joint termination byte is a final termination byte of the first parcel and the second parcel for processing. [C13] The method of C1, wherein the encoded video data comprises one or more syntax elements of a video bitstream. [C14] The method of C13, wherein the one or more syntax elements indicate one or more parameters defining a neural network for decoding the encoded video data. [C15] The method of C14, wherein the one or more parameters defining the neural network comprise at least one of weights of the neural network and an activation function of the neural network. [C16] 1. An apparatus for processing video data, comprising: a memory configured to store video data; one or more processors coupled to the memory; wherein the processor: Obtaining encoded video data; determining an intersection of a value between a value for a first termination byte of a first parcel of the encoded video data and a value for a second termination byte of a second parcel of the encoded video data; determining a co-termination byte for the first termination byte of the first parcel and the second termination byte of the second parcel, wherein a value for the co-termination byte is based on the intersection of values; generating entropy coded data including the joint termination byte for the first parcel and the second parcel; An apparatus configured to: [C17] The apparatus of C16, wherein the one or more processors are configured to use arithmetic coding to generate the entropy-coded data. [C18] 16. The apparatus of claim 15, wherein the value for the first termination byte comprises a first range of termination byte values ​​that are permitted to be decoded, the value for the second termination byte comprises a second range of termination byte values ​​that are permitted to be decoded, and the intersection of values ​​comprises values ​​that are in the first range and the second range. [C19] The apparatus of C16, wherein the one or more processors are configured to use binary coding to generate the entropy-coded data. [C20] 13. The apparatus of claim 19, wherein the value for the first termination byte comprises a first number of bits and the value for the second termination byte comprises a second number of bits, and the intersection of values ​​comprises at least one of a common value in the first number of bits and the second number of bits, and a subset of values ​​from the first number of bits and a subset of values ​​from the second number of bits. [C21] The apparatus of C20, wherein an order of the first number of bits and an order of the second number of bits are unchanged in the co-termination byte compared to an order of the first number of bits in the first termination byte and an order of the second number of bits in the second termination byte. [C22] The apparatus of C16, wherein the one or more processors are configured to perform parallel entropy coding of the first parcel and the second parcel to generate the entropy coded data. [C23] a first encoder configured to encode the first parcel; a second encoder configured to encode the second parcel; The apparatus of C22, further comprising: [C24] the one or more processors: performing parallel entropy decoding of the first parcel and the second parcel using the joint termination byte for the first parcel and the second parcel; The apparatus of C16, configured to perform [C25] the one or more processors: reading the first parcel in forward order; reading the second parcel in reverse order; 20. The apparatus of claim 19, configured to: [C26] The apparatus of C16, wherein the joint termination byte is a final termination byte of the first parcel and the second parcel for processing. [C27] The apparatus of C16, wherein the encoded video data comprises one or more syntax elements of a video bitstream. [C28] The apparatus of C27, wherein the one or more syntax elements indicate one or more parameters defining a neural network for decoding the encoded video data. [C29] 29. The apparatus of claim 28, wherein the one or more parameters defining the neural network comprise at least one of weights of the neural network and an activation function of the neural network. [C30] The device of C16, wherein the device is one of a mobile device, an extended reality device, or a television, and the device further comprises at least one of a display and a camera configured to capture one or more video frames.

Claims

1. 1. A method for processing video data, comprising: Obtaining encoded video data; determining a value intersection between a value for a first termination byte of a first parcel of the encoded video data and a value for a second termination byte of a second parcel of the encoded video data, wherein the value for the first termination byte comprises a first range of termination byte values ​​permitted for decoding and the value for the second termination byte comprises a second range of termination byte values ​​permitted for decoding, and the value intersection includes values ​​in the first range and the second range. determining a co-termination byte for the first termination byte of the first parcel and the second termination byte of the second parcel, wherein a value for the co-termination byte is based on the intersection of values; generating entropy coded data including the joint termination byte for the first parcel and the second parcel; A method comprising:

2. The method of claim 1 , wherein the entropy coded data is generated using arithmetic coding.

3. the entropy coded data is generated using binary coding; the values ​​for the first termination byte include a first number of bits and the values ​​for the second termination byte include a second number of bits, and the intersection of values ​​includes at least one of a common value between the first number of bits and the second number of bits, and a subset of values ​​from the first number of bits and a subset of values ​​from the second number of bits; 2. The method of claim 1, wherein an order of the first number of bits and an order of the second number of bits are unchanged in the co-termination byte compared to an order of the first number of bits in the first termination byte and an order of the second number of bits in the second termination byte.

4. generating the entropy coded data includes performing parallel entropy coding of the first parcel and the second parcel; The method of claim 1 , wherein the first parcel is encoded using a first encoder and the second parcel is encoded using a second encoder.

5. performing parallel entropy decoding of the first parcel and the second parcel using the joint termination byte for the first parcel and the second parcel; Furthermore, reading the first parcel in forward order; reading the second parcel in reverse order; Furthermore, converting the bytes of said second parcel to reverse order; The method of claim 1 further comprising:

6. 2. The method of claim 1, wherein the joint termination byte is the final termination byte of the first parcel and the second parcel for processing.

7. the encoded video data comprises one or more syntax elements of a video bitstream; the one or more syntax elements indicate one or more parameters defining a neural network for decoding the encoded video data; The method of claim 1 , wherein the one or more parameters defining the neural network comprise at least one of weights of the neural network and an activation function of the neural network.

8. 1. An apparatus for processing video data, comprising: a memory configured to store video data; one or more processors coupled to the memory; wherein the processor: Obtaining encoded video data; determining a value intersection between a value for a first termination byte of a first parcel of the encoded video data and a value for a second termination byte of a second parcel of the encoded video data, wherein the value for the first termination byte comprises a first range of termination byte values ​​permitted for decoding and the value for the second termination byte comprises a second range of termination byte values ​​permitted for decoding, and the value intersection includes values ​​in the first range and the second range. determining a co-termination byte for the first termination byte of the first parcel and the second termination byte of the second parcel, wherein a value for the co-termination byte is based on the intersection of values; generating entropy coded data including the joint termination byte for the first parcel and the second parcel; An apparatus configured to:

9. The apparatus of claim 8 , wherein the one or more processors are configured to use arithmetic coding to generate the entropy-coded data.

10. the one or more processors are configured to use binary coding to generate the entropy coded data; the values ​​for the first termination byte include a first number of bits and the values ​​for the second termination byte include a second number of bits, and the intersection of values ​​includes at least one of a common value between the first number of bits and the second number of bits, and a subset of values ​​from the first number of bits and a subset of values ​​from the second number of bits; 9. The apparatus of claim 8, wherein an order of the first number of bits and an order of the second number of bits are unchanged in the co-termination byte compared to an order of the first number of bits in the first termination byte and an order of the second number of bits in the second termination byte.

11. to generate the entropy coded data, the one or more processors are configured to perform parallel entropy coding of the first parcel and the second parcel; a first encoder configured to encode the first parcel; a second encoder configured to encode the second parcel; The apparatus of claim 8 further comprising:

12. the one or more processors: performing parallel entropy decoding of the first parcel and the second parcel using the joint termination byte for the first parcel and the second parcel; configured to: the one or more processors: reading the first parcel in forward order; reading the second parcel in reverse order; The apparatus of claim 8 configured to:

13. the joint termination byte is the final termination byte of the first parcel and the second parcel for processing; and / or the encoded video data comprises one or more syntax elements of a video bitstream; The apparatus of claim 8 , wherein the one or more syntax elements indicate one or more parameters that define a neural network for decoding the encoded video data.

14. 14. The apparatus of claim 13, wherein the one or more parameters defining the neural network comprise at least one of weights of the neural network and an activation function of the neural network.

15. 10. The device of claim 8, wherein the device is one of a mobile device, an extended reality device, or a television, and the device further comprises at least one of a display and a camera configured to capture one or more video frames.

Citation Information

Patent Citations

  • Further improved method and apparatus for image compression

    JP2021520728A

  • A further improved method and apparatus for image compression

    WO2019191811A1