Multi-threaded CABAC Decoding

By separating the CABAC engine from the parser and using a finite state machine to determine context probabilities, the method enables parallel processing across multiple cores, addressing the inefficiencies in existing CABAC decoding methods and achieving improved performance for high-bitrate streams.

JP7696019B2Active Publication Date: 2025-06-19SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023575575
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-06-07
Filing Date
2022-04-25
Publication Date
2025-06-19
Estimated Expiration
2042-04-25

AI Technical Summary

Technical Problem

Existing CABAC decoding methods are inefficient due to data dependency at the bin level, which prevents parallel processing across multiple processor cores and computing units, leading to performance deterioration when decoding high-bitrate CABAC streams.

Method used

The implementation of a parallel CABAC decoding method that separates the CABAC engine from the parser, utilizing a finite state machine to determine context probabilities and allowing for independent operation of the CABAC engine and syntax analyzer, enabling parallel processing across multiple cores.

Benefits of technology

This approach enables real-time decoding of high-bitrate unconstrained CABAC streams by leveraging multiple CPU cores and GPUs, improving decoding performance without requiring additional encoder information or reducing coding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007696019000002
    Figure 0007696019000002
  • Figure 0007696019000003
    Figure 0007696019000003
  • Figure 0007696019000004
    Figure 0007696019000004
Patent Text Reader

Abstract

A method, system, and computer-readable medium are described for improved decoding of CABAC coded media. A decoding loop includes decoding a coded binary element from a sequence of coded binary elements using a context probability to generate a decoded binary element. A next context probability for a next coded binary element in the sequence is determined from the decoded binary element, and the next context probability for decoding the next coded binary element is provided to the decoding loop for a next iteration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Aspects of the present disclosure relate to decoding of encoded media. In particular, aspects of the present disclosure relate to decoding of improved context-adaptive binary arithmetic coding (CABAC).

Background Art

[0002] As shown in FIG. 1, prior art CABAC decoding is a tight loop. For each binary decision (bin) 101 to be decoded, the CABAC decoding engine 102 must obtain from the bin probability table 107 the probability that the bin is 0 or 1. As the name context-adaptive binary arithmetic coding indicates, the probability depends on the context 106 of this bin. Using a VP9 decoder as an example, if all DCT coefficients are represented by the same integer bit depth, without context, the probability of 0 or 1 at each bit position is the same. For example, all most significant bits of the same bit depth integers share the same probability. To resolve the true probability of the DCT coefficient, a context must be determined.

[0003] To detect (106) the correct context of the bin, the decoder must execute a parser 104 to parse all the bins 103 that have been decoded so far and determine the decoded symbol 105. For example, assume that the parser detects that the bit depth of the current coefficient is 5 (4 value bits and 1 sign bit) and that four bins have been decoded for this coefficient. In this case, when the parser determines that the 4 bits are the coefficient, it determines that the next bin represents the sign bit of this coefficient. By knowing this information, the CABAC engine can select the correct probability for decoding the sign bit bin. The CABAC engine always waits for the parser to complete the last bin parsing before decoding the next bin. This results in data dependency at the bin level.

[0004] Modern general-purpose processors have multiple processor cores, and each processor core has multiple computing units. The multiple processor cores can utilize thread-level parallelism, and the multiple computing units can utilize instruction-level parallelism. However, since the decoding of CABAC bins is a dense loop, when a single loop is executed in multiple threads, a large inter-thread communication delay occurs for each iteration of the loop. Since the CABAC algorithm forms data dependencies at the bin level, there is no instruction-level parallelism for multiple computing units in the CABAC algorithm. As a result, the CABAC decoding loop cannot utilize multiple processor cores and multiple computing units per core. Therefore, the performance of CABAC decoding using a general-purpose processor tends to deteriorate.

[0005] CABAC entropy coding is common in video compression standards such as AVC (H.264), HEVC (H.265), VVC (H.266), VP9, and AV1. Currently, only dedicated hardware decoders can decode unconstrained CABAC streams over 100 Mbps in real time. A single consumer-grade processor core is not fast enough to decode such a stream. The only viable approach to enable real-time decoding at high bitrates with existing general-purpose processors is parallel decoding. Unfortunately, there is no parallel decoding suitable for CABAC streams. In video coding standards, multiple coding tools or constraints have been introduced to enable parallel decoding such as multi-slice, multi-tile, and wavefront. However, these tools or constraints may slightly reduce the coding efficiency, and not all encoders support these coding tools or accept the constraints. Ideally, the decoder should not assume that all input streams have such tools or constraints.

[0006] In such a context, aspects of the present disclosure arise. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The teachings of the present disclosure can be readily understood by considering the following detailed description in conjunction with the accompanying drawings.

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

[0009] The following detailed description includes many specific details for purposes of illustration, but one of ordinary skill in the art will recognize that many variations and modifications to the following details are within the scope of the invention. Accordingly, the exemplary embodiments of the invention described below are shown without loss of generality to the claimed invention and without imposing limitations on the claimed invention.

[0010] Aspects of the present disclosure enable parallel CABAC decoding independent of parallel decoding tools and video standards with introduced constraints. The implementation of parallel CABAC can utilize the computing power of multiple CPU cores or CPU / GPU to decode high-bitrate unconstrained CABAC streams in real time. Existing parallelization solutions divide the task of decoding pictures at the slice, row, or tile level using information from the encoder defined by the encoder standard to determine the split positions for parallelization. Some video coding standards cannot be parallelized because they do not obtain information from the encoder for splitting pictures for parallel decoding. Also, the information for splitting pictures for parallel decoding requires additional data to be encoded with the pictures. Thus, when a frame is encoded with information for splitting pictures, the coding efficiency decreases. Aspects of the present disclosure provide a method for enabling parallel CABAC decoding in video standards lacking information for splitting pictures.

[0011] Before describing an improved method for parallelizing CABAC decoding, it is useful to understand how digital pictures, e.g., video pictures, are encoded / decoded for streaming and storage applications. In the context of aspects of the present disclosure, video picture data can be decomposed into units of a size suitable for encoding and decoding. For example, in the case of video data, this video data can be decomposed into a plurality of pictures, each picture corresponding to a particular picture in a sequence of images. Each unit of video data may be decomposed into sub-units of different sizes. Generally, within each unit, there are some minimum or basic sub-units. In the case of video data, each video frame can be decomposed into a plurality of pixels, each pixel including luma (luminance) data and chroma (color) data.

[0012] As an example and without limitation, as shown in FIG. 5, a single picture 500 (e.g., a digital video frame) can be decomposed into one or more sections. As used herein, the term "section" may refer to a group of one or more luma samples or chroma samples within picture 500. A section can range from a single luma sample or chroma sample within the picture to the entire picture. Non-limiting examples of sections include slices (e.g., macroblock rows) 502, macroblocks 504, sub-macroblocks 506, blocks 508, and individual pixels 510. As illustrated in FIG. 5, each slice 502 includes one or more rows of macroblocks 504, or a portion of one or more such rows. The number of macroblocks within a row is determined by the size of the macroblocks, as well as the size and resolution of picture 500. For example, if each macroblock includes 16×16 chroma samples or luma samples, the number of macroblocks within each row can be determined by dividing the width of picture 500 (in chroma sample units or luma sample units) by 16. Each macroblock 504 can be decomposed into several sub-macroblocks 506. Each sub-macroblock 506 can be decomposed into several blocks 508, and each block can include several chroma samples or luma samples 510. As an example and without limitation, in a typical video coding scheme, each macroblock 504 may be decomposed into four sub-macroblocks 506. Each sub-macroblock may be decomposed into four blocks 508, and each block may include a 4×4 array of 16 chroma samples or luma samples 510.

[0013] Some codecs, such as H.265, allow a given picture to be decomposed into two or more sections of different sizes for encoding. Specifically, the H.265 standard introduces the concept of "tiles" for partitioning pictures. A tile is an independently decodable region of a picture and is encoded with some shared header information. Tiles can additionally be used for the purpose of spatial random access to local regions of a video picture. A typical tile configuration of a picture is composed by segmenting the picture into rectangular regions having approximately the same number of coding units (CUs) in each tile. A coding unit is similar to the macroblock (MB) of the H.264 standard. However, the size of the CU can be set by the encoder and can be larger than the macroblock. The size of the CU can be flexibly and adaptively adjusted according to the video content for the best partitioning of the picture.

[0014] Note that each picture can be either a frame or a field. A frame refers to a complete image. A field is a part of an image that is used to facilitate the display of an image on a particular type of display device. Generally, chroma samples or luma samples within an image are arranged in rows. To facilitate the display, the image may be split by alternately dividing the rows of pixels into two different fields. In that case, the rows of chroma samples or luma samples in the two fields can be interlaced to form a complete image. In some display devices, such as a cathode ray tube (CRT) display, the two fields can simply be displayed successively in sequence. The afterglow of the phosphor or other light-emitting elements used to illuminate the pixels in the display, combined with the persistence of vision, results in the two fields being perceived as a continuous image. In certain display devices, such as a liquid crystal display, it may be necessary to interlace the two fields into a single picture before display. Streaming data representing the encoded image may contain information indicating whether the image is a field or a frame, or in some standards, such information may be absent. Such information may be included in the header for the image.

[0015] Modern video coders / decoders (codecs), such as MPEG2, MPEG4, and H.264, generally encode a video frame as one of three basic types, known as intra-frame, predictive frame, and bi-directional predictive frame, typically referred to as I-frame, P-frame, and B-frame, respectively.

[0016] An I-frame is a picture that is coded without referring to any pictures other than itself. I-frames are used for random access and as references for the decoding of other P-frames or B-frames. I-frames may be generated by the encoder to create a random access point (which enables the decoder to start decoding properly from zero at a given picture position). An I-frame may be generated when the ability to distinguish the details of an image prohibits the generation of valid P-frames or B-frames. Since I-frames contain complete pictures, I-frames typically require more bits to encode than P-frames or B-frames. Video frames are often encoded as I-frames when a scene change is detected in the input video.

[0017] A P-frame requires the pre-decoding of some other picture(s) in order to be decoded. P-frames typically require fewer bits to encode than I-frames. A P-frame contains coding information regarding the difference with respect to the previous I-frame in decoding order. A P-frame typically refers to the preceding I-frame within a group of pictures (GoP). A P-frame may contain both picture data and motion vector displacements, as well as combinations of the two. In some standard codecs (such as MPEG-2), a P-frame also requires using only one previously decoded picture as a reference during decoding and that picture must precede the P-frame in display order. In H.264, a P-frame can use multiple previously decoded pictures as references during decoding and can have any arbitrary display order relationship to the picture(s) used for its prediction.

[0018] A B-frame requires prior decoding of either an I-frame or a P-frame in order to be decoded. Similar to a P-frame, a B-frame may contain both image data and motion vector displacements, and / or a combination of the two. A B-frame may include several prediction modes that form a prediction of a motion area (e.g., a segment of a frame such as a macroblock or a smaller area) by averaging predictions obtained using two different previously decoded reference areas. In some codecs (such as MPEG-2), a B-frame is not used as a reference for the prediction of other pictures. As a result, a lower quality encoding (usually using fewer bits) may be used for such B-pictures since the loss of detail does not negatively affect the prediction quality of subsequent pictures. In other codecs such as H.264, a B-frame may or may not be used as a reference for the decoding of other pictures (depending on the encoder's decision). Some codecs (such as MPEG-2) require exactly two previously decoded pictures to be used as references during decoding, with one of those pictures preceding the B-frame picture in display order and the other following it. In other codecs such as H.264, a B-frame may use one, two, or more than two previously decoded pictures as references during decoding and can have any arbitrary display order relationship to the picture(s) used for its prediction. A B-frame typically requires fewer bits for encoding than either an I-frame or a P-frame.

[0019] As used herein, the terms I-frame, B-frame, and P-frame may be applied to any streaming data unit having characteristics similar to an I-frame, B-frame, and P-frame, as described above, for example, with respect to the context of streaming video.

[0020] To encode a digital video picture, the encoder receives a plurality of digital images and encodes each image. Encoding of the digital picture can be advanced section by section. As used herein, image compression refers to applying data compression to a digital image. The purpose of image compression is to reduce the redundancy of the image data of a given image in order to enable the storage or transmission of the data of that image in an efficient compressed data format.

[0021] Entropy encoding is a coding method that assigns codes to signals such that the code length matches the probability of the signal. Typically, an entropy encoder is used to compress data by replacing symbols represented by codes of equal length with symbols represented by codes proportional to the negative logarithm of the probability.

[0022] CABAC is a form of entropy encoding used in the H.264 / MPEG-4 AVC and High Efficiency Video Coding (HEVC) standards. CABAC is characterized by providing much better compression than most other entropy encoding algorithms used in video encoding, and is one of the main elements that provides better compression ability than previous ones in the H.264 / AVC encoding method. However, it should be noted that the fact that CABAC uses arithmetic coding may require a larger amount of processing for decoding than Context Adaptive Variable Length Coding (CAVLC).

[0023] Research has shown that the following operations are required to perform CABAC bin decoding. These operations are referred to as CABAC core operations. 1. Select a context for the next bin according to all the bins decoded previously. 2. Determine the probability of the next bin. 3. Decode the next bin.

[0024] The insight of this disclosure is that not all parser tasks belong to the CABAC core operations. For example, in the decoding of DCT coefficients, after decoding all four value bit bins and one sign bit bin, the decoder reconstructs the signed integer and stores it in the DCT coefficient matrix. The process of integer reconstruction and storage may not be a CABAC core operation and may be moved to a separate thread that is decoupled from the thread performing CABAC decoding. CABAC bin decoding is a very dense loop with few operations required for arithmetic calculations and table lookups. In contrast, the parser is significantly larger. Each type of syntax symbol has a dedicated code block for reconstruction.

[0025] When the syntax of the coding system is context-free at the superblock level or the macroblock level, it has been found that there is a many-to-one relationship between the state at that level and the CABAC bin context. In other words, the machine state can be used to look up the context of CABAC bins at the superblock level or the macroblock level. When symbol reconstruction is removed from the CABAC decoding thread, a finite state machine can be defined to decode all symbols with a context-free grammar. By removing the syntax reconstruction function from the decoding cycle, at least all the cycle-heavy syntax elements can be parsed by the finite state machine. The cycle-heavy syntax elements include decoded symbols such as DCT coefficients, motion vectors, and block prediction modes. There are some symbols that are loop-dependent and cannot be reconstructed using the state machine. Typically, these symbols are at the picture level and should be noted that they can be decoded using symbol-dependent context lookups used in the prior art.

[0026] As shown in FIG. 2, due to the separation between the CABAC engine 208 and the parser 209, a loop can be formed within the CABAC engine 208 using the decoded binary sequence (which may also be referred to as "binary quantization of syntax elements" or "bin stream") without completely reconstructing the symbols (205). The CABAC engine 208 receives the CABAC encoded input stream 201, includes a CABAC decoding engine 202 that generates the decoded bin stream 203, a context state machine 206 that advances the state when the current binary element (also referred to herein as a "bin") is decoded and provides a context to the context probability lookup table 207. The context probability table 207 is used with the state from the state machine 206 and the current decoded binary element 203 to provide the CABAC decoding engine 202 with the probability that the next binary element will be 1 or 0. As used herein, a binary element or bin is a single bit of a bit stream. A binary sequence or bin is a set of binary elements and is an intermediate representation of the value of a syntax element from the binary quantization of the syntax element.

[0027] By resolving the context probability from the decoded binary element 203 and the state of the finite state machine 206, the program size can be reduced, and furthermore, since there is no dependency on syntax analysis, parallel processing becomes possible. After extracting syntax reconstruction from the CABAC decoding thread, the thread can easily fit into the level 1 instruction cache of the CPU cores of an existing multi-core processor. The context probability table 207 can be small enough to be loaded into the data cache of the CPU. The CABAC engine 208 can operate independently of the syntax analyzer 209 and decode bins in a tight loop. This further enables various degrees of parallelism. Multiple instances of the syntax analyzer 209 can be executed on separate threads of the processor or separate cores of the processor, and the CABAC engine 208 can supply bins to a separate core executing the syntax analyzer 209. In some alternative implementations, the CABAC engine 208 can be implemented on a central processing unit (CPU), and the syntax analyzer 209 can be implemented on one or more graphics processing units (GPUs). In some other alternative implementations, the CABAC engine 208 may be implemented on the CPU core of the processor, and the syntax analyzer 209 may be implemented on the GPU core of the processor.

[0028] The syntax analyzer 209 may implement a syntax analysis operation 204 that converts the decoded bin sequence 203 from the CABAC engine 208 into decoded symbols 205. The syntax analyzer includes dedicated code blocks for each type of syntax symbol in the encoding standard that can be executed in parallel to decode each symbol individually, or to increase parallelism, the syntax analyzer 209 may decode symbols simultaneously at the superblock, macroblock, or block level. Separating the syntax analyzer 209 from the CABAC engine 208 enables parallelization of the syntax analysis operation. Such parallelization may include the implementation of parallel threads that can be executed on one core or multiple cores, or the computational units within a core.

[0029] Returning to the CABAC engine 208, FIG. 3 shows an exemplary implementation of a finite state machine according to an aspect of the present disclosure. As shown, the finite state machine obtains an initial state 301 and advances the initial state to the current state. The system uses the current state of the finite state machine having the last decoded bin value (when the previous bin was decoded) to determine the next state (306). The next state is used to determine the context probability of the next bin using a look-up table 302. In some implementations, the context state machine 206 can perform two separate look-up operations, and the context state machine can include an internal look-up table used to look up the context from its own state. Using the current context obtained from the internal look-up table of the state machine 206, the state machine can then look up the bin probability from the context probability look-up table 207. The bin probability can be determined from the context using any suitable method. By way of example, and not limitation, one method is to define the bin probability as constant for a given context, and thus each context has a probability within the array. As an alternative example, and not limitation, another method is to have an array of probabilities indexed by the context. The CABAC engine may update the entries of the array according to the decoding results of the previous bin having the same context.

[0030] When the context probability is obtained, it is possible to supply the context probability to the CABAC decoding engine 202 (307), and the CABAC decoding engine 202 can use the context probability to determine the next bin value. The input to the CABAC decoding engine 202 is the input CABAC stream 201 and the context probability determined using the look-up table 302. After providing the context probability, the state machine performs a check (303) to determine whether the next state is the end state. If the next state is not the end state (No), the state machine changes the next state to the current state (304), obtains the bin value 306 and the current state 304, and continues the loop by, for example, using the look-up table 302 to determine the next context probability.

[0031] If the next state is the end state (Yes), the finite state enters the end program state 305. The CABAC decoding loop outputs the decoded bit sequence from all the previously decoded binary elements. For example, but not limited to, the CABAC decoding loop can output a decoded binary sequence having two or more previously decoded binary elements. The end program state 305 can include additional entries indicating the end of a block, macroblock, or superblock. The end of the state entry can include several binary numbers indicating the type of symbol and direct the decoded bit sequence to an appropriate parser. In some implementations, there are multiple parsers operating in parallel, and the end of the state entry can be used to direct the decoded bit sequence to the appropriate parser that matches the symbols to the corresponding decoded bit sequence. Alternatively, the parser can use the end flag of the entry to select the appropriate parsing operation.

[0032] Accordingly, a method for improved decoding of CABAC encoded media can include a decoding loop that can use context probabilities to decode encoded binary elements from a sequence of encoded binary elements to generate decoded binary elements. The next context probability is determined for the next encoded binary element in the sequence from the decoded binary elements and provided to the decoding loop for the next iteration. Determining the next context probability for the next encoded binary element can include advancing the state of a finite state machine configured to provide a context for determining the next context probability. Instructions for the finite state machine can be stored in the instruction cache of a processor. Further, a lookup table can be used by looking up the next context probability in the lookup table using the decoded binary elements and the state of the finite state machine. The lookup table may be stored in the data cache of the processor. In some implementations, the decoded binary sequence may be generated from a sequence of two or more previously decoded binary elements at the end state of the decoding loop. It should be understood that the two or more previously decoded binary elements can be all of the binary elements decoded from the encoded binary sequence. The decoding loop may be processed in a first processing thread, and parsing the syntax of the decoded binary sequence to generate decoded symbols from the parsed syntax may be performed in a second processing thread. In other implementations, the decoding loop may be processed on a first processor core, and parsing the syntax of the decoded binary sequence to generate decoded symbols from the parsed syntax may be executed on a second processor core. In yet other implementations, the decoding loop may be processed by a processor, and parsing the syntax of the decoded binary sequence to generate decoded symbols from the parsed syntax may be executed by a graphics processing unit. An improved method for decoding CABAC encoded media may include a plurality of parsers operating in parallel to parse the syntax of the decoded binary sequence and generate decoded symbols from the parsed syntax.An improved method for decoding CABAC symbolized media may further include decoding encoded discrete cosine transform coefficients, motion vectors, or binary elements for block prediction modes.

[0033] FIG. 4 shows an implementation of a finite state machine shown as a binary decision tree ending in a final symbol value determination. Each node in this binary decision tree represents a state of the state machine. The first state 401 has two binary choices that may have associated probabilities. If the bin value of the first state 401 is zero, the state machine proceeds to the end state 402 with symbol value 0. If the bin value of the first state 401 is 1, the state machine proceeds to the second state 403 to determine the second bin value. Similar to the first state, the second state also has two choices with associated context probabilities. If it is determined that the binary value of the second state 403 is 1, the state machine proceeds to the end state 404 with symbol value 4. If the value is determined to be 0, the state machine proceeds to the third state 405 to determine the third bin value. As described above, context probabilities can be used to determine state values. If the value is 1, the state proceeds to an end state and the symbol value is determined to be 3 (406). If the value is determined to be 0, the state proceeds to the fourth state 407 and the binary value is determined as before. Here, the fourth state is an end state, and thus both bin values provide an end symbol value. If the binary value is 0, the symbol value is determined to be 1 (408), and if the binary value is 1, the end symbol value is 2 (409). The states represent bin values that can be combined to reach a final value. For example, for a symbol value of 3, a binary number of 101 is generated by processing through the state tree. [Table 1] Table 1 above shows an example of a state machine as shown in FIG. 4 as a look-up table. The first column of Table 1 represents the current state of the state machine. The state tracks the context without the need to fully decode the symbols. Once the binary value of the current state is determined, the table can be used to determine the next state and the probability of the next state. For example, if the first binary value is 1, the state machine advances to the second bin shown in the next state column, and the probability that the next bin is 0 is provided in the final column. This probability is provided to the CABAC decoding engine to determine the value bin used in the table. Thus, the state machine can include a table similar to that shown in Table 1 to assist in context probability determination.

[0034] The context-dependent nature of the decoded symbols, which means that the context probabilities differ between symbols, may require a separate context table, such as Table 1, and a separate state machine for each symbol to be decoded. In some implementations, the number of tables can be reduced by replacing only some of the data in the table for a particular symbol. For example, without limitation, the context probabilities of some symbols may be approximately the same, and their tables may differ by only a few entries. Instead of loading the entire new table into the data cache to process the approximately same tables, only the entries with different symbols can be changed in the already loaded table. Thus, the cycles required to flush the data buffer and write the new table can be reduced.

[0035] FIG. 6 illustrates a block diagram of a computer system 600 that may be used to implement video coding according to an aspect of the present disclosure. The system 600 may generally include a main processor 603, a memory 604, and a graphics processing unit (GPU) 626 that communicate via a main bus 605. The processor 603 may include one or more processor cores, such as a single core, dual core, quad core, processor-coprocessor, cell processor, architecture, etc. Each core may include one or more processing threads that can process instructions in parallel with each other. The processor may include one or more integrated graphics cores that can function as a GPU for the system.

[0036] In some implementations of the present disclosure, the processor 603 can execute a CABAC engine 623 and a syntax analyzer 624. The CABAC engine 623 may include state machine instructions that are small enough in data size to fit within the instruction cache of the processor 603. The CABAC engine 623 may also include one or more context tables to convert CABAC-decoded binary syntax elements and states from the state machine into context probabilities for the next encoded binary syntax elements. For example, without limitation, the first of two lookup operations for determining context probabilities determines a context from the state of the machine state, and the second lookup determines a bin context probability from the context. The context table may be small enough in data size to fit within a data cache such as, for example, a level 1 data cache, a level 2 cache, or a level 3 cache. The syntax analyzer 624 can operate on a separate thread from the CABAC engine 623. For example, the CABAC engine can be executed on a first thread, and the syntax analyzer 624 can be executed in parallel on a second thread. Alternatively, the syntax analyzer 624 can be executed on a separate core of the processor 603 or a GPU core. The syntax analyzer can parse the syntax or grammar of the decoded binary sequence and generate decoded syntax elements or symbols. In some implementations, the syntax analyzer 627 can be loadable from memory to a graphics processing unit (GPU) 626. The syntax analyzer 627 may receive the decoded binary sequence from the CABAC engine 623 executing on the processor 603.

[0037] Memory 604 may be in the form of an integrated circuit, such as, for example, RAM, DRAM, and ROM. The memory may also be a main memory accessible by all of the processor cores within processor 603. In some embodiments, processor 603 may have local memory associated with one or more processor cores or one or more coprocessors. Decoder program 622 may be stored in memory 604 in the form of processor-readable instructions executable on processor 603. Decoder program 622 may be configured to decode CABAC-encoded signal data into a decoded picture, as described above, for example. Decoder program 622 may coordinate the operations of CABAC engine 623 and syntax analyzer 624. CABAC engine 623 can obtain CABAC-encoded binary elements and generate decoded binary elements. CABAC engine 623 creates two or more decoded binary elements, i.e., a decoded binary sequence parsed into one symbol or multiple symbols by the syntax analyzer. CABAC engine 623 may include a state machine loaded from memory 604 into the instruction cache of processor 603. The state machine may be one of a number of state machines 610 stored in memory 604 until the appropriate symbol of the state machine is decoded, at which point the instructions of the appropriate state machine are loaded into the instruction cache of processor 603. The CABAC engine may also include a context table loaded from memory 604 into the data cache of processor 603. The context table may be one of a number of context tables 621 stored in memory 604. Each symbol may have an associated context table for decoding the encoded binary syntax elements associated with that symbol using the state from the state machine and the currently decoded binary syntax elements, as described above. Memory 604 may also include a syntax analyzer program 609 that converts the decoded binary sequence into decoded symbols.The syntax parser program 609 may be executed by the processor 603, and at least a part of the syntax parser 609 may be loaded from the memory 604 into the instruction cache and / or data cache of the processor 603. In an implementation having a syntax parser executed on the GPU 627, the syntax parser 627 can receive the decoded binary syntax elements from the CABAC engine 623 that is executed on the processor 603 or stored in the buffer 608 within the memory 604. The buffer 608 can store the encoded data or other data generated or received during the decoding process in the memory 604.

[0038] System 600 may also include well-known support functions 606 such as input / output (I / O) element 607, power supply (P / S) 611, clock (CLK) 612, and cache 613. System 600 may optionally include a mass storage device 615 such as a disk drive, CD-ROM drive, tape drive, or the like for storing program 617 and / or data 618. Decoder program 622 and parser 609 may be stored in mass storage device 615 as program 617. Context table 621, state machine 610, and buffered data may also be stored in mass storage device 615 as data 618. Device 600 may also optionally include a user interface 616 and user input device 602 for facilitating interaction between system 600 and the user. User interface 616 may take the form of a cathode ray tube (CRT) or flat panel screen for displaying text, numbers, graphic symbols, or images. User input device 602 may include a keyboard, mouse, joystick, light pen, or other device that can be used with a graphical user interface (GUI). System 600 may also include a network interface 614 for enabling the device to communicate with other devices through a network 620 such as the Internet. System 600 can receive one or more frames of encoded streaming data (e.g., one or more encoded video frames) from other devices connected to network 620 via network interface 614. These components may be implemented in hardware, software, firmware, or some combination of two or more of these.

[0039] The foregoing is a complete description of preferred embodiments of the present invention, but various alternatives, modifications, and equivalents may be used. Accordingly, the scope of the present invention should not be determined with reference to the above description, but instead should be determined with reference to the appended claims, which follow the full scope of such equivalents. Whether or not preferred, any feature described herein may be combined with any other feature described herein, whether or not preferred. In the following claims, the indefinite article "A" or "An" refers to one or more of the items following the article, unless expressly stated otherwise. The appended claims should not be construed as including means-plus-function limitations, unless such a limitation is expressly recited in a given claim using the phrase "means for".

Claims

1. A method for decoding a CABAC-coded medium, using context probabilities to decode the encoded binary elements from a sequence of encoded binary elements to generate decoded binary elements; determining, from the decoded binary elements, the next context probability for the next encoded binary element in the sequence; providing the next context probability to a next iteration of the decoding loop for decoding the next encoded binary element; comprising the decoding loop being processed in a first processing thread, at an end state of the decoding loop, outputting a decoded binary sequence from two or more previously decoded elements, and parsing the syntax of the decoded binary sequence to generate decoded symbols from the parsed syntax in a second processing thread, further comprising the decoding loop.

2. Determining the next context probability for the next encoded binary element includes advancing the state of a finite state machine configured to provide a context for determining the next context probability, according to the method of claim 1.

3. The decoding loop further includes looking up the next context probability in a lookup table using the decoded binary elements and the state of the finite state machine, according to the method of claim 2.

4. The lookup table is stored in a data cache of a processor, according to the method of claim 3.

5. The decoding loop is processed on a first processor core, according to the method of claim 1.

6. At the end state of the decoding loop, output a decoded binary sequence from two or more previously decoded elements, and parse the syntax of the decoded binary sequence, and generate a decoded symbol from the parsed syntax by a second processor core. The method according to claim 5, further comprising:

7. At the end state of the decoding loop, output a decoded binary sequence from two or more previously decoded elements, and parse the syntax of the decoded binary sequence, and generate a decoded symbol from the parsed syntax by a graphics processing unit. The method according to claim 5, further comprising:

8. The method according to claim 1, further comprising decoding a binary sequence for an encoded discrete cosine transform coefficient, a motion vector, or a block prediction mode.

9. The method according to claim 1, further comprising executing a plurality of syntax analyzers in parallel.

10. The instructions for the finite state machine are stored in the instruction cache of the processor. The method according to claim 2.

11. A system for decoding a CABAC encoded medium, comprising: A processor; A memory coupled to the processor; Processor-executable instructions embodied in the memory, the instructions comprising: Using context probability, decode an encoded binary element from a sequence of encoded binary elements to generate a decoded binary element; Determine a next context probability for a next encoded binary element in the sequence from the decoded binary element; Provide the next context probability to a next iteration of the decoding loop to decode the next encoded binary element; Including The decoding loop is processed in a first processing thread of the processor, and the decoding loop further includes outputting a decoded binary sequence from two or more previously decoded elements in an end state of the decoding loop, and is configured to implement a method for decoding a CABAC-coded medium including the decoding loop, the instruction, and a syntax analyzer configured to parse the syntax of the decoded binary sequence and generate decoded symbols from the parsed syntax in a second processing thread of the processor; A system comprising.

12. Determining the next context probability for the next encoded binary element includes advancing the state of a finite state machine configured to provide a context for determining the next context probability, the system of claim 11.

13. The decoding loop further includes looking up the next context probability in a lookup table using the decoded binary element and the state of the finite state machine, the system of claim 12.

14. The lookup table is stored in a data cache of the processor, the system of claim 13.

15. The decoding loop is processed in a first processor core of the processor, and the decoding loop further includes outputting a decoded binary sequence from two or more previously decoded elements in an end state of the decoding loop, the system of claim 11.

16. The system of claim 15, further comprising a syntax analyzer configured to parse the syntax of the decoded binary sequence and generate decoded symbols from the parsed syntax in a second processor core of the processor.

17. The graphic processing apparatus further includes a syntax analyzer configured to parse the syntax of the decoded binary sequence and generate a decoded symbol from the parsed syntax in the graphic processing apparatus, and the decoding loop further includes outputting a decoded binary sequence from two or more previously decoded elements in an end state of the decoding loop. The system according to claim 11.

18. The decoding loop further includes outputting a decoded binary sequence from two or more previously decoded elements in an end state of the decoding loop, and the method further includes decoding a binary sequence for an encoded discrete cosine transform coefficient, a motion vector, or a block prediction mode. The system according to claim 11.

19. The system according to claim 11, further comprising a plurality of syntax analyzers executed in parallel.

20. The system according to claim 12, wherein the instructions of the finite state machine are stored in an instruction cache of the processor.

21. A non-transitory computer-readable medium in which computer-readable instructions are embodied, the instructions being configured to implement a method for decoding a CABAC encoded medium when executed by a computer, the method comprising: using context probabilities to decode an encoded sequence of binary elements to generate decoded binary elements; determining a next context probability for a next encoded binary element in the sequence from the decoded binary elements; providing the next context probability to a next iteration of the decoding loop for decoding the next encoded binary element; including the decoding loop is processed in a first processing thread, At the end state of the decoding loop, output a decoded binary sequence from two or more elements decoded previously, and perform syntax analysis on the syntax of the decoded binary sequence, and generate a decoded symbol from the syntax analyzed by a second processing thread, further including a non-transitory computer-readable medium including the decoding loop.

Citation Information

Patent Citations

  • Coefficient coding in video encoding

    JP2015508617A

  • advanced arithmetic coder

    JP2018521556A