A method, an apparatus and a computer program product for video encoding and video decoding
By using bin values to update probability state parameters across coding tree units, the solution addresses inefficiencies in updating arithmetic coding contexts, resulting in enhanced compression efficiency and video quality.
Patent Information
- Application Number
- PCT/EP2024/077322
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-16
- Filing Date
- 2024-09-27
- Publication Date
- 2025-05-22
Smart Images

Figure EP2024077322_22052025_PF_FP_ABST
Abstract
Description
[0001] A METHOD, AN APPARATUS AND A COMPUTER PROGRAM PRODUCT FOR VIDEO ENCODING AND VIDEO DECODING
[0002] Technical Field
[0003] The present solution generally relates to video encoding and video decoding.
[0004] Background
[0005] Video encoding is a process, where input video is transformed into a compressed format suited for storage or transmission. In video decoding, the opposite is performed, i.e., compressed video is uncompressed back into a viewable form. The encoding process comprises prediction, where pixel values of a certain picture area are predicted. Then a prediction error, i.e., difference between the predicted pixels and the original pixels is coded.
[0006] Summary
[0007] The embodiments discussed in the present description provides an improved prediction solution to be used in video encoding and decoding.
[0008] The scope of protection sought for various embodiments of the invention is set out by the independent claims. The embodiments and features, if any, described in this specification that do not fall under the scope of the independent claims are to be interpreted as examples useful for understanding various embodiments of the invention.
[0009] Various aspects include a method, an apparatus and a computer readable medium comprising a computer program stored therein, which are characterized by what is stated in the independent claims. Various embodiments are disclosed in the dependent claims.
[0010] According to a first aspect, there is provided an apparatus comprising i) means for determining a first coding tree unit to be encoded or decoded; ii) means for determining a coding unit belonging to the first coding tree unit; Hi) means for determining a bin value used by an arithmetic encoder or an arithmetic decoder to represent a syntax element of the coding unit; iv) means for determining if the bin value is going to be used to perform a location-based context probability update for a second coding tree unit, where the determining includes: a) determining a context class of the syntax element; b) determining if the context class belongs to a set of context classes selected to be used for the probability update for a second coding tree unit; c) determining if the number of bins with the same context class as the context class of the bin which are already included for the context probability update is below a selected maximum number; v) means for using the bin value to update a probability state parameter used in the second coding tree unit depending on the result of the means for determining on step iv).
[0011] According to a second aspect, there is provided a method, comprising i) determining a first coding tree unit to be encoded or decoded; ii) determining a coding unit belonging to the first coding tree unit;
[0012] Hi) determining a bin value used by an arithmetic encoder or an arithmetic decoder to represent a syntax element of the coding unit; iv) determining if the bin value is going to be used to perform a location-based context probability update for a second coding tree unit, where the determining includes: a) determining a context class of the syntax element; b) determining if the context class belongs to a set of context classes selected to be used for the probability update for a second coding tree unit; c) determining if the number of bins with the same context class as the context class of the bin which are already included for the context probability update is below a selected maximum number; v) using the bin value to update a probability state parameter used in the second coding tree unit depending on the result of the means for determining on step iv).
[0013] According to a third aspect, there is provided an apparatus comprising at least one processor, memory including computer program code, the memory and the computer program code configured to, with the at least one processor, cause the apparatus to perform at least the following: i) determine a first coding tree unit to be encoded or decoded; ii) determine a coding unit belonging to the first coding tree unit; Hi) determine a bin value used by an arithmetic encoder or an arithmetic decoder to represent a syntax element of the coding unit; iv) determine if the bin value is going to be used to perform a location-based context probability update for a second coding tree unit, where the determining includes: a) determining a context class of the syntax element; b) determining if the context class belongs to a set of context classes selected to be used for the probability update for a second coding tree unit; c) determining if the number of bins with the same context class as the context class of the bin which are already included for the context probability update is below a selected maximum number; v) use the bin value to update a probability state parameter used in the second coding tree unit depending on the result of the means for determining on step iv).
[0014] According to a fourth aspect, there is provided computer program product comprising computer program code configured to, when executed on at least one processor, cause an apparatus or a system to: i) determine a first coding tree unit to be encoded or decoded; ii) determine a coding unit belonging to the first coding tree unit;
[0015] Hi) determine a bin value used by an arithmetic encoder or an arithmetic decoder to represent a syntax element of the coding unit; iv) determine if the bin value is going to be used to perform a location-based context probability update for a second coding tree unit, where the determining includes: d) determining a context class of the syntax element; e) determining if the context class belongs to a set of context classes selected to be used for the probability update for a second coding tree unit; f) determining if the number of bins with the same context class as the context class of the bin which are already included for the context probability update is below a selected maximum number; v) use the bin value to update a probability state parameter used in the second coding tree unit depending on the result of the means for determining on step iv).
[0016] According to an embodiment, determining if the context class belongs to the set of context classes selected to be used for the probability update for the second coding tree includes determining a maximum amount of context classes considered for the update and determining if the number of included context classes for the first coding tree unit is below the maximum amount.
[0017] According to an embodiment, bin values of the first coding tree unit selected to be used for the probability update for the second coding tree unit are stored in a memory unit with an identifier representing the context class of the bin.
[0018] According to an embodiment, bin values of the first coding tree unit selected to be used for the probability update for the second coding tree unit are stored in a memory unit with a counter representing the number of bins stored for the context class of the bin.
[0019] According to an embodiment, number of the set of context classes to be used for the probability update for the second coding tree unit is limited to a predetermined number.
[0020] According to an embodiment, the pre-determined number is determined by using a set of given rules or from a video or image bitstream. According to an embodiment, binary shift operations are performed for the update.
[0021] According to an embodiment, the computer program product is embodied on a non-transitory computer readable medium.
[0022] Description of the Drawings
[0023] In the following, various embodiments will be described in more detail with reference to the appended drawings, in which
[0024] Fig. 1 shows an example of an encoding process;
[0025] Fig. 2 shows an example of a decoding process;
[0026] Fig. 3 is a flowchart illustrating a method according to an embodiment; and
[0027] Fig. 4 shows an apparatus according to an embodiment.
[0028] Description of Example Embodiments
[0029] The following description and drawings are illustrative and are not to be construed as unnecessarily limiting. The specific details are provided for a thorough understanding of the disclosure. However, in certain instances, well- known or conventional details are not described in order to avoid obscuring the description. Reference in this specification to “one embodiment” or “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the disclosure.
[0030] The present embodiments relate to encoding and decoding of digital video material. In the following, several embodiments will be described in the context of one video coding arrangement. It is to be noted, however, that the present embodiments are not necessarily limited to this particular arrangement. In particular, the present embodiments relate to tuning of arithmetic coding contexts with bin grouping. The arithmetic coding is a process where syntax elements are compressed or decompressed based on probability estimates for the syntax elements.
[0031] Before describing the embodiments further, a brief reference to evolution of video coding standardization is given. The present embodiments are suited within the context of next generation video coding standardization, e.g., H.267 video coding standard and the ECM (Enhanced Compression Model) exploration.
[0032] A video codec comprises an encoder and a decoder. The encoder transforms the input video into a compressed representation suited for storage / transmission. The decoder can un-compress the compressed video representation back into a viewable form. The encoder may discard some information in the original video sequence in order to represent the video in a more compact form (i.e. at lower bitrate).
[0033] Figure 1 shows an example of an encoding process for two-dimensional (2D) pictures, and Figure 2 shows an example of a decoding process for 2D pictures. In Figure 1 , the following have been illustrated:
[0034] - an image to be encoded (ln);
[0035] - a predicted representation of an image block (P'n);
[0036] - a prediction error signal (Dn);
[0037] - a reconstructed prediction error signal (D'n);
[0038] - a preliminary reconstructed image (l'n);
[0039] - a final reconstructed image (R'n);
[0040] - a transform (T) and inverse transform (T-1);
[0041] - a quantization (Q) and inverse quantization (Q-1);
[0042] - entropy encoding (E);
[0043] - a reference frame memory (RFM);
[0044] - inter prediction (Pinter);
[0045] - intra prediction (Pintra);
[0046] - mode selection (MS), and
[0047] - filtering (F).
[0048] In Figure 2 the following have been illustrated:
[0049] - a predicted representation of an image block (P'n);
[0050] - a reconstructed prediction error signal (D'n); - a preliminary reconstructed image (I'n);
[0051] - a final reconstructed image (R'n); an inverse transform (T-1 );
[0052] - an inverse quantization (Q-1 );
[0053] - an entropy decoding (E-1 );
[0054] - a reference frame memory (RFM);
[0055] - a prediction (either inter or intra) (P);
[0056] - and filtering (F).
[0057] An elementary unit for the input to an encoder and the output of a decoder, respectively, in most cases is a picture (also referred to as “an image”). A picture given as an input to an encoder may also be referred to as a source picture, and a picture decoded by a decoded may be referred to as a decoded picture or a reconstructed picture.
[0058] The source and decoded pictures are each comprised of one or more sample arrays, such as one of the following sets of sample arrays:
[0059] - Luma (Y) only (monochrome).
[0060] - Luma and two chroma (YCbCr or YCgCo).
[0061] - Green, Blue and Red (GBR, also known as RGB).
[0062] - Arrays representing other unspecified monochrome or tristimulus color samplings (for example, YZX, also known as XYZ).
[0063] A picture may be defined to be either a frame or a field. A frame comprises a matrix of luma samples and possibly the corresponding chroma samples. A field is a set of alternate sample rows of a frame and may be used as encoder input, when the source signal is interlaced. Chroma sample arrays may be absent (and hence monochrome sampling may be in use) or chroma sample arrays may be subsampled when compared to luma sample arrays.
[0064] The Advanced Video Coding standard (which may be abbreviated AVC or H.264 / AVC) was developed by the Joint Video Team (JVT) of the Video Coding Experts Group (VCEG) of the Telecommunications Standardization Sector of International Telecommunication Union (ITU-T) and the Moving Picture Experts Group (MPEG) of International Organization for Standardization (ISO) I International Electrotechnical Commission (IEC). The H.264 / AVC standard is published by both parent standardization organizations, and it is referred to as ITU-T Recommendation H.264 and ISO / IEC International Standard 14496-10, also known as MPEG-4 Part 10 Advanced Video Coding (AVC). There have been multiple versions of the H.264 / AVC standard, each integrating new extensions or features to the specification. These extensions include Scalable Video Coding (SVC) and Multiview Video Coding (MVC).
[0065] The High Efficiency Video Coding standard (which may be abbreviated HEVC or H.265 / HEVC) was developed by the Joint Collaborative Team - Video Coding (JCT-VC) of VCEG and MPEG. The standard is published by both parent standardization organizations, and it is referred to as ITU-T Recommendation H.265 and ISO / IEC International Standard 23008-2, also known as MPEG-H Part 2 High Efficiency Video Coding (HEVC). Extensions to H.265 / HEVC include scalable, multiview, three-dimensional, and fidelity range extensions, which may be referred to as SHVC, MV-HEVC, 3D-HEVC, and REXT, respectively. The references in this description to H.265 / HEVC, SHVC, MV-HEVC, 3D-HEVC and REXT that have been made for the purpose of understanding definitions, structures or concepts of these standard specifications are to be understood to be references to the latest versions of these standards that were available before the date of this application, unless otherwise indicated.
[0066] Versatile Video Coding (which may be abbreviated WC, H.266, or H.266 / VVC) is a video compression standard developed as the successor to HEVC. WC is specified in ITU-T Recommendation H.266 and equivalently in ISO / IEC 23090-3, which is also referred to as MPEG-I Part 3.
[0067] Some key definitions, bitstream and coding structures, and concepts of H.264 / AVC, HEVC, WC, and / or AV1 and some of their extensions are described in this section as an example of a video encoder, decoder, encoding method, decoding method, and a bitstream structure, wherein the embodiments may be implemented. The aspects of various embodiments are not limited to H.264 / AVC, HEVC, WC, and / or AV1 or their extensions, but rather the description is given for one possible basis on top of which the present embodiments may be partly or fully realized. Hybrid video codecs, for example ITU-T H.263, H.264 / AVC, HEVC, and WC, may encode the video information in two phases. At first, pixel values in a certain picture area (or “block”) are predicted for example by motion compensation means or by spatial means.
[0068] In motion compensation based prediction (which may be referred to as inter prediction, temporal prediction or motion-compensated temporal prediction or motion-compensated prediction or MCP) an area in one of the previously coded frames that corresponds closely to the block being coded is found and used for prediction. Inter prediction may reduce temporal redundancy.
[0069] In spatial prediction pixel values around the block to be coded are used. In the first phase, predictive coding may be applied, for example, as so-called sample prediction and / or so-called syntax prediction. In the sample prediction, pixel or sample values in a certain picture area or "block" are predicted. These pixel or sample values can be predicted, for example, using one or more of motion compensation or intra prediction mechanisms.
[0070] Intra prediction, where pixel or sample values can be predicted by spatial mechanisms, involve finding and indicating a spatial region relationship. Intra prediction utilizes the fact that adjacent pixels within the same picture are likely to be correlated. Intra prediction can be performed in spatial or transform domain, i.e., either sample values or transform coefficients can be predicted. Intra prediction is typically exploited in intra coding, where no inter prediction is applied.
[0071] In the syntax prediction, which may also be referred to as parameter prediction, syntax elements and / or syntax element values and / or variables derived from syntax elements are predicted from syntax elements (de)coded earlier and / or variables derived earlier. Non-limiting examples of syntax prediction are provided below.
[0072] In motion vector prediction, motion vectors e.g., for inter and / or inter-view prediction may be coded differentially with respect to a block-specific predicted motion vector. In many video codecs, the predicted motion vectors are created in a predefined way, for example by calculating the median of the encoded or decoded motion vectors of the adjacent blocks. Another way to create motion vector predictions, sometimes referred to as advanced motion vector prediction (AMVP), is to generate a list of candidate predictions from adjacent blocks and / or co-located blocks in temporal reference pictures and signalling the chosen candidate as the motion vector predictor. In addition to predicting the motion vector values, the reference index of previously coded / decoded picture can be predicted. The reference index is typically predicted from adjacent blocks and / or co-located blocks in temporal reference picture. Differential coding of motion vectors is typically disabled across slice boundaries.
[0073] The block partitioning, e.g., from a coding tree unit (CTU) to coding units (CUs) and down to prediction units (PUs), may be predicted.
[0074] In filter parameter prediction, the filtering parameters e.g., for sample adaptive offset may be predicted. Prediction approaches using image information from a previously coded image can also be called as inter prediction methods which may also be referred to as temporal prediction and motion compensation. Prediction approaches using image information within the same image can also be called as intra prediction methods.
[0075] In the second phase of encoding, the prediction error, i.e., the difference between the predicted block of pixels and the original block of pixels, is coded. This may be done by transforming the difference in pixel values using a specified transform (e.g., Discrete Cosine Transform (DCT) or a variant of it), quantizing the coefficients and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, encoder can control the balance between the accuracy of the pixel representation (picture quality) and size of the resulting coded video representation (file size of transmission bitrate).
[0076] In some video codecs, such as H.265 / HEVC, video pictures are divided into coding units (CU) covering the area of the picture. A CU consists of one or more prediction units (Pll) defining the prediction process for the samples within the CU and one or more transform units (TU) defining the prediction error coding process for the samples in the said CU. A CU may consist of a square block of samples with a size selectable from a predefined set of possible CU sizes. A CU with the maximum allowed size is typically named as LCU (largest coding unit) or CTU (coding tree unit) and the video picture is divided into non-overlapping CTUs. A CTU can be further split into a combination of smaller CUs, e.g. by recursively splitting the CTU and resultant CUs. Each resulting CU typically has at least one PU and at least one TU associated with it. Each PU and TU can be further split into smaller PUs and TUs in order to increase granularity of the prediction and prediction error coding processes, respectively. Each PU has prediction information associated with it defining what kind of a prediction is to be applied for the pixels within that PU (e.g. motion vector information for inter predicted PUs and intra prediction directionality information for intra predicted PUs). Similarly, each TU is associated with information describing the prediction error decoding process for the samples within the said TU (including e.g. DCT coefficient information). It is typically signaled at CU level whether prediction error coding is applied or not for each CU. In the case there is no prediction error residual associated with the CU, it can be considered there are no TUs for the said CU. The division of the image into CUs, and division of CUs into PUs and TUs is typically signaled in the bitstream allowing the decoder to reproduce the intended structure of these units.
[0077] The decoder reconstructs the output video by applying prediction means similar to the encoder to form a predicted representation of the pixel blocks (using the motion or spatial information created by the encoder and stored in the compressed representation) and prediction error decoding (inverse operation of the prediction error coding recovering the quantized prediction error signal in spatial pixel domain). After applying prediction and prediction error decoding means the decoder sums up the prediction and prediction error signals (pixel values) to form the output video frame. The decoder (and encoder) can also apply additional filtering means to improve the quality of the output video before passing it for display and / or storing it as prediction reference for the forthcoming frames in the video sequence.
[0078] Instead, or in addition to approaches utilizing sample value prediction and transform coding for indicating the coded sample values, a color palette based coding can be used. Palette based coding refers to a family of approaches for which a palette, i.e. , a set of colors and associated indexes, is defined and the value for each sample within a coding unit is expressed by indicating its index in the palette. Palette based coding can achieve good coding efficiency in coding units with a relatively small number of colors (such as image areas which are representing computer screen content, like text or simple graphics). In order to improve the coding efficiency of palette coding different kinds of palette index prediction approaches can be utilized, or the palette indexes can be run-length coded to be able to represent larger homogenous image areas efficiently. Also, in the case the CU contains sample values that are not recurring within the CU, escape coding can be utilized. Escape coded samples are transmitted without referring to any of the palette indexes. Instead their values are indicated individually for each escape coded sample.
[0079] In many video codecs, including H.264 / AVC, HEVC, and WC, motion information is indicated by motion vectors associated with each motion compensated image block. Each of these motion vectors represents the displacement of the image block in the picture to be coded (in the encoder) or decoded (at the decoder) and the prediction source block in one of the previously coded or decoded images (or pictures). In order to represent motion vectors efficiently those are typically coded differentially with respect to block specific predicted motion vectors. In typical video codecs the predicted motion vectors are created in a predefined way, for example calculating the median of the encoded or decoded motion vectors of the adjacent blocks. Another way to create motion vector predictions is to generate a list of candidate predictions from adjacent blocks and / or co-located blocks in temporal reference pictures and signaling the chosen candidate as the motion vector predictor. In addition to predicting the motion vector values, the reference index of previously coded / decoded picture can be predicted. The reference index is typically predicted from adjacent blocks and / or or co-located blocks in temporal reference picture. Moreover, typical high efficiency video codecs employ an additional motion information coding / decoding mechanism, often called merging / merge mode, where all the motion field information, which includes motion vector and corresponding reference picture index for each available reference picture list, is predicted and used without any modification / correction. Similarly, predicting the motion field information is carried out using the motion field information of adjacent blocks and / or colocated blocks in temporal reference pictures and the used motion field information is signaled among a list of motion field candidate list filled with motion field information of available adjacent / co-located blocks. Video codecs may support motion compensated prediction from one source image (uni-prediction) and two sources (bi-prediction). In the case of uniprediction a single motion vector is applied whereas in the case of bi-prediction two motion vectors are signaled and the motion compensated predictions from two sources are averaged to create the final sample prediction. In the case of weighted prediction the relative weights of the two predictions can be adjusted, or a signaled offset can be added to the prediction signal.
[0080] In addition to applying motion compensation for inter picture prediction, similar approach can be applied to intra picture prediction. In this case the displacement vector indicates where from the same picture a block of samples can be copied to form a prediction of the block to be coded or decoded. This kind of intra block copying methods can improve the coding efficiency substantially in presence of repeating structures within the frame - such as text or other graphics.
[0081] In video codecs the prediction residual after motion compensation or intra prediction may be first transformed with a transform kernel (like DCT) and then coded. The reason for this is that often there still exists some correlation among the residual and transform can in many cases help reduce this correlation and provide more efficient coding.
[0082] Many video encoders utilize Lagrangian cost functions to find optimal coding modes, e.g. the desired Macroblock mode and associated motion vectors. This kind of cost function uses a weighting factor A to tie together the (exact or estimated) image distortion due to lossy coding methods and the (exact or estimated) amount of information that is required to represent the pixel values in an image area:
[0083] C = D + AR (Eq. 1 )
[0084] Where C is the Lagrangian cost to be minimized, D is the image distortion (e.g. Mean Squared Error) with the mode and motion vectors considered, and R the number of bits needed to represent the required data to reconstruct the image block in the decoder (including the amount of data to represent the candidate motion vectors). The phrase along the bitstream (e.g. indicating along the bitstream) may be defined to refer to out-of-band transmission, signaling, or storage in a manner that the out-of-band data is associated with the bitstream. The phrase decoding along the bitstream or alike may refer to decoding the referred out- of-band data (which may be obtained from out-of-band transmission, signaling, or storage) that is associated with the bitstream. For example, an indication along the bitstream may refer to metadata in a container file that encapsulates the bitstream.
[0085] Scalable video coding refers to coding structure where one bitstream can contain multiple representations of the content at different bitrates, resolutions or frame rates. In these cases, the receiver can extract the desired representation depending on its characteristics (e.g., resolution that matches best the display device). Alternatively, a server or a network element can extract the portions of the bitstream to be transmitted to the receiver depending on e.g., the network characteristics or processing capabilities of the receiver. A scalable bitstream may consist of a “base layer” providing the lowest quality video available and one or more enhancement layers that enhance the video quality when received and decoded together with the lower layers. In order to improve coding efficiency for the enhancement layers, the coded representation of that layer typically depends on the lower layers. E.g., the motion and mode information of the enhancement layer can be predicted from lower layers. Similarly, the pixel data of the lower layers can be used to create prediction for the enhancement layer.
[0086] A scalable video codec for quality scalability (also known as Signal-to-Noise or SNR) and / or spatial scalability may be implemented as follows. For a base layer, a conventional non-scalable video encoder and decoder is used. The reconstructed / decoded pictures of the base layer are included in the reference picture buffer for an enhancement layer. In H.264 / AVC, HEVC, and similar codecs using reference picture list(s) for inter prediction, the base layer decoded pictures may be inserted into a reference picture list(s) for coding / decoding of an enhancement layer pictures similarly to the decoded reference pictures of the enhancement layer. Consequently, the encoder may choose a base-layer reference picture as inter prediction reference and may indicate its use e.g., with a reference picture index in the coded bitstream. The decoder decodes from the bitstream, for example from a reference picture index, that a base-layer picture is used as an inter-prediction reference for the enhancement layer. When a decoded base-layer picture is used as a prediction reference for an enhancement layer, it is referred to as an inter-layer reference picture.
[0087] In addition to quality scalability following scalability modes exist:
[0088] • Spatial scalability: Base layer pictures are coded at a lower resolution than enhancement layer pictures.
[0089] • Bit-depth scalability: Base layer pictures are coded at lower bit-depth (e.g., 8 bits) than enhancement layer pictures (e.g., 10 or 12 bits).
[0090] • Chroma format scalability: Enhancement layer pictures provide higher fidelity in chroma (e.g., coded in 4:4:4 chroma format) than base layer pictures (e.g., 4:2:0 format).
[0091] In all of the above scalability cases, base layer information could be used to code enhancement layer to minimize the additional bitrate overhead.
[0092] Scalability can be enabled in two basic ways. Either by introducing new coding modes for performing prediction of pixel values or syntax from lower layers of the scalable representation or by placing the lower layer pictures to the reference picture buffer (decoded picture buffer, DPB) of the higher layer. The first approach is more flexible and thus can provide better coding efficiency in most cases. However, the second, reference frame -based scalability, approach can be implemented very efficiently with minimal changes to single layer codecs while still achieving majority of the coding efficiency gains available. A reference frame -based scalability codec can be implemented by utilizing the same hardware or software implementation for all the layers, just taking care of the DPB management by external means.
[0093] To be able to utilize parallel processing, images can be split into independently codable and decodable image segments (slices or tiles). Slices typically refer to image segments constructed of certain number of basic coding units that are processed in default coding or decoding order. Tiles typically refer to image segments that have been defined as rectangular image regions that are processed at least to some extend as individual frames.
[0094] Video may be encoded in YUV or YCbCr color space as that is found to reflect some characteristics of human visual system and allows using lower quality representation for chrominance / color difference (Cb and Cr) channels as human perception is less sensitive to the chrominance fidelity those channels represent.
[0095] Typical video codecs, such as H.265 / HEVC and H.266 / WC standards use context adaptive binary arithmetic coder (CABAC) for the purpose of compressing and decompressing different syntax elements efficiently. Being a binary arithmetic coder, CABAC encodes and decodes syntax elements of size 1 bit. An alternative is to use non-binary arithmetic coders, where encoded and decoded symbols can represent syntax elements with more than one bit of information. The following example is applicable to both binary and non-binary arithmetic coders, although examples are given for the case of binary arithmetic coding for simplicity.
[0096] In case of context adaptive arithmetic coding, each syntax element has an estimate of probabilities for its possible values. Sometimes the estimated probability of a syntax element may depend on other data available during the coding process. For example, it may depend on the coding mode of a coding unit. Thus, it has been found beneficial to define multiple “contexts” for some syntax elements to be able to estimate the probability of that syntax element more accurately considering the “context”, or the environment, where the syntax element appears. The probability estimates thus provide the context or a model that describes the behavior of the syntax element or elements associated with the context. A syntax element may have two contexts, one if the syntax element appears in an Intra coded block and another if the syntax element appears in an Inter coded block.
[0097] The efficiency of the arithmetic coder depends heavily on the accuracy of the probability estimates generated for the contexts. For example, the H.266 standard estimates the probabilities of syntax element being 0 or 1 using two estimators with different characteristics and calculates the active probability for the context using an average of the two estimators. Estimators may be configured to output estimates when activated. In that case, the estimators have different adaptation rates: The first adaptation rate is reacting faster to the changes in the encoded and decoded syntax element values. The second adaptation rate reacts slower, estimating more the long-term trend of the values of the syntax element. The first estimate can then be considered to have a shorter adaptation window and the second estimate a longer one.
[0098] Encoding and decoding of image and video content may be done in a nested tree order within a coding tree unit (CTU), starting from the top-left corner of the CTU and ending to the bottom-right corner of the CTU. As a result, when done with encoding and decoding a specific CTU, the probability estimates of the arithmetic coder are tuned for the data present in the bottom-right corner area of the CTU. When progressing to a new CTU, the encoding or decoding process continues again from the top-left corner of the new CTU, and the probability estimates or context states of the arithmetic coder might not be ideal, as the processing has jumped to a different spatial location in the pictures. In different contexts and in different image or video codec specifications, such coding tree units may naturally be called with different names, such as superblocks, or macroblocks, or with other applicable names.
[0099] A publication “JVET-AE0058: “AHG12: Spatial CABAC tuning” from July 2023 describe a method to update states of an arithmetic encoder or decoder based on discontinuities in scanning of coding units within a picture. The publication discloses that bins or syntax element values can be stored in a memory to be used to update probability states of later coding units or coding tree units. Examples are given on how to store the bin or syntax element values, for example, by storing a bin value and a context index together to be able to know at a later stage what context the bin is associated with. As another example, an encoder or a decoder can store bins in a larger table that has entries for every possible context supported by the implementation, and each of those entries can contain slots for storing bins for a specific context.
[0100] To be able to adjust the probabilities of the arithmetic coding engine towards their optimal values in the local neighborhood of syntax elements to be coded or decoded, some memory is needed to store those bins or syntax element values in different neighborhoods in the picture. As the amount of such data can be relatively large, it needs to be stored efficiently to be able minimize the amount of memory required for the storage (translating directly to implementation cost, especially in hardware-based implementations). It is an aim of the present embodiments to provide a solution that selects a subset of arithmetic coding context classes to be included in the set of context classes to be stored for future use to adjust the probability states of the arithmetic coder. Various examples provide different approaches to select the subset of the arithmetic coding context classes. For example, N first context classes (wherein N is a positive integer) representing syntax elements of bottom coding units in each coding tree units can be selected. It is further taught how the bins associated with the selected context classes can be efficiently represented for storage using minimal amount of memory for the storage operation.
[0101] According to a first embodiment, an update process for the state variables of an arithmetic coder is triggered when starting to encode or decode the current coding tree unit (CTU). For each, or for some determined set of CTUs, a selected number of bin values associated with different context classes can be stored to be used by the update process. In the first embodiment, the number of context classes for which bin values are stored is limited to N. N can be a pre-determined number, such as 64, 96 or 128, or a number that is derived using a set of given rules, or it can be indicated in a video or image bitstream. Context classes, for which bin values are stored, are determined by selecting the first N context classes in the order syntax elements with their associated context classes appear in the bitstream. In this embodiment, the bottom coding units (CUs) of the CTU are included as candidate CUs for the process, although also other selection can be made in alternative embodiments. To allow efficient storage of the bin values, a maximum number of bins M for each context class that are used in the state variable update process is limited. The limit can be a pre-determined number, such as 3, 4, 5, 6, 7 or 8; or it can be can be derived in different ways. For example, it can be indicated in a video or image bitstream.
[0102] Update of the state variables of the arithmetic coder can be performed in different ways using the selected bin values. For example, it can be done as presented in afore-mentioned publication “JVET-AE0058: “AHG12: Spatial CABAC tuning” from July 2023, or in other manner considering stored bin values of one or more CTUs. The first embodiment can be described as follows by using pseudo-code, where N is the maximum number of context classes, whose bins can be stored, and M is the maximum number of bins that can be stored for each such context class: for a CTU in picture: numberOfActiveContextClasses = 0 arrayOfActiveContextClasses = empty() arrayOfStoredBins = empty() for a CU located on the bottom of the CTU: for a bin used to represent syntax elements of the CU: contextclass = contextClassOf(bin) if contextclass is already in arrayOfActiveContextClasses: if numberOfBinsStored for contextclass < M: store bin to arrayOfStoredBins else if numberOfActiveContextClasses < N: include contextclass in arrayOfActiveContextClasses increase numberOfActiveContextClasses by 1 store bin to arrayOfStoredBins
[0103] The elements of arrayOfStoredBins can be defined to include three aspects or elements of their own: 1 ) a context identifier (ID) defining the context class uses in arithmetic coding of bins stored in the element of arrayOfStoredBins, 2) a counter indicating how many bins of this context class has been stored in arrayOfStoredBins, and 3) array of bin values indicating the stored bin values of this context class. An example implementation of an element of arrayOfStoredBins can be given, for example, as follows, where the first element with uint16_t identifier is representing a 16-bit context ID, the second element with uint8_t identifier is representing a 8-bit counter, and the third element with uint8_t identifier is representing an 8-bit array of stored bins. struct BinStorageElement
[0104] { uintl6_t ctxld; uint8_t count; uint8_t bins;
[0105] } As an alternative example, if the number of context categories supported by the arithmetic coder can be represented by a 10-bit identifier, and the maximum number of bins that can be stored (M) is determined to be 5, and thus a 3-bit counter is enough to represent the number of stored bins for each context class, the following selection can be made using the same notation: struct BinStorageElement
[0106] { uintl0_t ctxld; uint3_t count; uint5_t bins;
[0107] }
[0108] As a further alternative example, the bins used for the state variable update for a given context class can be selected so that the number of bins with value of 0 (mO) will not exceed a threshold value M and the number of bins with a value of 1 (ml) will not exceed the threshold value M. In addition, any bins for that context class that follow the selected bins in the set of candidate coding units once either mO or ml has reached M should be omitted from the update process to avoid misrepresenting the actual probabilities of occurrences of ones and zeros due to saturation of one of the bin categories. In this case, the selection of a bin to be included in the state variable update process can be presented by the following pseudo-code: for a bin used to represent syntax elements of the CU: contextclass = contextClassOf(bin) if contextclass is already in arrayOfActiveContextClasses: if mO for contextclass < M and ml for contextclass < M: store bin to arrayOfStoredBins else if numberOfActiveContextClasses < N: include contextclass in arrayOfActiveContextClasses increase numberOfActiveContextClasses by 1 store bin to arrayOfStoredBins
[0109] In the above example, the order of the bins is lost and it can be beneficial to use alternative methods for updating the context state parameters. For example, the difference between the number of bins with value 0 and value 1 can be calculated, and the absolute value of that difference can determine how many context updates are performed to update the context state parameter of the arithmetic coder and the bine value with larger number of occurrences can be selected as the bin value used in the update process. As an example, the following pseudo-code can be used to determine the number of updates numUpdates to be performed, and the bin value binValue is used to update the probability state parameters: numUpdates = abs(ml - mO); binValue = ml > mO ? 1 : 0;
[0110] An update process for one state parameter stateParameter can be represented as: numUpdates = abs(ml - mO); binValue = ml > mO ? 1 : 0; if binValue is 0: stateParameter = stateParameter - d * numUpdates else: stateParameter = stateParameter + d * numUpdates
[0111] As a further example, the update can be implemented using binary shift operations. Also, in that case, the amount of update to the state parameter can be modified based on the determined numUpdates parameters and using nominal context class dependent rate parameter rateO and binary masks MASK_X and MASK_Y which can be selected for example depending on the bitdepth used to represent different probability states, as follows: numUpdates = abs(ml - mO); binValue = ml > mO ? 1 : 0; stateParameter -= numUpdates * ( (stateParameter » rateO) & MASK_X ); if binValue is 1: stateParameter += numUpdates * ( (MASK_Y » rateO) & MASK_X ); According to an embodiment when updating probability estimates using stored bins, the arithmetic coder is configured to read the stored bins associated with one context class in the same order those were originally stored.
[0112] According to an embodiment, a bin value of a first coding tree unit is determined to be used to perform a context probability update for a second coding tree unit if the context class of the bin belongs N first context classes found in a selected set of candidate coding units or in a selected set of candidate syntax elements.
[0113] In an embodiment a bin value of a first coding tree unit is determined to be used to perform a context probability update for a second coding tree unit if the context class of the bin belongs N first context classes found in a selected set of candidate coding units or in a selected set of candidate syntax elements and if the bin belongs to M first bins of the same context class in the same selected set of candidate coding units or candidate syntax elements.
[0114] According to an embodiment a bin value of a first coding tree unit is determined to be used to perform a context probability update for a second coding tree unit if the bin belongs to a coding unit located on the bottom of the first coding tree unit.
[0115] According to an embodiment the bins of a first coding tree unit selected to be used for probability update for a second coding tree unit are stored in a memory unit together with a counter representing the number of bins stored for the context class of the bin.
[0116] According to an embodiment the bins of a first coding tree unit selected to be used for probability update for a second coding tree unit are stored in a memory unit together with an identifier representing the context class of the bins.
[0117] According to an embodiment the bins of a first coding tree unit selected to be used for probability update for a second coding tree unit are stored in a memory unit together with an identifier representing the context class of the bin and a counter representing the number of bins stored for the context class of the bin. According to an embodiment the number of context classes for which bins of a first coding tree unit are stored to be used to update context probabilities of a second coding tree unit is limited to N, where N can be a pre-determined number, such as 64, 96 or 128, a number that is derived using a set of given rules, or it can be indicated in a video or image bitstream.
[0118] According to an embodiment the maximum number of bins for each context class that are used in the state variable update process is limited to a predetermined number, such as 3, 4, 5, 6, 7 or 8.
[0119] The method according to an embodiment is shown in Figure 6. The method generally comprises (1 ) determining a first coding tree unit to be encoded or decoder; (2) determining a coding unit belonging to the first coding tree unit; (3) determining a bin value used by an arithmetic encoder or decoder to represent a syntax element of the coding unit; (4) determining if the bin value is going to be used to perform a location-based context probability update for a second coding tree unit, where the determining includes: a) determining a context class of the syntax element; b) determining if the context class belongs to a set of context classes selected to be used for the probability update for a second coding tree unit; c) determining if the number of bins with the same context class as the context class of the bin which are already included for the context probability update is below a selected maximum number; and (5) using the bin value to update a probability state parameter used in the second coding tree unit depending on the result of step (4).
[0120] The first coding tree unit is at least partially different from the second coding tree unit, wherein the second coding tree unit comprises a previously encoded or decoded coding tree unit.
[0121] In the method according to another embodiment, the step (4b) of determining if the context class belongs to the set of context classes selected to be used for the probability update for the second coding tree include determining a maximum amount of context classes considered for the update and determining if the number of included context classes for the first coding tree unit is below the maximum amount. Each of the steps can be implemented by a respective module of a computer system.
[0122] An apparatus according to an embodiment comprises means for determining a first coding tree unit to be encoded or decoder; means for determining a coding unit belonging to the first coding tree unit; means for determining a bin value used by an arithmetic encoder or decoder to represent a syntax element of the coding unit; means for determining if the bin value is going to be used to perform a location-based context probability update for a second coding tree unit, where the determining includes: a) determining a context class of the syntax element; b) determining if the context class belongs to a set of context classes selected to be used for the probability update for a second coding tree unit; c) determining if the number of bins with the same context class as the context class of the bin which are already included for the context probability update is below a selected maximum number; and means for using the bin value to update a probability state parameter used in the second coding tree unit depending on the result from means for determining if the bin value is going to be used to perform a location-based context probability update. The means comprises at least one processor, and a memory including a computer program code, wherein the processor may further comprise processor circuitry. The memory and the computer program code are configured to, with the at least one processor, cause the apparatus to perform the method of Figure 3 according to various embodiments.
[0123] Figure 4 illustrates an example of an electronic apparatus 400, being an example of video coding system. In some embodiments, the apparatus may be a mobile terminal or a user equipment of a wireless communication system or a camera device. The apparatus 400 may also be comprised at a local or a remote server or a graphic processing unit of a computer. The apparatus may also be comprised as part of a head-mounted display device
[0124] The apparatus may be configured to perform various functions, such as for example, gathering information by one or more sensors, encoding and / or decoding information, receiving and / or transmitting information, analyzing information gathered or received by the apparatus. An apparatus configured to encode a video scene may optionally comprise one or more microphones for capturing the scene and / or one or more cameras for capturing information about the physical environment in which the scene is captured. Alternatively, the apparatus configured for encoding may be configured to receive information about an environment in which a scene is captured and / or a simulated environment. An apparatus configured to decode and / or render the video scene may be configured to receive a bitstream comprising encoded video. An apparatus configured to decode and / or render the video scene may comprise one or more speakers / audio transducers and / or displays, and / or may be configured to transmit a decoded scene or signals to a device comprising one or more speakers / audio transducers and / or displays. An apparatus configured to decode and / or render the video scene may comprises a user equipment , a head-mounted display, or another device capable of rendering to a user an AR; VR and / or MR experience.
[0125] The apparatus 400 comprises one or more processors 410 and one or more memories 420 and one or more transceivers interconnected through one or more buses. The one or more memories 420 store computer instructions, for example in respective modules (Modulel , Module2, ModuleN). The one or more memories may store data in the form of image, video and / or audio data, and / or may also store instructions to be executed by the processors or the processor circuitry. The one or more processors may comprise a central processing unit (CPU) and / or a graphical processing unit (GPU). The one or more buses may be address, data or control buses, and may include interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment. The apparatus also comprises a codec 430 that is configured to implement various embodiments relating to present solution. According to some embodiments, the apparatus may comprise an encoder or a decoder. The apparatus 400 also comprises a communication interface 440 which is suitable for generating wireless communication signals for example for communication with a cellular communications network, a wireless communications system, or a wireless local area network, and thus enabling data transfer over data transfer network 450.
[0126] The apparatus 400 may comprise a display in the form of a liquid crystal display. In other embodiments of the invention the display may be any suitable display technology suitable to display an image or video. The apparatus 400 may further comprise a keypad. In other embodiments of the invention any suitable data or user interface mechanism may be employed. For example, the user interface may be implemented as a virtual keyboard or data entry system as part of a touch-sensitive display. The apparatus 400 may comprise a microphone or any suitable audio input which may be a digital or analogue signal input. The apparatus 400 may further comprise an audio output device which in embodiments of the invention may be any one of: an earpiece, speaker, or an analogue audio or digital audio output connection. The apparatus 400 may also comprise a battery (or in other embodiments of the invention the device may be powered by any suitable mobile energy device such as solar cell, fuel cell or clockwork generator). The apparatus may further comprise a camera capable of recording or capturing images and / or video. The camera may be a multi-lens camera system having at least two camera sensors. The camera is capable of recording or detecting individual frames which are then passed to the codec 430 or to processor 410. The apparatus may receive the video and / or image data for processing from another device prior to transmission and / or storage.
[0127] The apparatus 400 may further comprise e.g., the other functional units disclosed in any of the Figures 1 - 2 for implementing any of the present embodiments.
[0128] The various embodiments can be implemented with the help of computer program code that resides in a memory and causes the relevant apparatuses to carry out the method. For example, a device may comprise circuitry and electronics for handling, receiving and transmitting data, computer program code in a memory, and a processor that, when running the computer program code, causes the device to carry out the features of an embodiment. Yet further, a network device like a server may comprise circuitry and electronics for handling, receiving and transmitting data, computer program code in a memory, and a processor that, when running the computer program code, causes the network device to carry out the features of various embodiments.
[0129] The various embodiments can be implemented in a system comprising multiple communication devices which can communicate through one or more networks. The system may comprise any combination of wired or wireless networks including, but not limited to a wireless cellular telephone network, a wireless local area network (WLAN) such as defined by any of the IEEE 802.x standards, a BLUETOOTH™ personal area network, an Ethernet local area network, a token ring local area network, a wide area network, and / or the Internet. A wireless network may implement network virtualization, which is the process of combining hardware and software network resources and network functionality into a single, software based administrative entity, a virtual network.
[0130] It may also be noted that operations of example embodiments of the present disclosure maybe carried out by a plurality of cooperating devices (e.g. centralized Radio Access Network “cRAN”).
[0131] The system may include both wired and wireless communication devices and / or electronic devices suitable for implementing example embodiments of the present disclosure.
[0132] For example, the system may comprise a mobile telephone network and a representation of the internet. Connectivity to the internet may include, but is not limited to, long range wireless connections, short range wireless connections, and various wired connections including, but not limited to, telephone lines, cable lines, power lines, and similar communication pathways.
[0133] The example communication devices of the system may include, but are not limited to, an apparatus, a combination of a personal digital assistant (PDA) and a mobile telephone, a PDA, an integrated messaging device (IMD), a desktop computer, a notebook computer, and a head-mounted display (HMD). The apparatus according to present embodiments may comprise any of such example communication devices. In an example embodiment of the present disclosure, more than one of these devices, or a plurality of one or more of these devices, may perform the method according to various embodiments.
[0134] If desired, the different functions discussed herein may be performed in a different order and / or concurrently with other. Furthermore, if desired, one or more of the above-described functions and embodiments may be optional or may be combined.
[0135] Although various aspects of the embodiments are set out in the independent claims, other aspects comprise other combinations of features from the described embodiments and / or the dependent claims with the features of the independent claims, and not solely the combinations explicitly set out in the claims. It is also noted herein that while the above describes example embodiments, these descriptions should not be viewed in a limiting sense. Rather, there are several variations and modifications, which may be made without departing from the scope of the present disclosure as, defined in the appended claims.
Claims
Claims:
1. An apparatus, comprising i) means for determining a first coding tree unit to be encoded or decoded; ii) means for determining a coding unit belonging to the first coding tree unit;Hi) means for determining a bin value used by an arithmetic encoder or an arithmetic decoder to represent a syntax element of the coding unit; iv) means for determining if the bin value is going to be used to perform a location-based context probability update for a second coding tree unit, where the determining includes: a) determining a context class of the syntax element; b) determining if the context class belongs to a set of context classes selected to be used for the probability update for a second coding tree unit; c) determining if the number of bins with the same context class as the context class of the bin which are already included for the context probability update is below a selected maximum number; v) means for using the bin value to update a probability state parameter used in the second coding tree unit depending on the result of the means for determining on step iv).
2. The apparatus according to claim 1 , wherein determining if the context class belongs to the set of context classes selected to be used for the probability update for the second coding tree includes determining a maximum amount of context classes considered for the update and determining if the number of included context classes for the first coding tree unit is below the maximum amount.
3. The apparatus according to claim 1 or 2, further comprising means for storing bin values of the first coding tree unit selected to be used for the probability update for the second coding tree unit in a memory unit with an identifier representing the context class of the bin.
4. The apparatus according to any of the claims 1 to 3, further comprising means for storing bin values of the first coding tree unit selected to be used for the probability update for the second coding tree unit in a memory unit with a counter representing the number of bins stored for the context class of the bin.
5. The apparatus according to any of the claims 1 to 4, wherein number of the set of context classes to be used for the probability update for the second coding tree unit is limited to a pre-determined number.
6. The apparatus according to claim 5, further comprising means for deriving the pre-determined number using a set of given rules or from a video or image bitstream.
7. The apparatus according to any of the claims 1 to 6, further comprising means for performing binary shift operations for the update.
8. A method, comprising: i) determining a first coding tree unit to be encoded or decoded; ii) determining a coding unit belonging to the first coding tree unit;Hi) determining a bin value used by an arithmetic encoder or an arithmetic decoder to represent a syntax element of the coding unit; iv) determining if the bin value is going to be used to perform a location-based context probability update for a second coding tree unit, where the determining includes: a) determining a context class of the syntax element; b) determining if the context class belongs to a set of context classes selected to be used for the probability update for a second coding tree unit; c) determining if the number of bins with the same context class as the context class of the bin which are already included for the context probability update is below a selected maximum number;v) using the bin value to update a probability state parameter used in the second coding tree unit depending on the result of the means for determining on step iv).
9. The method according to claim 8, wherein determining if the context class belongs to the set of context classes selected to be used for the probability update for the second coding tree includes determining a maximum amount of context classes considered for the update and determining if the number of included context classes for the first coding tree unit is below the maximum amount.
10. The method according to claim 8 or 9, further comprising storing bin values of the first coding tree unit selected to be used for the probability update for the second coding tree unit in a memory unit with an identifier representing the context class of the bin.11 . The method according to any of the claims 8 to 10, further comprising storing bin values of the first coding tree unit selected to be used for the probability update for the second coding tree unit in a memory unit with a counter representing the number of bins stored for the context class of the bin.
12. The method according to any of the claims 8 to 11 , wherein number of the set of context classes to be used for the probability update for the second coding tree unit is limited to a pre-determined number.
13. The method according to claim 12, further comprising deriving the predetermined number using a set of given rules or from a video or image bitstream.
14. The method according to any of the claims 8 to 13, further comprising performing binary shift operations for the update.
15. An apparatus comprising at least one processor, memory including computer program code, the memory and the computer program code configured to, with the at least one processor, cause the apparatus to perform at least the following:i) determine a first coding tree unit to be encoded or decoded; ii) determine a coding unit belonging to the first coding tree unit; Hi) determine a bin value used by an arithmetic encoder or an arithmetic decoder to represent a syntax element of the coding unit; iv) determine if the bin value is going to be used to perform a location-based context probability update for a second coding tree unit, where the determining includes: a) determining a context class of the syntax element; b) determining if the context class belongs to a set of context classes selected to be used for the probability update for a second coding tree unit; c) determining if the number of bins with the same context class as the context class of the bin which are already included for the context probability update is below a selected maximum number; v) use the bin value to update a probability state parameter used in the second coding tree unit depending on the result of the means for determining on step iv).
Citation Information
Patent Citations
Acceleration of context adaptive binary arithmetic coding (CABAC) in video codecs
US11336921B2