Methods for context-based video coding

BR112025022032A2Pending Publication Date: 2026-09-15
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
BR112025022032
Authority / Receiving Office
BR · BR
Patent Type
Applications
Publication Date
2026-09-15

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

2 / 57 of videos are developed by standardization organizations. With increasingly advanced video encoding technologies being adopted in video standards, the encoding efficiency of the new video encoding standards is becoming ever greater. SUMMARY OF DESCRIPTION

[0004] The embodiments of the present invention provide methods and apparatus for initializing a set of context model probabilities for a layer in context-adaptive binary arithmetic coding (CABAC).

[0005] According to some exemplary embodiments, a decoding method is provided including: selecting, from a plurality of predefined sets of probability parameters, a first set of probability parameters to initiate one or more context models for a layer B; and performing entropy decoding of the slice B based on one or more context models and the first set of probability parameters, wherein the selection is based on a layer B encoding condition or a bitstream signal.

[0006] According to some exemplary embodiments, a coding method is provided including: selecting, from a plurality of predefined sets of probability parameters, a first set of probability parameters to initiate one or more context models for a layer B; and performing entropy coding of the slice B based on one or more context models and the first set of probability parameters, wherein the selection is based on a layer B coding condition or a bitstream signal.

[0007] According to some exemplary embodiments, a non-transient, computer-readable storage medium is provided that stores a bitstream of a video. The bitstream includes: Petition 870250092909, dated 10 / 10 / 2025, page 9 / 100 3 / 57 select, from a plurality of predefined sets of probability parameters, a first set of probability parameters to initiate one or more context models for a layer B; and perform entropy encoding or decoding of layer B based on one or more context models and the first set of probability parameters, wherein the selection is based on a layer B encoding condition or a bitstream signal. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The embodiments and various aspects of the present invention are illustrated in the following detailed description and in the accompanying figures. Several features in the figures are not shown to scale.

[0009] FIG. 1 is a schematic diagram illustrating structures of an exemplary video sequence, according to some embodiments of the present invention.

[0010] FIG. 2A is a schematic diagram illustrating an exemplary encoding process of a hybrid video encoding system, according to some embodiments of the present invention.

[0011] FIG. 2B is a schematic diagram illustrating another exemplary encoding process of a hybrid video encoding system, according to some embodiments of the present invention.

[0012] FIG. 3A is a schematic diagram illustrating an exemplary decoding process of a hybrid video encoding system, according to some embodiments of the present invention.

[0013] FIG. 3B is a schematic diagram illustrating another exemplary decoding process of a hybrid video encoding system, according to some embodiments of the present invention. Petition 870250092909, dated 10 / 10 / 2025, page 10 / 100 4 / 57

[0014] FIG. 4 is a block diagram of an exemplary apparatus for encoding or decoding video, according to some embodiments of the present invention.

[0015] FIG. 5 is a schematic diagram illustrating a context-adaptive binary arithmetic coding (CABAC) mechanism, according to some embodiments of the present invention.

[0016] FIG. 6 illustrates an illustrative table of code words used for binary encoding, according to some embodiments of the present invention.

[0017] FIG. 7 illustrates an exemplary process for updating the Range and Low variables in a binary arithmetic encoding (BAE) stage of the CABAC mechanism of FIG. 5, according to some embodiments of the present invention.

[0018] FIG. 8 illustrates an exemplary process for updating the Range and Low variables in a BAE stage of the CABAC mechanism of FIG. 5, according to some embodiments of the present invention.

[0019] FIG. 9 illustrates an illustrative set of context model probability parameters, according to some embodiments of the present invention.

[0020] FIG. 10 illustrates another exemplary set of context model probability parameters, according to some embodiments of the present invention.

[0021] FIG. 11 illustrates four exemplary sets of context model probability parameters, according to some embodiments of the present invention.

[0022] FIG. 12 illustrates a flowchart of an exemplary method for decoding a bitstream associated with a video, according to some embodiments of the present invention.

[0023] FIG. 13 illustrates a flowchart of another exemplary method for decoding a bitstream associated with a video, according to Petition 870250092909, dated 10 / 10 / 2025, p. 11 / 100 5 / 57 with some embodiments of the present invention.

[0024] FIG. 14 illustrates a flowchart of an exemplary method for encoding a bitstream associated with a video, according to some embodiments of the present invention.

[0025] FIG. 15 illustrates a flowchart of another exemplary method for encoding a bitstream associated with a video, according to some embodiments of the present invention. DETAILED DESCRIPTION

[0026] Detailed reference will now be made to the exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings in which the same numbers in different drawings represent identical or similar elements, except where otherwise represented. The implementations defined in the following description of the exemplary embodiments do not represent all implementations consistent with the invention. Instead, they are only examples of apparatus and methods consistent with aspects related to the invention as cited in the accompanying claims. The specific aspects of the present invention are described in more detail below. The terms and definitions provided in this document prevail if they conflict with terms or definitions incorporated by reference.

[0027] The present invention provides methods for initializing the probability parameters of context models for use in context-adaptive binary arithmetic coding (CABAC). According to some disclosed embodiments, the initial probability parameters can be selected from a plurality of predefined sets of probability parameters. In some embodiments, the plurality of predefined sets of probability parameters can be pre-stored in the encoder and decoder, and the selection can be derived in both the encoder and the decoder. Petition 870250092909, dated 10 / 10 / 2025, page 12 / 100 6 / 57 without explicit signaling. In some modes, the plurality of predefined sets of probability parameters can be pre-stored in the encoder and decoder, and the selection can be signaled in a bitstream. In some modes, the initial probability parameters can be selected without referring to previously encoded images. In some modes, the initial probability parameters can be selected based on the content of the current layer in which CABAC is applied.

[0028] The Joint Video Experts Team (JVET) of the ITU-T standard The Video Coding Expert Group (ITU-T VCEG) and the ISO / IEC Moving Picture Expert Group (ISO / IEC MPEG) are currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of the VVC standard is to achieve the same subjective quality as the HEVC / H.265 standard using half the bandwidth.

[0029] To achieve the same subjective quality as the standard While HEVC / H.265 uses half the bandwidth, JVET is developing technologies beyond the HEVC standard using the Joint Exploration Model (JEM) reference software. Because the encoding technologies have been incorporated into JEM, JEM has achieved significantly superior encoding performance compared to HEVC.

[0030] The VVC standard was recently developed and continues to incorporate more encoding technologies that provide better compression performance. VVC is based on the same hybrid video encoding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, H.263, etc.

[0031] A video is a collection of still images (or “frames”) arranged in a temporal sequence to store visual information. A video capture device (e.g., a camera) Petition 870250092909, dated 10 / 10 / 2025, page 13 / 100 7 / 57 can be used to capture and store these images in a time-series sequence, and a video playback device (e.g., a television, a computer, a smartphone, a laptop, a video player, or any end-user terminal with a display function) can be used to display such images in time-series. Similarly, in some applications, a video capture device can transmit the captured video to the video playback device (e.g., a computer with a monitor) in real time, such as for surveillance, conferencing, or live streaming.

[0032] To reduce the storage space and transmission bandwidth required by such applications, video can be compressed before storage and transmission and decompressed before display. Compression and decompression can be implemented by software running on a processor (e.g., a processor in a generic computer) or specialized hardware. The module for compression is generally referred to as an “encoder,” and the module for decompression is generally referred to as a “decoder.” The encoder and decoder can be collectively referred to as a “codec.” The encoder and decoder can be implemented as any variety of hardware, software, or suitable combination thereof.For example, the hardware implementation of the encoder and decoder may include circuits such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. The software implementation of the encoder and decoder may include program code, computer-executable instructions, firmware, or any algorithm implemented by a computer. Petition 870250092909, dated 10 / 10 / 2025, page 14 / 100 8 / 57 suitable computer or fixed process on a computer-readable medium. Video compression and decompression can be implemented by various algorithms or standards, such as the MPEG-1, MPEG-2, MPEG-4, H.26x series or similar. In some applications, the codec may decompress the video from a first encoding standard and then compress the decompressed video again using a second encoding standard, in which case the codec may be referred to as a “transcoder”.

[0033] The video encoding process can identify and retain useful information that can be used to reconstruct an image and discard irrelevant information for reconstruction. If the discarded and irrelevant information cannot be fully reconstructed, the aforementioned encoding process can be referred to as “lossy”. Otherwise, the same process can be referred to as “lossless”. Most encoding processes occur with data loss, which is a trade-off to reduce the required storage space and transmission bandwidth.

[0034] Useful information from an image being encoded (referred to as a “current image”) includes changes with respect to a reference image (e.g., a previously encoded and reconstructed image). Such changes may include changes in position, brightness, or color of the pixels, with position changes being the most relevant. Changes in the position of a group of pixels representing an object may reflect the object's movement between the reference image and the current image.

[0035] An encoded image without reference to another image (i.e., it is its own reference image) is referred to as an “I-image”. An image is referred to as a “P-image” if some or all of it is encoded. Petition 870250092909, dated 10 / 10 / 2025, p. 15 / 100 9 / 57 blocks (for example, blocks that generally refer to portions of the video image) in the image are predicted using intraprediction or interprediction with a reference image (for example, uniprediction). An image is referenced with a “B-image” if at least one block in it is predicted with two referenced images (for example, bi-prediction).

[0036] FIG. 1 illustrates structures of an exemplary video sequence 100, according to some embodiments of the present invention. The video sequence 100 can be a live video or a video that has been captured and archived. The video 100 can be a real-life video, a computer-generated video (e.g., computer game video), or a combination thereof (e.g., a real-life video with augmented reality effects). The video sequence 100 can be introduced from a video capture device (e.g., a camera), a video archive (e.g., a video file stored on a storage device) containing previously captured video, or a video transmission interface (e.g., a video broadcast transceiver) to receive video from a video content provider.

[0037] As shown in FIG. 1, the video sequence 100 may include a series of images arranged temporally along a timeline, including images 102, 104, 106, and 108. Images 102–106 are continuous, and there are more images between images 106 and 108. In FIG. 1, image 102 is an I-image; the reference image is image 102 itself. Image 104 is a P-image; the reference image is image 102, as indicated by the arrow. Image 106 is a B-image; the reference images are images 104 and 108, as indicated by the arrows. In some embodiments, the reference image of an image (e.g., image 104) Petition 870250092909, dated 10 / 10 / 2025, p. 16 / 100 10 / 57 may not be immediately before or after the image. For example, the reference image for image 104 may be an image before image 102. It should be noted that the reference images for images 102-106 are only examples, and the present invention does not limit the embodiments of the reference images as exemplified in FIG. 1.

[0038] Typically, video codecs do not encode or decode an entire image at once due to the computational complexity of such tasks. Instead, they divide the image into basic segments and encode or decode the image segment by segment. Such basic segments are referred to as basic processing units (“BPUs”) in the present invention. For example, structure 110 in FIG. 1 shows an exemplary structure of a sequence of 100 video images (e.g., any of the images 102-108). In structure 110, an image is divided into 4^4 basic processing units, the boundaries of which are shown by dashed lines. In some applications, the basic processing units may be referred to as “macroblocks” in some video coding standards (e.g., MPEG, H.261, H.263, or H.264 / AVC families), or as “coding tree units” (“CTUs”) in some other video coding standards (e.g., H.(265 / HEVC or H.266 / VVC). The basic processing units can have variable sizes in an image, such as 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, or any arbitrary pixel shape and size. The sizes and shapes of the basic processing units can be selected for an image based on the balance between encoding efficiency and levels of detail to maintain in the basic processing unit.

[0039] The basic processing units can be logical units, which can include a group of different data types. Petition 870250092909, dated 10 / 10 / 2025, page 17 / 100 11 / 57 Videos stored in computer memory (e.g., in a video frame buffer). For example, a basic processing unit for a color image might include a luma component (Y) representing achromatic brightness information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements, where the luma and chroma components might be the same size as the basic processing unit. The luma and chroma components might be referred to as “coding tree blocks” (“CTBs”) in some video encoding standards (e.g., H.265 / HEVC or H.266 / VVC). Any operation performed on a basic processing unit can be repeatedly performed on each of its luma and chroma components.

[0040] Video encoding has multiple stages of operations, examples of which are shown in FIGS. 2A-2B and FIGS. 3A-3B. For each stage, the size of the basic processing units may still be too large for processing, so they may be further divided into segments referred to as “basic processing subunits” in the present invention. In some embodiments, the basic processing subunits may be referred to as “blocks” in some video encoding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC), or as “encoding units” (“CUs”) in some other video encoding standards (e.g., H.265 / HEVC or H.266 / VVC). A basic processing subunit may be the same size as or smaller than the basic processing unit.Similar to basic processing units, basic processing subunits are also logical units, which may include a group of different types of video data (e.g., Y, Cb, Cr and associated syntax elements) stored in computer memory (e.g., Petition 870250092909, dated 10 / 10 / 2025, page 18 / 100 12 / 57 in a video frame buffer). Any operation performed on a basic processing subunit can be repeatedly performed on each of its luma and chroma components. It should be noted that each division can be performed to additional levels depending on processing needs. It should also be noted that different stages can divide the basic processing units using different schemes.

[0041] For example, in a mode decision stage (an example of which is shown in FIG. 2B), the encoder can decide which prediction mode (e.g., intra-image prediction or inter-image prediction) to use for a basic processing unit, which may be too large to make such a decision. The encoder can divide the basic processing unit into multiple basic processing subunits (e.g., CUs as in H.265 / HEVC or H.266 / VVC) and decide on a prediction type for each individual basic processing subunit.

[0042] For another example, in a prediction stage (an example of which is shown in FIGS. 2A-2B), the encoder can perform prediction operations at the level of basic processing subunits (e.g., CUs). However, in some cases, a basic processing subunit may still be too large to process. The encoder can further divide the basic processing subunit into smaller segments (e.g., referred to as “prediction blocks” or “PBs” in H.265 / HEVC or H.266 / VVC), at the level at which prediction operations can be performed.

[0043] For another example, in a transform stage (an example of which is shown in FIGS. 2A-2B), the encoder can perform a transform operation for residual basic processing subunits (e.g., CUs). However, in some cases, a basic processing subunit may still be too large for Petition 870250092909, dated 10 / 10 / 2025, page 19 / 100 13 / 57 processing. The encoder can further divide the basic processing subunit into smaller segments (e.g., referred to as “transform blocks” or “TBs” in H.265 / HEVC or H.266 / VVC), at the level at which the transform operation can be performed. It should be noted that the division schemes of the same basic processing subunit may be different in the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and the transform blocks of the same CU may have different sizes and numbers.

[0044] In structure 110 of FIG. 1, the basic processing unit 112 is further divided into basic processing subunits 3^3, whose boundaries are shown in dashed lines. Different basic processing units of the same image may be divided into basic processing subunits in different schemes.

[0045] In some implementations, to provide parallel processing capability and resilience to errors in video encoding and decoding, an image can be divided into regions for processing, so that, for one region of the image, the encoding or decoding process does not depend on any information from any other regions of the image. In other words, each region of the image can be processed independently. By doing so, the codec can process different regions of an image in parallel, thus increasing encoding efficiency. Furthermore, when data from one region is corrupted during processing or lost in network transmission, the codec can correctly encode or decode other regions of the same image without depending on corrupted or lost data, thus providing error resilience. In some video encoding standards, an image can be divided into different types of regions. For example, H.265 / HEVC and Petition 870250092909, dated 10 / 10 / 2025, page 20 / 100 14 / 57 H.266 / VVC provides two types of regions: "layers" and "blocks". It should also be noted that different images in the 100-video sequence may have different partitioning schemes for dividing an image into regions.

[0046] For example, in FIG. 1, structure 110 is divided into three regions 114, 116, and 118, whose boundaries are shown as solid lines within structure 110. Region 114 includes four basic processing units. Each of regions 116 and 118 includes six basic processing units. It should be noted that the basic processing units, basic processing subunits, and regions of structure 110 in FIG. 1 are only examples, and the present invention does not limit their embodiments.

[0047] FIG. 2A illustrates a schematic diagram of an exemplary 200A encoding process, consistent with the embodiments of the description. For example, the 200A encoding process can be performed by an encoder. As shown in FIG. 2A, the encoder can encode a video sequence 202 into a video bitstream 228 according to the 200A process. Similar to the video sequence 100 in FIG. 1, the video sequence 202 can include a set of images (referred to as “original images”) arranged in a temporal order. Similar to the structure 110 in FIG. 1, each original image of the video sequence 202 can be divided by the encoder into basic processing units, basic processing subunits, or regions for processing. In some embodiments, the encoder can perform the 200A process at the basic processing unit level for each individual image of the video sequence 202.For example, the encoder can execute process 200A iteratively, where the encoder can encode a basic processing unit in one iteration of process 200A. In some embodiments, the encoder can execute process 200A in... Petition 870250092909, dated 10 / 10 / 2025, page 21 / 100 15 / 57 parallel for regions (e.g., regions 114-118) of each original image in the 202 video sequence.

[0048] In FIG. 2A, the encoder can feed a basic processing unit (referred to as an “original BPU”) of an original video sequence image 202 to prediction stage 204 to generate prediction data 206 and predicted BPU 208. The encoder can subtract the predicted BPU 208 from the original BPU to obtain the residual BPU 210. The encoder can feed the residual BPU 210 to the transform stage 212 and quantization stage 214 to generate quantized transform coefficients 216. The encoder can feed prediction data 206 and quantized transform coefficients 216 to the binary encoding stage 226 to generate the video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 can be referred to as a "path of progress".During process 200A, after quantization stage 214, the encoder can feed the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220 to generate the reconstructed residual BPU 222. The encoder can add residual BPU 222 to the predicted BPU 208 to generate prediction reference 224, which is used in prediction stage 204 for the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A can be referenced as a “reconstruction path”. The reconstruction path can be used to ensure that the encoder and decoder use the same reference data for prediction.

[0049] The encoder can execute process 200A iteratively to encode each original BPU of the original image (in the forward path) and generate the predicted reference 224 to encode the next original BPU of the original image (in the reconstruction path). After encoding the original BPUs of the original image, the encoder can proceed to Petition 870250092909, dated 10 / 10 / 2025, page 22 / 100 16 / 57 encode the next image in the video sequence 202.

[0050] Referring to process 200A, the encoder may receive the video sequence 202 generated by a video capture device (e.g., a camera). The term “receive” as used in this document may refer to receiving, inserting, acquiring, retrieving, obtaining, reading, accessing, or any action in any way to insert data.

[0051] In prediction stage 204, in a current iteration, the encoder may receive an original BPU and prediction reference 224, and execute a prediction operation to generate prediction data 206 and predicted BPU 208. Prediction reference 224 may be generated from the reconstruction path of the previous iteration of process 200A. The purpose of prediction stage 204 is to reduce information redundancy by extracting prediction data 206 which can be used to reconstruct the original BPU as predicted BPU 208 from prediction data 206 and prediction reference 224.

[0052] Ideally, the predicted BPU 208 can be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 is generally slightly different from the original BPU. To account for these differences, after generating the predicted BPU 208, the encoder can subtract it from the original BPU to generate the residual BPU 210. For example, the encoder can subtract values ​​(e.g., grayscale values ​​or RGB values) from pixels of predicted BPU 208 from corresponding pixel values ​​of the original BPU. Each pixel of residual BPU 210 may have a residual value as a result of such subtraction between the corresponding pixels of the original BPU and the predicted BPU 208. Compared to the original BPU, the prediction data 206 and residual BPU 210 may have few bits, but they can be used to reconstruct the original BPU without significant quality deterioration. Thus, the original BPU is compressed.

[0053] To further compress the residual BPU 210, at the stage of Petition 870250092909, dated 10 / 10 / 2025, page 23 / 100 17 / 57 transform 212, the encoder can reduce the spatial redundancy of the residual BPU 210 by decomposing it into a set of two-dimensional “base patterns,” each base pattern being associated with a “transform coefficient.” The base patterns can be the same size (e.g., the size of the residual BPU 210). Each base pattern can represent a frequency variation component (e.g., brightness variation frequency) of the residual BPU 210. None of the base patterns can be reproduced from any combinations (e.g., linear combinations) of any other base patterns. In other words, the decomposition can decompose variations of the residual BPU 210 into a frequency domain.This decomposition is analogous to a discrete Fourier transform of a function, where the basis patterns are analogous to the basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are analogous to the coefficients associated with the basis functions.

[0054] Different transform algorithms can use different basis patterns. Several transform algorithms can be used in the 212 transform stage, such as, for example, a discrete cosine transform, a discrete sine transform, or similar. The transform in the 212 transform stage is reversible. That is, the encoder can restore the residual BPU 210 by an inverse transform operation (referred to as an “inverse transform”). For example, to restore a pixel from residual BPU 210, the inverse transform can multiply the corresponding pixel values ​​of the basis patterns by their respective associated coefficients and add the products to produce a weighted sum. For a video encoding pattern, both the encoder and the decoder can use the same transform algorithm (thus, the same basis patterns). Thus, the encoder can register only Petition 870250092909, dated 10 / 10 / 2025, page 24 / 100 18 / 57 are the transform coefficients, by which the decoder can reconstruct the residual BPU 210 without receiving the base patterns from the encoder. Compared to the residual BPU 210, the transform coefficients may have few bits, but they can be used to reconstruct the residual BPU 210 without significant quality deterioration. Thus, the residual BPU 210 is additionally compressed.

[0055] The encoder can further compress the transform coefficients at quantization stage 214. In the transform process, different basis patterns can represent different frequencies of variation (e.g., frequencies of brightness variation). Since human eyes are generally better at recognizing low-frequency variation, the encoder can disregard high-frequency variation information without causing significant quality deterioration in the decoding. For example, at quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (referred to as a “quantization scaling factor”) and rounding the quotient to its nearest integer.Following this operation, some transform coefficients of the high-frequency base patterns can be converted to zero, and the transform coefficients of the low-frequency base patterns can be converted to smaller integers. The encoder can disregard the zero-valued 216 quantized transform coefficients, whereby the transform coefficients are further compressed. The quantization process is also reversible, in which the 216 quantized transform coefficients can be reconstructed into the transform coefficients in an inverse quantization operation (referred to as “inverse quantization”).

[0056] Since the encoder disregards the remainder of such divisions in the rounding operation, the quantization stage 214 Petition 870250092909, dated 10 / 10 / 2025, page 25 / 100 19 / 57 may contain data loss. Typically, the quantization stage 214 can contribute to the greatest loss of information in process 200A. The greater the loss of information, the fewer bits of the quantized transform coefficients 216 are needed. To obtain different levels of loss of information, the encoder can use different values ​​of the quantization parameter or any other parameter of the quantization process.

[0057] In the binary coding stage 226, the encoder can encode the prediction data 206 and the quantized transform coefficients 216 using a binary coding technique, such as, for example, context-adaptive binary arithmetic coding (CABAC), entropy coding, variable-length coding, arithmetic coding, Huffman coding, or any other lossless or lossy compression algorithm.

[0058] For example, the CABAC encoding process at the 226 binary encoding stage may include a binarization step, a context-based modeling step, and a binary arithmetic encoding step. If the syntax element is not binary, the encoder first maps the syntax element into a binary sequence. The encoder may select a context-based encoding mode or a bypass encoding mode. In some embodiments, for context-based encoding mode, the probability model of the bin to be encoded is selected by the “context,” which refers to previously encoded syntax elements. Then, the bin and the selected context model are passed through an arithmetic encoding mechanism, which encodes the bin and updates the corresponding probability distribution of the context model.In some modes, for bypass encoding mode, without selecting the probability model by "context", bins are encoded with one. Petition 870250092909, dated 10 / 10 / 2025, page 26 / 100 20 / 57 fixed probability (e.g., a probability equal to 0.5). In some embodiments, the bypass encoding mode is selected for specific bins in order to accelerate the entropy encoding process with negligible loss of encoding efficiency.

[0059] In some embodiments, in addition to the prediction data 206 and quantized transform coefficients 216, the encoder may encode other information in the binary encoding stage 226, such as, for example, a prediction mode used in the prediction stage 204, parameters of the prediction operation, a type of transform in the transform stage 212, parameters of the quantization process (e.g., quantization parameters), an encoder control parameter (e.g., a bit rate control parameter), or the like. The encoder may use the output data from the binary encoding stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 may be further packaged for network transmission.

[0060] Referring to the reconstruction path of process 200A, at inverse quantization stage 218, the encoder can perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. At inverse transform stage 220, the encoder can generate reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder can add reconstructed residual BPU 222 to the predicted BPU 208 to generate prediction reference 224, which is used in the next iteration of process 200A.

[0061] It should be noted that other variations of process 200A can be used to encode video sequence 202. In some embodiments, the stages of process 200A can be performed by the encoder in different orders. In some embodiments, one or more stages of process 200A can be combined into a single stage. Petition 870250092909, dated 10 / 10 / 2025, page 27 / 100 21 / 57 stage. In some embodiments, a single stage of process 200A may be divided into multiple stages. For example, the transform stage 212 and the quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may omit one or more stages in FIG. 2A.

[0062] FIG. 2B illustrates a schematic diagram of another exemplary coding process 200B, consistent with the modalities of the description. Process 200B can be modified from process 200A. For example, process 200B can be used by an encoder to conform to a hybrid video coding standard (e.g., H.26x series). Compared with process 200A, the forward path of process 200B additionally includes a mode decision stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes an in-circuit filter stage 232 and a buffer 234.

[0063] Generally, prediction techniques can be categorized into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-image prediction or “intraprediction”) can use pixels from one or more surrounding BPUs already encoded in the same image to predict the current BPU. That is, the 224 prediction reference in spatial prediction can include the surrounding BPUs. Spatial prediction can reduce the inherent spatial redundancy of the image. Temporal prediction (e.g., inter-image prediction or “interprediction”) can use regions from one or more already encoded images to predict the current BPU. That is, the 224 prediction reference in temporal prediction can include the encoded images. Temporal prediction can reduce the inherent temporal redundancy of the images. Petition 870250092909, dated 10 / 10 / 2025, page 28 / 100 22 / 57

[0064] Referring to process 200B, in the forward path, the encoder performs the prediction operation in the spatial prediction stage 2042 and in the temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder may perform intraprediction. For an original BPU of an image being encoded, the prediction reference 224 may include one or more surrounding BPUs that were encoded (in the forward path) and reconstructed (in the reconstructed path) in the same image. The encoder may generate predicted BPU 208 by extrapolating the surrounding BPUs. The extrapolation technique may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, or similar. In some embodiments, the encoder may perform pixel-level extrapolation, such as extrapolating corresponding pixel values ​​for each pixel of the predicted BPU 208.The surrounding BPUs used for extrapolation can be located relative to the original BPU from various directions, such as in a vertical direction (e.g., at the top of the original BPU), a horizontal direction (e.g., to the left of the original BPU), a diagonal direction (e.g., below and to the left, below and to the right, above and to the left, or above and to the right of the original BPU), or any direction defined in the video encoding standard used. For intraprediction, the 206 prediction data may include, for example, locations (e.g., coordinates) of the surrounding BPUs used, sizes of the surrounding BPUs used, extrapolation parameters, a direction of the surrounding BPUs used relative to the original BPU, or similar.

[0065] For example, in the temporal prediction stage 2044, the encoder can perform interprediction. For an original BPU of a current image, the prediction reference 224 can include one or more images (referenced as “reference images”) that were encoded (in the forward path) and reconstructed (in the reconstructed path). Petition 870250092909, dated 10 / 10 / 2025, page 29 / 100 23 / 57 In some embodiments, a reference image can be encoded and the BPU reconstructed from the BPU. For example, the encoder can add reconstructed residual BPU 222 to predicted BPU 208 to generate a reconstructed BPU. When all reconstructed BPUs from the same image are generated, the encoder can generate a reconstructed image as a reference image. The encoder can perform a “motion estimation” operation to search for a matching region within a scope (referred to as a “search window”) of the reference image. The location of the search window in the reference image can be determined based on the location of the original BPU in the current image. For example, the search window can be centered at a location with the same coordinates in the reference image as the original BPU in the current image and can be extended to a predetermined distance.When the encoder identifies (for example, using a pel-recursive algorithm, a block-matching algorithm, or similar) a region similar to the original BPU in the search window, the encoder can designate such a region as the matching region. The matching region may have different dimensions (for example, be smaller than, equal to, larger than, or in a different format) than the original BPU. Since the reference image and the current image are temporally separated on the timeline (for example, as shown in FIG. 1), the matching region can be considered to “move” to the location of the original BPU as time passes. The encoder can record the direction and distance of such movement as a “motion vector”. When multiple reference images are used (for example, an image 106 in FIG.1) The encoder can search for a matching region and determine its associated motion vector for each reference image. In some modalities, the encoder... Petition 870250092909, dated 10 / 10 / 2025, page 30 / 100 24 / 57 can assign weights to the pixel values ​​of the matching regions of the respective corresponding reference images.

[0066] Motion estimation can be used to identify various types of motion, such as, for example, translations, rotations, zooms, or the like. For interprediction, prediction data 206 may include, for example, locations (e.g., coordinates) of the matching region, motion vectors associated with the matching region, the number of reference images, weights associated with the reference images, or the like.

[0067] To generate the predicted BPU 208, the encoder can perform a “motion compensation” operation. Motion compensation can be used to reconstruct the predicted BPU 208 based on prediction data 206 (e.g., the motion vector) and the prediction reference 224. For example, the encoder can move the matching region of the reference image according to the motion vector, whereby the encoder can predict the original BPU of the current image. When multiple reference images are used (e.g., one image 106 in FIG. 1), the encoder can move the matching regions of the reference images according to the respective motion vectors and average pixel values ​​of the matching regions.In some modes, if the encoder has weights assigned to the pixel values ​​of the matching regions of the respective corresponding reference images, the encoder may add a weighted sum of pixel values ​​from the moved matching regions.

[0068] In some modalities, interprediction can be unidirectional or bidirectional. Unidirectional interpredictions can use one or more reference images in the same temporal direction with respect to the current image. For example, image 104 in FIG. 1 is a unidirectionally interpreted image, in which the reference image (by Petition 870250092909, dated 10 / 10 / 2025, p. 31 / 100 25 / 57 example, image 102) precedes image 104. Bidirectional interpredictions can use one or more reference images in both temporal directions relative to the current image. For example, image 106 in FIG. 1 is a bidirectionally interpreted image, in which the reference images (e.g., images 104 and 108) are in both temporal directions relative to image 104.

[0069] Still with reference to the path of progress of the process 200B, after the spatial prediction stage 2042 and temporal prediction stage 2044, in the mode decision stage 230, the encoder can select a prediction mode (e.g., one from intraprediction or interprediction) for the current iteration of process 200B. For example, the encoder can perform a rate distortion optimization technique, in which the encoder can select a prediction mode to minimize a value of a cost function depending on the bit rate of a candidate prediction mode and distortion of the reconstructed reference image under the candidate prediction mode. Depending on the selected prediction mode, the encoder can generate the predicted BPU 208 and predicted data 206.

[0070] In the reconstruction path of process 200B, if the intraprediction mode was selected in the advance path, after generating prediction reference 224 (e.g., the current BPU that was encoded and reconstructed in the current image), the encoder can directly feed prediction reference 224 to spatial prediction stage 2042 for later use (e.g., for extrapolation of a next BPU from the current image). The encoder can feed prediction reference 224 to circuit filter stage 232, where the encoder can apply a circuit filter to prediction reference 224 to reduce or eliminate distortion (e.g., blocking artifacts) introduced during the encoding of prediction reference 224. The encoder can apply various circuit filter techniques in the stage of Petition 870250092909, dated 10 / 10 / 2025, page 32 / 100 26 / 57 circuit filter 232, such as, for example, unblocking, adaptive sample compensations (SAO), adaptive circuit filters (ALF), or the like. The circuit-filtered reference image may be stored in a buffer 234 (or “decoded image buffer”) for later use (e.g., to be used as an interprediction reference image for a future video sequence image 202). The encoder may store one or more reference images in a buffer 234 to be used in the temporal prediction stage 2044. In some embodiments, the encoder may encode circuit filter parameters (e.g., a circuit filter intensity) in the binary encoding stage 226, along with quantized transform coefficients 216, prediction data 206, and other information.

[0071] FIG. 3A illustrates a schematic diagram of an exemplary decoding process 300A, consistent with the modalities of the description. Process 300A may be a decompression process corresponding to compression process 200A in FIG. 2A. In some modalities, process 300A may be similar to the reconstruction path of process 200A. A decoder may decode the bitstream of videos 228 into video transmission 304 according to process 300A. Video transmission 304 may be very similar to the video sequence 202. However, due to information loss in the compression and decompression process (e.g., quantization stage 214 in FIGS. 2A-2B), video transmission 304 is generally not identical to the video sequence 202. Similar to processes 200A and 200B in FIGS. In versions 2A-2B, the decoder can execute process 300A at the basic processing unit (BPU) level for each image encoded in the 228-bit video stream.For example, the decoder can execute process 300A iteratively, where the decoder can decode a basic unit. Petition 870250092909, dated 10 / 10 / 2025, page 33 / 100 27 / 57 of processing in a 300A process iteration. In some embodiments, the decoder can execute the 300A process in parallel for regions (e.g., regions 114-118) of each image encoded in the 228 bitstream.

[0072] In FIG. In version 3A, the decoder can feed a portion of the video bitstream 228 associated with a basic processing unit (referred to as a “coded BPU”) of an image encoded in binary decoding stage 302. In binary decoding stage 302, the decoder can decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder can feed the quantized transform coefficients 216 into inverse quantization stage 218 and inverse transform stage 220 to generate reconstructed residual BPU 222. The decoder can feed prediction data 206 into prediction stage 204 to generate predicted BPU 208. The decoder can add reconstructed residual BPU 222 to predicted BPU 208 to generate predicted reference 224. In some embodiments, the predicted reference 224 can be stored in a buffer (e.g., an image buffer). (decoded in a computer memory).The decoder can feed predicted reference 224 into prediction stage 204 to perform a prediction operation in the next process iteration 300A.

[0073] The decoder can execute process 300A iteratively to decode each encoded BPU of the encoded image and generate predicted reference 224 to decode the next encoded BPU of the encoded image. After decoding all encoded BPUs of the encoded image, the decoder can send the image to the video transmission 304 to display and proceed with decoding on the next encoded image in the video bitstream 228.

[0074] In the binary decoding stage 302, the decoder can perform an inverse operation of the binary encoding technique. Petition 870250092909, dated 10 / 10 / 2025, page 34 / 100 28 / 57 used by the encoder (e.g., CABAC, entropy coding, variable-length coding, arithmetic coding, Huffman coding, or any other lossless data compression algorithm). In some embodiments, in addition to the prediction data 206 and quantized transform coefficients 216, the decoder may decode other information in the binary decoding stage 302, such as, for example, a prediction mode, prediction operation parameters, a transform type, quantization process parameters (e.g., quantization parameters), an encoder control parameter (e.g., a bit rate control parameter), or the like. In some embodiments, if the video bitstream 228 is transmitted over a network in packets, the decoder may unpack the video bitstream 228 before feeding it into the binary decoding stage 302.

[0075] FIG. 3B illustrates a schematic diagram of another exemplary decoding process 300B, consistent with the modalities of the description. Process 300B can be modified from process 300A. For example, process 300B can be used by a decoder to conform to a hybrid video coding standard (e.g., H.26x series). Compared with process 300A, process 300B additionally divides the prediction stage 204 into spatial prediction stage 2042 and temporal prediction stage 2044, and additionally includes the circuit filter stage 232 and buffer stage 234.

[0076] In process 300B, for a basic encoded processing unit (referred to as a “current BPU”) of an encoded image (referred to as a “current image”) being decoded, the prediction data 206 decoded by the binary decoding stage 302 by the decoder may include various data types, depending on which prediction mode was used to encode the BPU. Petition 870250092909, dated 10 / 10 / 2025, page 35 / 100 29 / 57 current from the coder. For example, if intraprediction was used by the coder to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., an indicative value) of the intraprediction, parameters of the intraprediction operation, or similar. The parameters of the intraprediction operation may include, for example, locations (e.g., coordinates) of one or more surrounding BPUs used as a reference, sizes of the surrounding BPUs, extrapolation parameters, a direction of the surrounding BPUs with respect to the original BPU, or similar. In another example, if interprediction was used by the coder to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., an indicative value) of the interprediction, parameters of the interprediction operation, or similar.The parameters of the interprediction operation may include, for example, the number of reference images associated with the current BPU, weights respectively associated with the reference images, locations (e.g., coordinates) of one or more matching regions in the respective reference images, one or more motion vectors respectively associated with the matching regions, or similar.

[0077] Based on the prediction mode indicator, the decoder can decide whether to perform a spatial prediction (e.g., intraprediction) at spatial prediction stage 2042 or a temporal prediction (e.g., interprediction) at temporal prediction stage 2044. The details of performing such a spatial or temporal prediction are described in FIG. 2B and will not be repeated herein. After performing such a spatial or temporal prediction, the decoder can generate the predicted BPU 208. The decoder can add the predicted BPU 208 and the reconstructed residual BPU 222 to generate prediction reference 224, as described in FIG. 3A. Petition 870250092909, dated 10 / 10 / 2025, page 36 / 100 30 / 57

[0078] In process 300B, the decoder can feed the predicted reference 224 to spatial prediction stage 2042 or temporal prediction stage 2044 to perform a prediction operation in the next iteration of process 300B. For example, if the current BPU is decoded using intraprediction in spatial prediction stage 2042, after generating prediction reference 224 (e.g., the decoded current BPU), the decoder can directly feed prediction reference 224 into spatial prediction stage 2042 for later use (e.g., for extrapolation of a next BPU from the current image).If the current BPU is decoded using interprediction in the time prediction stage 2044, after generating the prediction reference 224 (e.g., a reference image in which all BPUs have been decoded), the decoder can feed the prediction reference 224 into the circuit filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder can apply a circuit filter to the prediction reference 224, in a manner described in FIG. 2B. The circuit-filtered reference image can be stored in buffer 234 (e.g., decoded image buffer in computer memory) for later use (e.g., to be used as an interprediction reference image for a future encoded video bitstream image 228). The decoder can store one or more reference images in buffer 234 for use in time prediction stage 2044.In some modes, the prediction data may additionally include circuit filter parameters (e.g., a circuit filter intensity). In some modes, the prediction data includes circuit filter parameters when the prediction mode indicator of prediction data 206 indicates that interprediction was used to encode the current BPU. Petition 870250092909, dated 10 / 10 / 2025, page 37 / 100 31 / 57

[0079] FIG. 4 is a block diagram of an example device 400 for encoding or decoding video, consistent with the embodiments described. As shown in FIG. 4, the device 400 may include the processor 402. When the processor 402 executes the instructions described here, the device 400 may become a specialized machine for encoding or decoding videos. The processor 402 may be any type of circuit capable of handling or processing information.For example, the 402 processor may include any combination of any number of a central processing unit (or “CPU”), a graphics processing unit (or “GPU”), a neural processing unit (“NPU”), a microcontroller unit (“MCU”), an optical processor, a programmable logic controller, a microcontroller, a microprocessor, a digital signal processor, an intellectual property (IP) core, a Programmable Logic Array (PLA), a programmable logic array (PAL), a Generic Logic Array (GAL), a Programmable Complex Logic Device (CPLD), a Field Programmable Gate Array (FPGA), a System on a Chip (SoC), an Application-Specific Integrated Circuit (ASIC), or the like. In some embodiments, the 402 processor may also be an array of processors as a single logic component. For example, as shown in FIG.4. The 402 processor may include multiple processors, including the 402a processor, the 402b processor, and the 402n processor.

[0080] Device 400 may also include memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, or the like). For example, as shown in FIG. 4, the stored data may include program instructions (e.g., program instructions to implement the stages in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequences). Petition 870250092909, dated 10 / 10 / 2025, page 38 / 100 32 / 57 202, video bitstream 228, or video transmission 304). The 402 processor can access program instructions and data for processing (e.g., via bus 410), and execute program instructions to perform an operation or manipulation on the data for processing. The 404 memory may include a high-speed random-access storage device or a non-volatile storage device. In some embodiments, the 404 memory may include any combination of any number of random-access memory (RAM), read-only memory (ROM), optical disk, magnetic disk, hard disk, solid-state drive, flash drive, card, digital secure card (SD), memory stick, compact flash card (CF), or similar. The 404 memory may also be a group of memories (not shown in FIG. 4) grouped as a single logical component.

[0081] The 410 bus can be a communication device that transfers data between components within the 400 appliance, such as an internal bus (e.g., a CPU-memory bus), an external bus (e.g., a universal serial bus port, a peripheral component interconnect express port), or the like.

[0082] To facilitate explanation without causing ambiguity, the 402 processor and other data processing circuits are collectively referred to as a “data processing circuit” in this description. The data processing circuit may be implemented entirely as hardware, or as a combination of software, hardware, or firmware. Furthermore, the data processing circuit may be a single standalone module or may be combined wholly or partially into any other component of the 400 appliance. Petition 870250092909, dated 10 / 10 / 2025, page 39 / 100 33 / 57

[0083] The 400 device may additionally include the 406 network interface to provide wired or wireless communication with a network (for example, the Internet, an intranet, a local area network, a mobile communications network, or the like). In some embodiments, the 406 network interface may include any combination of any number of a network interface controller (NIC), a radio frequency (RF) module, a transponder, a transceiver, a modem, a router, a gateway, a wired network adapter, a wireless network adapter, a Bluetooth adapter, an infrared adapter, a near-field communication (“NFC”) adapter, a cellular network chip, or the like.

[0084] In some embodiments, the 400 apparatus may optionally include a peripheral interface 408 to provide a connection for one or more peripheral devices. As shown in FIG. 4, the peripheral device may include, but is not limited to, a cursor control device (e.g., a mouse, a touchpad, or a touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light-emitting diode display), a video input device (e.g., a camera or an input interface coupled to a video file), or the like.

[0085] It should be noted that video codecs (e.g., a codec executing a 200A, 200B, 300A, or 300B process) can be implemented as any combination of any software or hardware modules in the 400 device. For example, some or all stages of the 200A, 200B, 300A, or 300B process can be implemented as one or more modules of the 400 device, such as program instructions that can be loaded into memory 404. For another example, some or all stages of the 200A, 200B, 300A, or 300B process can be implemented as one or more modules of Petition 870250092909, dated 10 / 10 / 2025, page 40 / 100 34 / 57 device hardware 400, such as a specialized data processing circuit (e.g., an FPGA, an ASIC, an NPU, or similar).

[0086] The present invention provides methods for initializing context model probabilities used in CABAC. FIG. 5 is a schematic diagram illustrating a context-adaptive binary arithmetic coding (CABAC) mechanism 500, according to some embodiments of the present invention. Consistent with the disclosed embodiment, the CABAC mechanism 500 can be used by an encoder to conduct binary coding. For example, the CABAC mechanism 500 can be used in the binary coding stage 226 in FIG. 2A or FIG. 2B.

[0087] With reference to FIG. 5, the CABAC 500 mechanism includes three elementary stages: binarization 502, context modeling (CM) 504 and binary arithmetic encoding (BAE) 506.

[0088] 502 binarization is a data preprocessing procedure. In 502 binarization, non-binary syntax elements are encoded into a string of binary symbols called “bins”. An individual binary symbol is simply referred to as a bin. Several methods can be used to binarize input syntax elements, such as table mapping, unary encoding, truncated unary encoding, fixed-length encoding, k-th order unary exponential Golomb encoding (UEG-k), etc.

[0089] As an example, UEG-k can be used to binarize motion vector differences as follows. In this example, the mvd value of a motion vector component is assumed to be given. For the prefix part of the UEG-k binary string, you use a TU binarization with a cutoff value of S=9. If mvd equals zero, the binary string includes only the prefix codeword “0”. If the condition | mvd | > 9 remains, the suffix is ​​constructed as a word. Petition 870250092909, dated 10 / 10 / 2025, p. 41 / 100 35 / 57 code EG3 for the value of |mvd| - 9, whose mvd sign is appended using the sign bit “1” for a negative mvd and the sign bit “0” for the remainder. For mvd values ​​with 0 < |mvd| < 9, the suffix includes only the sign bit. It is possible to note that the component of a motion vector difference represents the prediction error with a quarter-sample accuracy, the prefix part corresponds to a maximum error component of ±2 samples. With the selection of the Exp-Golomb k = 3 parameter, the suffix code words are given, so that a geometric increase in the prediction error in the units of 2 samples is captured by a linear increase in the length of the corresponding suffix code word.

[0090] The UEG-k binarization of the absolute values ​​of the transform coefficient levels (abs_level) is specified by the cutoff value S = 14 for the TU prefix part and the order k = 0 for the EGk suffix part. Note that the binarization and subsequent encoding process are applied to the syntax element coeff_abs_value_minus1 = abs_level - 1, since transform coefficient levels with a value of zero are encoded using a significance map. Constructing a binary string for a given value of coeff_abs_value_minus1 is similar to constructing UEG-k binary strings for motion vector difference components, except that there is no sign attached to the suffix. The table in FIG. 6 shows the corresponding binary strings for abs_level values ​​from 1 to 20, where the prefix parts are highlighted in grayed-out columns.

[0091] Referring again to FIG. 5, context-based modeling 504 calculates a context index (ctxIdx) for each bin. This context index is used to find a context model that is stored as the probability state tables. These tables are updated with each bin and reset at the beginning of each layer.

[0092] For example, a number (for example, 399) of models of Petition 870250092909, dated 10 / 10 / 2025, page 42 / 100 36 / 57 contexts can be stored in the encoder. A context index (ctxIdx) is used to track the sum of the context index offset (ctxIdxOffset) and the context index increment (ctxIdxInc). The exception to the above implementation is the calculation of the context index for residual syntax elements, where it is the sum of ctxIdxOffset, ctxIdxInc, and the context block category offset (ctxBlockCatOffset). The ctxBlockCatOffset depends on the context block category of the macroblock currently being encoded. The ctxIdxOffset is determined by the syntax element type and layer type. The ctxIdxInc differs for each bin of the encoded syntax element, therefore it is dependent on the bin index or bin index (binIdx). The calculation of ctxIdxInc is dependent on surrounding information in some cases.For residual syntax elements, the calculation of ctxIdxInc also depends on the scan position of the current element being encoded and the number of previously encoded coefficients.

[0093] Each bin value and its corresponding ctxIdx is sent to the binary arithmetic encoder module 506. The binary arithmetic encoder module 506 stores information such as the most probable symbol (MPS) and the probability of that state. This consists of context information that can be accessed with the ctxIdx. In the binary arithmetic encoder module 506, there are two possible symbols, namely 0 and 1. If one of the symbols is the most probable symbol, then the other symbol becomes the least probable symbol (LPS). A context memory or context table can be used to store context information consisting of the probability state index and the most probable state value (e.g., ranging from 0-399).

[0094] The regular 508 encoding engine performs arithmetic encoding, in which an encoding range is set and updated based on the probability of MPS and LPS. The codeword of Petition 870250092909, dated 10 / 10 / 2025, page 43 / 100 37 / 57 arithmetic encoding is generated from the recursive division of the interval. Two variables are used to keep track of the interval: the variable Low and the variable Range, which are referenced as codILow and codIRange, respectively. FIG. 7 shows the values ​​of Range and Low, and when they are updated.

[0095] The initial value of Range is 510, and it is a 9-bit register. The initial value of Low is 0, and it is a 10-bit register. rMPS and rLPS represent the two corresponding subranges of MPS and LPS, respectively. If the input bin is equal to MPS, rMPS is chosen as the new range; otherwise, rLPS is selected. When it is found that the updated Range is outside the range 256 and 511 inclusive, the renormalization procedure is employed. The renormalization procedure is where most of the codeword is constructed.

[0096] The probability state Plps is required to compute the value of rLPS. This value ranges from 0 to 0.5 and is quantized into 64 discrete probability states. These states are indexed by a variable pStateIdx ranging from 0 to 63. The transition to the next state based on the current bin is shown in FIG. 8. The lookup table is used to update the probability state.

[0097] As another multiplication operation in this state, the calculation of rLPS = Range * Rlps is also converted into a lookup table. The Range is quantized into four Rq values. The rLPS product is also quantized into 256 values ​​based on Rq and pStateIdx. As a result, the computation of rLPS can be done simply by looking up a two-dimensional table, where Rq and pStateIdx are the two indices.

[0098] Two other encoding methods are also used in binary arithmetic encoding. In the 510 bypass encoding mechanism, the context shaping stage is skipped. This means that the previous Range and Low values ​​are used and renormalization to Petition 870250092909, dated 10 / 10 / 2025, page 44 / 100 38 / 57 the bypass method is invoked. In the bypass encoding mechanism, the probability of the two symbols is considered to be equal to 0.5.

[0099] In addition, a termination coding mechanism (not shown in FIG. 5) can be invoked when the end of the layer syntax element is encountered or when the mb type is of the IPCM variety. At this stage, no context model is chosen. LPS is set to 1 and rLPS is set to 2. Otherwise, the renormalization is the same. When the end of a layer is encountered, then a cleaning algorithm is also invoked.

[0100] Every time a new layer starts, an initialization algorithm is invoked that resets all 399 context models, based on the layer type and the cabac init idc value. The layer's QP variable is used to calculate the exact context.

[0101] The CABAC 500 mechanism above (FIG. 5) is described in connection with an encoder. It is understood that the inverse operations of the CABAC 500 mechanism can be performed in the binary decoding stage 302 (FIG. 3A or FIG. 3B) of a decoder.

[0102] In VVC, the probability of each context model in the context-adaptive binary arithmetic coding (CABAC) mechanism is initialized according to the layer type (e.g., layer I, layer P, layer B). When encoding or decoding a bin, the context model probability is updated. To improve the accuracy of the probability estimate, a multi-hypothesis probability update model is supported. Two probabilities are associated with each context model and are updated independently, with different adaptation rates. The adaptation rates of the probabilities for each context model are pre-trained based on the statistics of the associated bins. The probability used to encode a bin is the average of the estimates of the two hypotheses. Petition 870250092909, dated 10 / 10 / 2025, page 45 / 100 39 / 57

[0103] In the Enhanced Compression Model (ECM), which is used to further improve the encoding performance of the VVC standard, the multi-hypothesis probability update model is further improved. The adaptation rates of the two probabilities associated with each context model are different for each layer type and are initialized according to the layer type. Furthermore, to improve the accuracy of the probability estimate for statistical variations in different regions, the adaptation rates are adjusted by two delta parameters in a context lookup table and retrieved by a previously encoded bin used as an index. The previously encoded bin is used as an index to obtain the adjustment parameters from a lookup table. In addition, the simple calculation of averaging the two probabilities is extended to weighted average calculation.Three different sets of weights are predetermined for each context model in the I, B, and P layer types.

[0104] In ECM, for interlayers (e.g., B and P layers), the context model probability can be predicted from previously encoded images that have the same layer type, QP, and temporal ID.

[0105] In the current ECM design, the probability of each context model is initialized according to the layer type or is predicted from previously encoded images. However, this may not be ideal since each image may have different contents. The present invention provides methods to solve this problem.

[0106] In some embodiments, it is proposed to independently select a context model probability set from a plurality of context model probability sets for each layer. The plurality of context model probability sets contains N sets, where N is an integer. Petition 870250092909, dated 10 / 10 / 2025, p. 46 / 100 40 / 57 positive. A context model probability set can contain any subset of {initial probability, adaptation rates, weights applied to two probabilities, and adjustment parameters}. For each layer, selecting a context model probability set from the plurality of context model probability sets can be based on layer type, QP, temporal ID, low delay condition, rate cost, etc. The selection can be derived in both the encoder and decoder, without signaling. The selection can also be signaled in the bitstream.

[0107] As an example, FIG. 9 shows a set of context model probability parameters 902, which contains initial probabilities, adaptation rates and weights applied to the probabilities that are associated with the two context models.

[0108] As another example, FIG. 10 shows a set of context model probability parameters 1002, which contains initial probability, adaptation rates, weights applied to the two probabilities, and adjustment parameters associated with each context model. Note that the adjustment parameters can be shared among the plurality of context model probability sets.

[0109] As another example, FIG. 11 shows four sets of predefined context model probability parameters including a first set of context model probability parameters 1102 for layer B and non-low delay condition, a second set of context model probability parameters 1104 for layer P, a third set of context model probability parameters 1106 for layer I, and a fourth set of context model probability parameters 1108 for layer B and low delay condition. For each layer, the context model probability is selected according to its type. Petition 870250092909, dated 10 / 10 / 2025, page 47 / 100 Layer 41 / 57 and low delay condition. This selection is derived in both the encoder and decoder without signaling in the bitstream.

[0110] As another example, X+1 sets of context model probability parameters are defined, where X is determined based on the temporal ID number. A first set of context model probability parameters for non-interlayer. A second set of context model probability parameters for interlayer with temporal ID equal to 0. A third set of context model probability parameters for interlayer with temporal ID equal to 1. Similarly, a k° set of context model probability parameters for interlayer with temporal ID equal to k-2. For each interlayer, a set of context model probability parameters is selected according to the temporal ID, where for non-interlayer, the first set of context model probability parameters is always used.

[0111] As another example, there are five sets of context model probability parameters including a first, a second, a third, and a fourth set predefined for layer I, layer P, layer B with non-low delay condition, and layer B with low delay condition, respectively. The fifth set is predicted from previously encoded images. For layer I, the first set of context model probability parameters is used, while for non-layer I (i.e., interlayer), it is selected from the second, third, fourth, and fifth sets. The selection is based on the rate cost. Rate costs are calculated using each set of context model probability parameters, and the set with the minimum rate cost is used. A parameter for interlayer is signaled in the bitstream to indicate Petition 870250092909, dated 10 / 10 / 2025, page 48 / 100 42 / 57 car which set of context model probability parameters is selected.

[0112] Similarly, in another example, there are five sets of context model probability parameters including a first, a second, a third, and a fourth set predefined for layer I, layer P, layer B with non-low delay condition, and layer B with low delay condition, respectively. The fifth set is predicted from previously encoded images. For interlayer, it can be selected when to use the fifth set predicted from previously encoded images. If it is determined not to use the fifth set, the set is selected according to the layer type and low delay condition.

[0113] In some embodiments, it is proposed to select a set of context model probability parameters from a plurality of context model probability parameter sets associated with each video sequence. The plurality of context model probability parameter sets contains N sets, where N is a positive integer. A set of context model probability parameters may contain any subset of {initial probability, adaptation rates, weights applied to two probabilities, and adjustment parameters}. The selection may be signaled in a bitstream associated with the video sequence.

[0114] As an example, four sets of context model probability parameters are predefined. The four sets include: a first set of context model probability parameters for layer B, a second set of context model probability parameters for layer P, a third set of context model probability parameters for the Petition 870250092909, dated 10 / 10 / 2025, page 49 / 100 43 / 57 Layer I, and a fourth set of context model probability parameters for Layer B. Each set of context model probability parameters contains initial probability, adaptation rates, and weights applied to the two probabilities. The adjustment parameters are shared among (i.e., are equal for) the four sets of context model probability parameters. For each layer, the corresponding context model is determined according to its layer type. Additionally, an indicator can be signaled at the SPS level in the bitstream to indicate which of the first and fourth sets of context model probability parameters is used for Layer B.

[0115] Similarly, in another example, each set of context model probability parameters contains initial probability, adaptation rates, weights applied to the two probabilities, and adjustment parameters.

[0116] Similarly, in another example, the level indicator of The SPS indicator in the bitstream indicates which of the first or fourth set of context model probability parameters is used for Layer B and is determined according to the low-delay condition in the encoder. If it is determined as a low-delay condition, the fourth set of context model probability parameters is used to initialize the Layer B context model, and the SPS level indicator is coded to indicate that the fourth set of context model probability parameters is selected. Otherwise, if it is determined as a non-low-delay condition, the first set of context model probability parameters is used to initialize the Layer B context model, and the SPS level indicator is coded to indicate that the first set of context model probability parameters is selected. Petition 870250092909, dated 10 / 10 / 2025, p. 50 / 100 44 / 57

[0117] As another example, four sets of context model probability parameters are predefined. The four sets of context model probability parameters include: a first set of context model probability parameters for layer B, a second set of context model probability parameters for layer P, a third set of context model probability parameters for layer I, and a fourth set of context model probability parameters for layer B. Each set of context model probability parameters contains initial probability, adaptation rates, and weights applied to the two probabilities. The adjustment parameters are shared among (i.e., are equal for) the four sets of context model probability parameters. Furthermore, for layers B and P, it is determined when to predict their context models from previously encoded images.If the context models for layers B and P are not predicted from previously encoded images, the second set of context model probability parameters is used for layer P, and an SPS level indicator is signaled to indicate which of the first or fourth set of context model probability parameters is used for layer B. Otherwise, if the context models for layers B and P are predicted from previously encoded images, the predefined sets of context model probability parameters are not used. In contrast to layers B and P, layers I always use the third set of context model probability parameters, without relying on previously encoded images.

[0118] In some modalities, in addition to or as an alternative to the SPS level indicator indicating that one of the first set or the fourth set of model probability parameters Petition 870250092909, dated 10 / 10 / 2025, p. 51 / 100 Context 45 / 57 is used for layer B; an indicator can be signaled at the PPS level, image header, or layer header of a bitstream to indicate which of the first set or fourth set of context model probability parameters is used for layer B.

[0119] In some embodiments, there are N sets of predefined context model probability parameters for layer type I, layer type P, and layer type B. Three parameters are flagged to indicate which of the N sets of context model probability parameters is / are used for layer type I, P, and B, respectively.

[0120] In some modalities, there are two sets of predefined context model probability parameters for each layer type I, layer type P, and layer type B. For each layer type, an indicator is signaled at the SPS level to indicate which of the two associated sets, in relation to the respective layer type, is used.

[0121] The methods described above for initializing the probability parameters of context models can be implemented in an encoder and a decoder. FIG. 12 illustrates a flowchart of an exemplary method 1200 for decoding a bitstream associated with a video, according to some embodiments of the present invention. Method 1200 can be executed by a decoder (e.g., by process 300A of FIG. 3A or 300B of FIG. 3B) or executed by one or more software or hardware components of an apparatus (e.g., apparatus 400 of FIG. 4), when executing CABAC. For example, one or more processors (e.g., processor 402 of FIG. 4) can execute method 1200. In some embodiments, method 1200 can be implemented by a computer program product, embedded in a computer-readable medium, including Petition 870250092909, dated 10 / 10 / 2025, page 52 / 100 46 / 57 computer executable instructions, such as program code, executed by computers (e.g., device 400 of FIG. 4). As shown in FIG. 12, method 1200 includes the following steps 1210-1220.

[0122] In step 1210, a processor (e.g., processor 402 in FIG. 4) selects, for a first layer, a first set of probability parameters to start one or more context models used in CABAC. The first layer can be a current layer that the processor is decoding. The first set of probability parameters can be selected from a plurality of predefined sets of probability parameters.

[0123] The first set of probability parameters selected can be used to initialize the context models. For example, two context models can be used to run CABAC. Before running CABAC from a new layer, the processor needs to initialize the probability parameters used by the two context models.

[0124] FIG. 9 provides an example of a probability parameter set. As shown in FIG. 9, the probability parameter set 902 may include one or more of: an initial probability for use in one or more context models, an adaptation rate of one or more context models, a plurality of weights that is associated, respectively, with a plurality of probabilities, or an adjusted probability for use in one or more context models. The initial probability is an initial probability value used for a context model in the CABAC of a layer. If two context models are used in the CABAC, the probability parameter set 902 may include two initial probabilities, one for each of the two context models. The adaptation rate Petition 870250092909, dated 10 / 10 / 2025, page 53 / 100 47 / 57 is the rate to adapt the probability state of a context model. Again, if two context models are used in CABAC, the 902 probability parameter set may include two adaptation rates, one for each of the two context models. Weights are used to weight the context models. If two context models are used in CABAC, the 902 probability parameter set may include two weights, one for each of the context models. Although the above description assumes that two context models are used in a single-layer CABAC, it is possible that more context models (e.g., three or four) are used. In that case, the 902 probability parameter set may include more initial probabilities, adaptation rates, and / or weights.

[0125] Furthermore, after initializing the CABAC of a layer, the initial probabilities of the context models can be adjusted. The probability parameter set 1002 in FIG. 10 additionally includes one or more adjustment parameters. An adjustment parameter is an adjusted probability that can be used for a context model after the model initialization.

[0126] In some embodiments, the processor can select the first set of probability parameters from a plurality of predefined sets of probability parameters, without requiring explicit signaling. For example, the selection can be based on one or more layer types, a quantization parameter (QP), a temporal identifier, a low delay condition, or a cost rate associated with the first layer. Examples for performing selection based on these parameters are given above, which are not repeated in this document.

[0127] Referring again to FIG. 12, in step 1220, the processor performs the entropy decoding of the first layer based Petition 870250092909, dated 10 / 10 / 2025, page 54 / 100 48 / 57 in one or more context models and in the first set of probability parameters.

[0128] In some embodiments, the selection of the initial probability parameter set can be signaled in a bitstream. FIG. 13 illustrates a flowchart of an exemplary method 1300 for decoding a bitstream associated with a video, according to some embodiments of the present invention. The method 1300 can be executed by a decoder (e.g., by process 300A of FIG. 3A or 300B of FIG. 3B) or executed by one or more software or hardware components of a device (e.g., device 400 of FIG. 4), when executing CABAC. For example, one or more processors (e.g., processor 402 of FIG. 4) can execute method 1300. In some embodiments, method 1300 can be implemented by a computer program product, embedded in a computer-readable medium, including computer-executable instructions, such as program code, executed by computers (e.g., device 400 of FIG. 4). As shown in FIG.13, method 1300 includes the following steps 1310-1320.

[0129] In step 1310, a processor (e.g., processor 402 in FIG. 4) selects, based on an indicator or a signaled parameter in a bitstream, a first set of probability parameters from a plurality of predefined sets of probability parameters. The first set of probability parameters is used to initialize the context model probabilities used in CABAC from a first layer.

[0130] For example, the indicator can be flagged in an SPS, a PPS, an image header, or a layer header. For example, if the indicator is flagged in an SPS, the set of probability parameters referenced by the indicator can be used by all layers associated with the SPS. Petition 870250092909, dated 10 / 10 / 2025, page 55 / 100 49 / 57

[0131] As another example, the indicator or parameter may have a value dependent on whether a low delay condition or a non-low delay condition is used to encode the first layer. For example, if the low delay condition is enabled, a value of the indicator or parameter (e.g., “2”) may refer to a first among a plurality of predefined probability parameter sets; and if the non-low delay condition is enabled, the same value of the indicator or parameter (e.g., “2”) may refer to a second among a plurality of predefined probability parameter sets.

[0132] In step 1320, the processor performs first-layer entropy decoding based on one or more context models and the first set of probability parameters.

[0133] The encoding methods corresponding to the decoding methods described above can be performed by an encoder. For example, FIG. 14 illustrates a flowchart of an exemplary method 1400 for encoding a bitstream associated with a video, according to some embodiments of the present invention. Method 1400 can be performed by an encoder (for example, by process 200A of FIG. 2A or 200B of FIG. 2B) or performed by one or more software or hardware components of a device (for example, device 400 of FIG. 4), when performing CABAC. For example, one or more processors (e.g., processor 402 of FIG. 4) can execute method 1400. In some embodiments, method 1400 can be implemented by a computer program product, embedded in a computer-readable medium, including computer-executable instructions, such as program code, executed by computers (e.g., device 400 of FIG. 4).As shown in FIG. 14, method 1400 includes the following steps 14101420. Petition 870250092909, dated 10 / 10 / 2025, page 56 / 100 50 / 57

[0134] In step 1410, a processor (e.g., processor 402 in FIG. 4) selects, for a first layer, a first set of probability parameters to initiate one or more context models used in CABAC. In some embodiments, the selection may be based on at least one layer type, a quantization parameter (QP), a temporal identifier, a low delay condition, or a cost rate associated with the first layer.

[0135] In step 1420, the processor performs first-layer entropy encoding based on one or more context models and the first set of probability parameters.

[0136] FIG. 15 illustrates a flowchart of an exemplary method 1500 for encoding a bitstream associated with a video, according to some embodiments of the present invention. As in method 1400, method 1500 can also be performed by a processor of an encoder. As shown in FIG. 15, method 1500 includes the following steps 1510-1520.

[0137] In step 1510, the processor encodes, in a bit stream, an indicator or a parameter indicating a first set of probability parameters. The indicator or parameter is associated with a layer, and signals that the first set of probability parameters is selected to execute CABAC of the first layer. The selection can be based on at least one layer type, a quantization parameter (QP), a temporal identifier, a low delay condition, or a cost rate associated with the first layer (step 1410 in FIG. 14).

[0138] Referring again to FIG. 15, in step 1520, the processor performs the first layer entropy encoding based on one or more context models and the first set of probability parameters. Petition 870250092909, dated 10 / 10 / 2025, p. 57 / 100 51 / 57

[0139] Note that the methods disclosed can be freely combined.

[0140] In some embodiments, a non-transient computer-readable storage medium is also provided that stores a bitstream. The context model probability set selected for a layer can be signaled in the bitstream.

[0141] In some embodiments, a non-transient computer-readable storage medium is also provided including instructions, and the instructions can be executed by a device (such as the disclosed encoder and decoder) to perform the methods described above.Common forms of non-transient media include, for example, a floppy disk, a floppy disk, a hard disk, a solid-state drive, magnetic tape or any other magnetic data storage media, a CD-ROM, any other optical data storage media, any other physical media with hole patterns, a RAM, a PROM and EPROM, a FLASH-EPROM or any other flash memory, NVRAM, a cache, a register, any other memory chip or cartridge, and network versions thereof. The device may include one or more processors (CPUs), an input / output interface, a network interface and / or memory.

[0142] The modalities can be further described using the following clauses:

[0143] 1. Method for decoding a bitstream associated with a sequence of videos, the method comprising: Select, from a plurality of predefined sets of probability parameters, a first set of probability parameters to initialize one or more context models for a layer B; and perform entropy decoding of the base layer B. Petition 870250092909, dated 10 / 10 / 2025, page 58 / 100 52 / 57 set in one or more context models and in the first set of probability parameters, where the selection is based on a layer B encoding condition or a signal in the bitstream.

[0144] 2. Method, according to clause 1, wherein the first set of probability parameters comprises at least one of: an initial probability for use with one or more context models, an adaptation rate of one or more context models, a plurality of weights that is associated, respectively, with a plurality of probabilities, or an adjusted probability for use with one or more context models.

[0145] 3. Method, according to clause 1, wherein the encoding condition comprises at least one of a quantization parameter (QP), a temporal identifier, a low delay condition, a non-low delay condition, or a rate cost associated with the first layer.

[0146] 4. Method, according to clause 1, in which the signal comprises an indicator or a parameter in the bit stream.

[0147] 5. Method, according to clause 4, wherein the selection is based on an indicator signaled in a sequence parameter set (SPS), an image parameter set (PPS), an image header or a layer header.

[0148] 6. Method, according to clause 4, in which the indicator or parameter has a value dependent on whether a low delay condition or a non-low delay condition is used to encode layer B. Petition 870250092909, dated 10 / 10 / 2025, page 59 / 100 53 / 57

[0149] 7. Method, according to clause 1, wherein the plurality of predefined sets of probability parameters comprises two predefined sets of probability parameters for layer B.

[0150] 8. Method for encoding a bitstream associated with a sequence of videos, the method comprising: Select, from a plurality of predefined sets of probability parameters, a first set of probability parameters to initiate one or more context models for a layer B; and perform layer B entropy coding based on one or more context models and the first set of probability parameters, wherein the selection is based on a layer B coding condition or a signal in the bitstream.

[0151] 9. Method, according to clause 8, wherein the first set of probability parameters comprises at least one of: an initial probability for use with one or more context models, an adaptation rate of one or more context models, a plurality of weights that is associated, respectively, with a plurality of probabilities, or an adjusted probability for use with one or more context models.

[0152] 10. Method, according to clause 8, wherein the encoding condition comprises at least one of a quantization parameter (QP), a time identifier, a delay condition Petition 870250092909, dated 10 / 10 / 2025, p. 60 / 100 54 / 57 low, a non-low delay condition, or a rate cost associated with the first tier.

[0153] 11. Method, according to clause 8, additionally comprising: To encode, in the bitstream, an indicator or parameter associated with layer B, the indicator or parameter indicating that the first set of probability parameters is selected.

[0154] 12. Method, according to clause 11, where the indicator is signaled in a sequence parameter set (SPS), an image parameter set (PPS), an image header or a layer header.

[0155] 13. Method, according to clause 11, in which the indicator or parameter has a value dependent on whether a low delay condition or a non-low delay condition is used to encode the first layer.

[0156] 14. Method, according to clause 8, wherein the plurality of predefined sets of probability parameters comprises two predefined sets of probability parameters for layer B.

[0157] 15. Non-transient, computer-readable storage media that stores a bitstream of a video for processing, according to: select, from a plurality of predefined sets of probability parameters, a first set of probability parameters to initialize one or more context models for a layer B; and perform entropy encoding or decoding of layer B based on one or more context models and the first set of probability parameters. Petition 870250092909, dated 10 / 10 / 2025, page 61 / 100 55 / 57 where the selection is based on a Layer B encoding condition or a signal in the bitstream.

[0158] 16. Non-transient computer-readable storage media, according to clause 15, wherein the first set of probability parameters comprises at least one of: an initial probability for use with one or more context models, a fitting rate of one or more context models, a plurality of weights that is associated, respectively, with a plurality of probabilities, or an adjusted probability for use with one or more context models.

[0159] 17. Non-transient computer-readable storage media, according to clause 15, wherein the encoding condition comprises at least one of a quantization parameter (QP), a temporal identifier, a low delay condition, a non-low delay condition, or a rate cost associated with the first layer.

[0160] 18. Non-transient computer-readable storage medium, according to clause 15, wherein the signal comprises an indicator or a parameter in the bitstream.

[0161] 19. Non-transient computer-readable storage media, according to clause 18, wherein the selection is based on a signaled indicator in a sequence parameter set (SPS), an image parameter set (PPS), an image header or a layer header.

[0162] 20. Non-transient computer-readable storage media, according to clause 18, wherein the indicator or parameter has a value dependent on whether a delay condition exists. Petition 870250092909, dated 10 / 10 / 2025, page 62 / 100 A 56 / 57 low or non-low delay condition is used to encode layer B.

[0163] 21. Non-transient computer-readable storage media, according to clause 15, wherein the plurality of predefined sets of probability parameters comprises two predefined sets of probability parameters for layer B.

[0164] It should be noted that relational terms in this document, such as “first” and “second,” are used only to differentiate one entity or operation from another entity or operation, and do not require or imply any actual relationship or sequence between those entities or operations. Furthermore, the words “comprising,” “having,” “containing,” and “including,” and other similar forms are intended to be equivalent in meaning and to be open-ended, so that an item or items following any of these words do not signify a comprehensive list of such item or items, or be limited only to the item or items listed.

[0165] As used in this document, except where specifically stated otherwise, the term “or” encompasses all possible combinations, except where it is impractical. For example, if it is stated that a database may include A or B, then, unless specifically stated otherwise or if it is impractical, the database may include A, or B, or A and B. As in the second example, if it is stated that a database may include A, B, or C, then, unless specifically stated otherwise or if it is impractical, the database may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.

[0166] It is appreciated that the modalities described above can be implemented by hardware, or software (program codes), or Petition 870250092909, dated 10 / 10 / 2025, p. 63 / 100 57 / 57 a combination of hardware and software. If implemented by software, they can be stored on computer-readable media as described above. The software, when executed by the processor, can perform the disclosed methods. The computing units and other functional units described in the present invention can be implemented by hardware, or software, or a combination of hardware and software. A person skilled in the art will also understand that multiple modules / units described above can be combined as one module / unit, and each of the modules / units described above can be further divided into a plurality of submodules / subunits.

[0167] In the preceding descriptive report, embodiments were described with reference to various specific details that may vary from implementation to implementation. Certain adaptations and modifications of the described embodiments may be made. Other embodiments may be apparent to those skilled in the art from consideration of the descriptive report and practice of the invention disclosed in this document. The descriptive report and examples are intended to be considered only as illustrative, with the true scope and spirit of the invention being indicated by the following claims. It is also intended that the sequences of steps shown in the figures are for illustrative purposes only and are not limited to any particular sequence of steps. Thus, those skilled in the art may understand that these steps can be performed in a different order when implementing the same method.

[0168] In the drawings and descriptive report, illustrative embodiments were disclosed. However, many variations and modifications can be made to these embodiments. Consequently, although specific terms are employed, they are used only in a generic and descriptive sense and not for the purpose of limitation. Petition 870250092909, dated 10 / 10 / 2025, page 64 / 100 1 / 57 METHODS FOR CONTEXT-BASED VIDEO CODING CROSS-REFERENCE TO RELATED ORDERS

[0001] The present application is based on and claims priority of U.S. Provisional Patent Application No. 63 / 496,012, filed April 13, 2023, U.S. Provisional Patent Application No. 63 / 618,884, filed January 8, 2024, and U.S. Patent Application No. 18 / 628,723, filed April 6, 2024, the contents of which are incorporated by reference herein in their entirety. TECHNICAL FIELD

[0002] The present invention relates generally to video processing and, more specifically, to methods and apparatus for initializing a set of context model probabilities for a layer in context-adaptive binary arithmetic coding (CABAC). BACKGROUND

[0003] A video is a collection of still images (or “frames”) that capture visual information. To reduce storage memory and transmission bandwidth, a video can be compressed before storage or transmission, and decompressed before display. The compression process is generally referred to as encoding, and the decompression process is generally referred to as decoding. There are several video encoding formats that use standardized video encoding technologies, generally based on prediction, transformation, quantization, entropy coding, and loop filtering. Video encoding standards, such as the High Efficiency Video Coding Standard (HEVC / H.265), Versatile Video Coding Standard (VVC / H.266), and AVS standards, specify specific encoding formats. Petition 870250092909, dated 10 / 10 / 2025, page 8 / 100

Claims

1 / 5 CLAIMS 1. A method for decoding a bitstream associated with a video sequence, characterized in that it comprises: selecting, from a plurality of predefined sets of probability parameters, a first set of probability parameters to initiate one or more context models for a layer B; and performing layer B entropy decoding based on one or more context models and the first set of probability parameters, wherein the selection is based on a layer B encoding condition or a signal in the bitstream.

2. Method, according to claim 1, characterized in that the first set of probability parameters comprises at least one of: an initial probability for use with one or more context models, an adaptation rate of one or more context models, a plurality of weights that is associated, respectively, with a plurality of probabilities, or an adjusted probability for use with one or more context models.

3. A method according to claim 1, characterized in that the encoding condition comprises at least one of a quantization parameter (QP), a temporal identifier, a low delay condition, a non-low delay condition, or a rate cost associated with the first layer.

4. Method according to claim 1, characterized in that the signal comprises an indicator or a parameter in the bit stream. Petition 870250092909, dated 10 / 10 / 2025, page 65 / 100 2 / 5 5. A method according to claim 4, characterized in that the selection is based on an indicator signaled in a sequence parameter set (SPS), an image parameter set (PPS), an image header, or a layer header.

6. Method, according to claim 4, characterized in that the indicator or parameter has a value dependent on whether a low delay condition or a non-low delay condition is used to encode layer B.

7. Method, according to claim 1, characterized in that the plurality of predefined sets of probability parameters comprises two predefined sets of probability parameters for layer B.

8. A method for encoding a bitstream associated with a video sequence, characterized in that it comprises: selecting, from a plurality of predefined sets of probability parameters, a first set of probability parameters to initiate one or more context models for a layer B; and performing entropy encoding of layer B based on one or more context models and the first set of probability parameters, wherein the selection is based on a layer B encoding condition or a signal in the bitstream.

9. Method according to claim 8, characterized in that the first set of probability parameters comprises at least one of: an initial probability for use with one or more context models, an adaptation rate of one or more context models, Petition 870250092909, dated 10 / 10 / 2025, page 66 / 100 3 / 5 a plurality of weights that is associated, respectively, with a plurality of probabilities, or an adjusted probability for use with one or more context models.

10. A method according to claim 8, characterized in that the encoding condition comprises at least one of a quantization parameter (QP), a temporal identifier, a low delay condition, a non-low delay condition, or a rate cost associated with the first layer.

11. Method, according to claim 8, characterized in that it further comprises: encoding, in the bitstream, an indicator or parameter associated with layer B, the indicator or parameter indicating that the first set of probability parameters is selected.

12. Method, according to claim 11, characterized in that the indicator is signaled in a sequence parameter set (SPS), an image parameter set (PPS), an image header or a layer header.

13. Method, according to claim 11, characterized in that the indicator or parameter has a value dependent on whether a low delay condition or a non-low delay condition is used to encode the first layer.

14. Method according to claim 8, characterized in that the plurality of predefined sets of probability parameters comprises two predefined sets of probability parameters for layer B.

15. Non-transient computer-readable storage media characterized in that it stores a bitstream of a video for processing, according to: selecting, from a plurality of predefined sets of probability parameters, a first set of probability parameters to initiate one or more context models for a layer B; and performing layer B entropy encoding or decoding based on one or more context models and the first set of probability parameters, wherein the selection is based on a layer B encoding condition or a signal in the bitstream.

16. Non-transient computer-readable storage media according to claim 15, characterized in that the first set of probability parameters comprises at least one of: an initial probability for use with one or more context models, an adaptation rate of one or more context models, a plurality of weights that are associated, respectively, with a plurality of probabilities, or an adjusted probability for use with one or more context models.

17. Non-transient computer-readable storage media according to claim 15, characterized in that the encoding condition comprises at least one of a quantization parameter (QP), a temporal identifier, a low delay condition, a non-low delay condition, or a rate cost associated with the first layer.

18. Non-transient computer-readable storage medium according to claim 15, characterized in that the signal comprises an indicator or a parameter in the bitstream.

19. Non-transient computer-readable storage media according to claim 18, characterized in that the selection is based on an indicator signaled in a sequence parameter set (SPS), an image parameter set (PPS), an image header, or a layer header.

20. Non-transient computer-readable storage media according to claim 15, characterized in that the indicator or parameter has a value dependent on whether a low-delay condition or a non-low-delay condition is used to encode layer B. Petition 870250092909, dated 10 / 10 / 2025, pp. 69 / 100