Method and apparatus for CTU-level inheritance of CABAC context initialization

By implementing CABAC state inheritance at the CTU level for video coding, the solution addresses the inefficiencies in initializing CABAC context models, improving coding efficiency and preserving frame-level parallelism.

JP2025515239APending Publication Date: 2025-05-14TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024518166
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-04-18
Filing Date
2023-04-21
Publication Date
2025-05-14

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently initializing context-adaptive binary arithmetic coding (CABAC) context models, particularly in maintaining coding efficiency while preserving frame-level parallelism.

Method used

The proposed solution involves CABAC state inheritance at the CTU level, where the CABAC states of encoded CTUs are stored and used to initialize context models for subsequent CTUs, rather than relying on picture or slice-level inheritance.

Benefits of technology

This approach enhances coding efficiency by allowing context model initialization based on specific CTU states, while also maintaining frame-level parallelism by avoiding dependencies introduced by picture or slice-level inheritance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025515239000001_ABST
    Figure 2025515239000001_ABST
Patent Text Reader

Abstract

In the method, a video bitstream is received that includes a current coding tree unit (CTU) in a current picture and a source CTU in a source picture. A value of a syntax element is determined. The syntax element indicates whether the current CTU is coded using context-adaptive binary arithmetic coding (CABAC). Context parameters of a context model of the current CTU are derived based on (i) predefined context model initialization information and (ii) CABAC state information of a source CTU corresponding to the current CTU. The CABAC state information is stored at a CTU level, not at a picture level or slice level. A context model of the current CTU is determined based on the derived context parameters. The current CTU is reconstructed based on the determined context model.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] Incorporation by Reference This application claims the benefit of priority to U.S. Patent Application No. 18 / 136,134, entitled "CTU-Level Inheritance for CABAC Context Initialization," filed April 18, 2023, which claims the benefit of priority to U.S. Provisional Application No. 63 / 334,592, entitled "CTU-Level Inheritance for CABAC Context Initialization," filed April 25, 2022. The disclosures of the prior applications are incorporated herein by reference in their entireties.

[0002] Technical Field This disclosure describes embodiments generally related to video coding. [Background technology]

[0003] The background discussion provided herein is intended to generally indicate the context of the present disclosure. The work of the inventors identified in this application, to the extent that their work is described in this background section, as well as aspects of this specification that may not qualify as prior art as of the filing date, are not admitted expressly or impliedly as prior art to the present disclosure.

[0004] Image / video compression can help transmit image / video files across various devices, storage, and networks with minimal quality degradation. In some examples, video codec technology can compress video based on spatial and temporal redundancy. In one example, video codecs can use a technique called intra prediction, which can compress images based on spatial redundancy. For example, intra prediction can use reference data from the current picture being reconstructed for sample prediction. In another example, video codecs can use a technique called inter prediction, which can compress images based on temporal redundancy. For example, inter prediction can predict samples of a current picture from a previously reconstructed picture using motion compensation. Motion compensation is commonly denoted by a motion vector (MV). Summary of the Invention [Problem to be solved by the invention]

[0005] Aspects of the present disclosure provide a method and apparatus for video encoding / decoding. In some examples, an apparatus for video decoding / encoding includes a receiving circuit and a processing circuit. [Means for solving the problem]

[0006] According to an aspect of the present disclosure, a method of video decoding performed in a video decoder is provided. In the method, a video bitstream including a current coding tree unit (CTU) in a current picture and a source CTU in a source picture is received. A value of a syntax element is determined. The syntax element indicates whether the current CTU is encoded using context-adaptive binary arithmetic coding (CABAC). Context parameters of a context model of the current CTU are derived based on (i) predefined context model initialization information and (ii) CABAC state information of a source CTU corresponding to the current CTU. The CABAC state information is stored at a CTU level, not at a picture level or slice level. A context model of the current CTU is determined based on the derived context parameters. The current CTU is reconstructed based on the determined context model. In some embodiments, the source CTU is one of (i) a first CTU of the source picture and (ii) one or more adjacent neighbor CTUs of the first CTU of the source picture.

[0007] In one example, a source CTU that corresponds to a current CTU is a co-located CTU of the current CTU in the source picture, that is located at the same relative position in the source picture as the current CTU in the current picture.

[0008] In one example, the source CTU that corresponds to the current CTU is one of the co-located CTU of the current CTU in the source picture, the left neighbor CTU of the co-located CTU, and the top neighbor CTU of the co-located CTU. The co-located CTU is located in the same relative position in the source picture as the current CTU in the current picture.

[0009] In some embodiments, the position of the source CTU that corresponds to the current CTU is indicated by position information included in the video bitstream, the position information being included at one of the sequence level or picture level.

[0010] In some embodiments, the source CTU position indicates the start position of the Nth CTU row in the source picture.

[0011] In some embodiments, context parameters of the context model of the current CTU are derived based on predefined context model initialization information and the location of the source CTU outside the source picture.

[0012] In response to the current CTU being located in the nth partition of the current picture, the source CTU is located at a predefined location within the same nth partition of the source picture.

[0013] In one example, the source CTU is determined as (i) the slice containing the co-located CTU in the source picture that corresponds to the first CTU in the current picture, and (ii) the CTU located in one of the first slices of the source picture.

[0014] In one example, the source picture is determined as one of the following: the closest previous picture with the same temporal identification (ID) as the current picture, the closest previous picture with the same temporal ID and the same picture quantization parameter (QP) as the current picture, the reference picture with the smallest reference index in the reference list of the current picture, and the reference picture in the reference list of the current picture with the smallest temporal distance from the current picture.

[0015] In one example, in response to the source picture and a co-located CTU corresponding to the current CTU in the source picture being available, context parameters of a context model for the current CTU are derived based on CABAC state information of a source CTU corresponding to the current CTU. In one example, in response to the source picture and a co-located CTU corresponding to the current CTU in the source picture being unavailable, context parameters of a context model for the current CTU are derived based on predefined context model initialization information.

[0016] In some embodiments, whether the context parameters of the context model of the current CTU are based on the CABAC state information of the source CTU corresponding to the current CTU is determined according to CABAC inheritance information, which is included in the encoding information.

[0017] According to another aspect of the present disclosure, there is provided an apparatus including a processing circuit, the processing circuit being configured to perform any of the described methods for video decoding / encoding.

[0018] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding, cause the computer to perform any of the described methods for video decoding / encoding. [Brief description of the drawings]

[0019] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.

[0020] [Figure 1] FIG. 1 is a schematic diagram of an example block diagram of a communication system (100).

[0021] [Diagram 2] FIG. 2 is a schematic diagram of an example block diagram of a decoder.

[0022] [Diagram 3] FIG. 2 is a schematic diagram of an example block diagram of an encoder;

[0023] [Figure 4] 1 is an example flowchart for decoding bins based on context-adaptive binary arithmetic coding (CABAC), in accordance with some embodiments of the present disclosure.

[0024] [Diagram 5] 1 shows a flowchart outlining a decoding process according to some embodiments of the present disclosure.

[0025] [Figure 6] 1 shows a flowchart outlining an encoding process according to some embodiments of the present disclosure.

[0026] [Figure 7] 1 is a schematic diagram of an exemplary computer system according to an embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0027] 1 illustrates a block diagram of a video processing system (100) in some examples. The video processing system (100) is an example of an application of the disclosed subject matter, which is a video encoder and video decoder in a streaming environment. The disclosed subject matter is equally applicable to other video-enabled applications including, for example, video conferencing, digital TV, streaming services, and storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0028] The video processing system (100) includes a capture subsystem (113) that may include a video source (101), such as a digital camera, and generates a stream of uncompressed video pictures (102). In one example, the stream of video pictures (102) includes samples taken by the digital camera. The stream of video pictures (102), shown as thick lines to emphasize the large amount of data compared to the encoded video data (104) (or encoded video bitstream), may be processed by an electronic device (120) that includes a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (104) (or encoded video bitstream), shown as thin lines to emphasize the small amount of data compared to the stream of video pictures (102), may be stored in a streaming server (105) for future use. One or more streaming client subsystems, such as the client subsystems (106) and (108) of FIG. 1, can access the streaming server (105) to retrieve copies (107) and (109) of the encoded video data (104). The client subsystem (106) can include a video decoder (110), for example within an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and generates an outgoing stream of video pictures (111) that can be rendered on a display (112) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (104), (107), and (109) (e.g., a video bitstream) can be encoded according to some video encoding / compression standard.Examples of these standards include ITU-T Recommendation H.265. In one example, the developing video coding standard is informally known as Universal Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.

[0029] It should be noted that the electronic devices (120) and (130) may include other components (not shown). For example, the electronic device (120) may include a video decoder (not shown), and the electronic device (130) may include a video encoder (not shown).

[0030] 2 shows an example block diagram of a video decoder (210). The video decoder (210) can be included in an electronic device (230). The electronic device (230) can include a receiver (231) (e.g., a receiving circuit). The video decoder (210) can be used in place of the video decoder (110) in the example of FIG. 1.

[0031] The receiver (231) may receive one or more coded video sequences to be decoded by the video decoder (210). In some embodiments, one coded video sequence is received at a time, and the decoding of each coded video sequence is independent of the decoding of the other coded video sequences. The coded video sequences may be received from a channel (201), which may be a hardware / software link to a storage device that stores the encoded video data. The receiver (231) may receive the encoded video data along with other data, e.g., coded audio data and / or auxiliary data streams, which may be forwarded to a respective usage entity (not shown). The receiver (231) may separate the coded video sequences from the other data. To address network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other applications, the buffer memory may be external to the video decoder (210) (not shown). In still other applications, there may be a buffer memory (not shown) external to the video decoder (210), for example to handle network jitter, and yet another buffer memory (215) internal to the video decoder (210), for example to handle playback timing. If the receiver (231) receives data from a storage / forwarding device with sufficient bandwidth and controllability, or from an isosynchronous network, the buffer memory (215) may be unnecessary or small. For use with best-effort packet networks such as the Internet, the buffer memory (215) may be necessary, may be relatively large, may be advantageously adaptively sized, and may be implemented, at least in part, in an operating system or similar element (not shown) external to the video decoder (210).

[0032] The video decoder (210) may include a parser (220) that reconstructs symbols (221) from the coded video sequence. The categories of symbols include information used to manage the operation of the video decoder (210) and potentially information for controlling a rendering device such as a rendering device (212) (e.g., a display screen) that is not an integral part of the electronic device (230) but may be coupled to the electronic device (230), as shown in FIG. 2. The control information for the rendering device may be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (220) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow various principles including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) can extract from the coded video sequence a set of subgroup parameters for at least one subgroup of pixels in the video decoder based on the at least one parameter corresponding to the group. The subgroup can include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The parser (220) can also extract information from the coded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0033] The parser (220) can perform an entropy decoding / parsing operation on the video sequence received from the buffer memory (215) to generate symbols (221).

[0034] The reconstruction of the symbols (221) can involve several different units, depending on the type of coded video picture or part thereof (inter / intra picture, inter / intra block, etc.) and other factors. Which units are involved and how can be controlled by subgroup control information parsed from the coded video sequence by the parser (220). The flow of such subgroup control information between the parser (220) and the following units is not shown for clarity.

[0035] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into several functional units, as described below. In a practical implementation operating under commercial constraints, many of these units will interact closely with each other and may be, at least in part, integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.

[0036] The first unit is a scaler / inverse transform unit (251), which receives quantized transform coefficients and control information from the parser (220) as symbols (221), including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (251) can output a block containing sample values ​​that can be input to an aggregator (255).

[0037] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to intra-coded blocks. Intra-coded blocks are blocks that do not use prediction information from a previously reconstructed picture, but may use prediction information from a previously reconstructed part of the current picture. Such prediction information may be provided by an intra picture prediction unit (252). In some cases, the intra picture prediction unit (252) generates a block of the same size and shape as the block being reconstructed using surrounding already reconstructed information fetched from a current picture buffer (258). The current picture buffer (258) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. The aggregator (255) adds, possibly on a sample-by-sample basis, the prediction information generated by the intra prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251).

[0038] In other cases, the output samples of the scaler / inverse transform unit (251) may relate to an inter-coded, potentially motion-compensated block. In such cases, the motion compensation prediction unit (253) may access a reference picture memory (257) to fetch samples used for prediction. After motion compensating the fetched samples according to the symbols (221) related to the block, these samples may be added by an aggregator (255) to the output of the scaler / inverse transform unit (251) (in this case referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (257) from which the motion compensation prediction unit (253) fetches the prediction samples may be controlled by motion vectors, which are available to the motion compensation prediction unit (253) in the form of symbols (221), which may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​fetched from the reference picture memory (257) when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, and the like.

[0039] The output samples of the aggregator (255) can be subjected to various loop filtering techniques in a loop filtering unit (256). Video compression techniques can include in-loop filter techniques controlled by parameters contained in the coded video sequence (also called coded video bitstream) and made available to the loop filter unit (256) as symbols (221) from the parser (220). Video compression can also be responsive to meta-information obtained during decoding of a previous portion (in decoding order) of the coded picture or coded video sequence, and can also be responsive to previously reconstructed and loop filtered sample values.

[0040] The output of the loop filter unit (256) can be a sample stream that can be output to a rendering device (212) and can also be stored in a reference picture memory (257) for use in future inter-picture prediction.

[0041] Certain coded pictures, once fully reconstructed, can be used as reference pictures for future predictions. For example, once the coded picture corresponding to the current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and a new current picture buffer can be reallocated before starting reconstruction of the next coded picture.

[0042] The video decoder (210) may perform decoding operations according to a given video compression technique or standard, for example, ITU-T Recommendation H.265. The coded video sequence may conform to a syntax specified by the video compression technique or standard being used. This means that the coded video sequence conforms to both the syntax of the video compression technique or standard and the profile described in the video compression technique or standard. Specifically, the profile may select certain tools from all tools available in the video compression technique or standard as the only tools available under the profile. Also, as a requirement for compliance, the complexity of the coded video sequence may be within a range defined by the level of the video compression technique or standard. In some cases, the level constrains the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limitations set by the level may be further constrained in some cases through a Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.

[0043] In some embodiments, the receiver (231) can receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence(s). The additional data may be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) improvement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0044] 3 shows an example block diagram of a video encoder (303). The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) can be used in place of the video encoder (103) in the example of FIG. 1.

[0045] The video encoder (303) can receive video samples from a video source (301) (which is not part of the electronic device (320) in the example of FIG. 3) that can capture video images to be encoded by the video encoder (303). In another example, the video source (301) is part of the electronic device (320).

[0046] The video source (301) may provide a source video sequence to be encoded by the video encoder (303) in the form of a digital video sample stream that may be of any suitable bit depth (e.g. 8-bit, 10-bit, 12-bit, ...), any color space (e.g. BT.601 YCrCB, RGB, ...), and any suitable sampling structure (e.g. YCrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (301) may be a storage device that stores pre-prepared videos. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as a number of individual pictures that give motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, each pixel may contain one or more samples depending on the sampling structure, color space, etc. in use. Those skilled in the art will readily appreciate the relationship between pixels and samples. The following description focuses on samples.

[0047] According to an embodiment, the video encoder (303) may encode and compress pictures of a source video sequence into an encoded video sequence (343) in real-time or under any other time constraint as required. Enforcing an appropriate encoding rate is one function of the controller (350). In some embodiments, the controller (350) controls and is operatively coupled to other functional units as described below. Coupling is not shown for clarity. Parameters set by the controller (350) may include rate control related parameters (picture skip, quantizer, lambda value for rate-distortion optimization techniques, ...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (350) may be configured to have other suitable functions for the video encoder (303) optimized for a certain system design.

[0048] In some embodiments, the video encoder (303) is configured to operate in an encoding loop. As a very simplified description, in one example, the encoding loop can include a source coder (330) (e.g., responsible for generating symbols, such as a symbol stream, based on an input picture to be encoded and a reference picture(s)) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to generate sample data in a manner similar to that which a (remote) decoder would generate. The reconstructed sample stream (sample data) is input to a reference picture memory (334). Since the decoding of the symbol stream results in bit-accurate results that are independent of the decoder location (local or remote), the contents in the reference picture memory (334) are also bit-accurate between the local and remote encoders. In other words, the predictive part of the encoder "sees" exactly the same sample values ​​as the decoder "sees" when using prediction during decoding. This basic principle of reference picture synchrony (and the resulting drift if synchrony cannot be maintained, eg, due to channel errors) is also used in several related techniques.

[0049] The operation of the "local" decoder (333) may be the same as a "remote" decoder, such as the video decoder (210) already described in detail in connection with Figure 2. However, with brief reference also to Figure 2, because symbols are available and the encoding / decoding of symbols into an encoded video sequence by the entropy coder (345) and parser (220) may be lossless, the entropy decoding portion of the video decoder (210), including the buffer memory (215) and parser (220), may not be fully implemented in the local decoder (333).

[0050] In some embodiments, decoder techniques, except for parsing / entropy decoding, present in a decoder are present in the same or substantially the same functional form in a corresponding encoder. Thus, the disclosed subject matter focuses on the operation of the decoder. A description of the encoder techniques can be omitted since they are the inverse of the decoder techniques described generically. In certain areas, more detailed descriptions are provided below.

[0051] In operation, in some examples, the source coder (330) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from a video sequence designated as “reference pictures.” In this manner, the coding engine (332) codes differences between pixel blocks of the input picture and pixel blocks of reference pictures that may be selected as predictive references for the input picture.

[0052] The local video decoder (333) can decode the coded video data of the pictures that may be designated as reference pictures based on the symbols generated by the source coder (330). The operation of the coding engine (332) can advantageously be a lossy process. If the coded video data can be decoded by a video decoder (not shown in FIG. 3), the reconstructed video sequence can be a replica of the source video sequence, typically with some errors. The local video decoder (333) can replicate the decoding process that may be performed on the reference pictures by the video decoder and store the reconstructed reference pictures in the reference picture memory (334). In this way, the video encoder (303) can locally store copies of reconstructed reference pictures that have a common content with the reconstructed reference pictures that would be obtained by the far-end video decoder (in the absence of transmission errors).

[0053] The predictor (335) can perform a prediction search for the coding engine (332). That is, for a new picture to be coded, the predictor (335) can search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or types of metadata, such as reference picture motion vectors, block shapes, etc., that may serve as suitable prediction references for the new picture. The predictor (335) can operate on a sample block, pixel block basis to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (335), the input picture can have prediction references drawn from multiple reference pictures stored in the reference picture memory (334).

[0054] The controller (350) can manage the encoding operations of the source coder (330), including, for example, setting the parameters and subgroup parameters used to encode the video data.

[0055] The output of all the functional units mentioned above may undergo entropy coding in an entropy coder (345), which converts the symbols produced by the various functional units into a coded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, or arithmetic coding.

[0056] The transmitter (340) can buffer the coded video sequence produced by the entropy coder (345) and prepare it for transmission over a communication channel (360), which can be a hardware / software link to a storage device that stores the encoded video data. The transmitter (340) can merge the coded video data from the video encoder (303) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).

[0057] The controller (350) can manage the operation of the video encoder (303). During encoding, the controller (350) can assign a certain encoding picture type to each encoded picture, which can affect the encoding technique that can be applied to each picture. For example, pictures may often be assigned as one of the following picture types:

[0058] An intra picture (I picture) may be a picture that can be coded and decoded without using other pictures in a sequence as a source of prediction. Some video encoders allow various types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of these variations of I pictures and their respective uses and characteristics.

[0059] A predictive picture (P picture) may be a picture that can be encoded and decoded using intra- or inter-prediction, using at most one motion vector and reference index to predict sample values ​​for each block.

[0060] A bidirectionally predictive picture (B picture) may be a picture that can be coded and decoded using intra- or inter-prediction, using at most two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, a multi-predictive picture can use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0061] A source picture may generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks as determined by the coding assignment applied to the respective picture of the block. For example, blocks of an I picture may be non-predictively coded or predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture may be predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Blocks of a B picture may be predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.

[0062] The video encoder (303) may perform encoding operations according to a given video encoding technique or standard, such as ITU-T Recommendation H.265. In its operations, the video encoder (303) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancies in the input video sequence. Thus, the encoded video data may conform to a syntax specified by the video encoding technique or standard being used.

[0063] In some embodiments, the transmitter (340) can transmit additional data along with the encoded video. The source coder (330) can include such data as part of the coded video sequence. The additional data may include other forms of redundant data, such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0064] A video may be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation in a given picture, while inter-picture prediction exploits correlation (temporal or other) between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is divided into blocks. If a block in the current picture is similar to a reference block in a previously coded and still buffered reference picture in the video, the block in the current picture can be coded by a vector called a motion vector. The motion vector points to a reference block in the reference picture, and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0065] In some embodiments, bi-prediction techniques can be used in inter-picture prediction. According to the bi-prediction techniques, two reference pictures, such as a first reference picture and a second reference picture, both of which precede a current picture in a video in decoding order (but may be past and future in display order, respectively), are used. A block in the current picture can be coded by a first motion vector that points to a first reference block in the first reference picture and a second motion vector that points to a second reference block in the second reference picture. A block can be predicted by a combination of the first reference block and the second reference block.

[0066] Furthermore, merge mode techniques can be used in inter-picture prediction to improve coding efficiency.

[0067] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed on a block-by-block basis. For example, according to the HEVC standard, pictures in a sequence of video pictures are divided into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. In general, a CTU includes three coding tree blocks (CTBs), which are one luma CTB and two chroma CTBs. Each CTU can be recursively quad-tree partitioned into one or more coding units (CUs). For example, a CTU of 64×64 pixels can be partitioned into one CU of 64×64 pixels, or four CUs of 32×32 pixels, or sixteen CUs of 16×16 pixels. In one example, each CU is analyzed to determine a prediction type (such as an inter prediction type or an intra prediction type) for the CU. The CU is partitioned into one or more prediction units (PUs) depending on the temporal and / or spatial predictability. In general, each PU includes a luma prediction block (PB) and two chroma PBs. In an embodiment, the prediction operation in coding (encoding / decoding) is performed in units of prediction blocks. Taking a luma prediction block as an example of a prediction block, the prediction block includes a matrix of values ​​(e.g., luma values) for pixels of 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0068] It should be noted that the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using any suitable technology. In one embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using one or more processors executing software instructions.

[0069] This disclosure includes embodiments related to context-adaptive binary arithmetic coding (CABAC) initialization. For example, CABAC initialization can be based on information of CTUs in previously coded pictures.

[0070] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC1 / SC29 / WG11) published the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (version 1), 2014 (version 2), 2015 (version 3) and 2016 (version 4). In 2015, these two standardization bodies jointly formed the Joint Video Exploration Team (JVET) to explore the possibility of developing the next video coding standard beyond HEVC. In October 2017, these two standardization bodies announced a Joint Call for Proposals on Video Compression with Capability beyond HEVC (CfP). By February 15, 2018, a total of 22 CfP responses for Standard Dynamic Range (SDR), 12 CfP responses for High Dynamic Range (HDR), and 12 CfP responses for 360 video category were submitted. In April 2018, all CfP responses were evaluated at the 122 MPEG / 10th JVET meeting. As a result of the meeting, JVET formally launched the standardization process for next-generation video coding beyond HEVC, the new standard was named Versatile Video Coding (VVC), and JVET was renamed Joint Video Experts Team. In 2020, ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC1 / SC29 / WG11) published the VVC video coding standard (version 1).

[0071] The HEVC CABAC engine uses a table-based probability transition process between 64 different representative probability states. In HEVC, the current interval range (e.g., ivlCurrRange), which represents the state of the coding engine, can be quantized to a set of four values ​​before calculating the new interval range. HEVC state transitions can be implemented using a table that contains all 64 × 4 8-bit pre-computed values ​​to approximate the value of ivlCurrRange * pLPS(pStateIdx), where pLPS is the probability of the least probable symbol (LPS) and pStateIdx is the index of the current state. Also, the decoding decision can be implemented using a pre-computed look-up table (LUT). For example, the LPS range (e.g., ivlLpsRange) can be determined based on the LUT in equation (1). The LPS range, ivlLpsRange, can further be used to update the current interval range (eg, ivlCurrRange) and to calculate the output bin value (eg, binVal). ivlLpsRange=rangeTabLps[pStateIdx][qRangeIdx] Formula (1)

[0072] For example, the probability in VVC can be linearly represented by the probability index pStateIdx. Therefore, all calculations can be done in formulas without LUT operations. To improve the accuracy of the probability estimation, a multi-hypothesis probability update model can be applied. The pStateIdx used in the interval subdivision of the binary arithmetic coder can be a combination of two probabilities pStateIdx0 and pStateIdx1. The two probabilities can be associated with each context model of CABAC and can be updated independently with different adaptation rates. The adaptation rates of pStateIdx0 and pStateIdx1 for each context model can be pre-trained based on the statistics of the associated bin. The probability estimate of pStateIdx can be the average of the estimates from the two hypotheses (e.g. pStateIdx0 and pStateIdx1).

[0073] FIG. 4 illustrates an example flowchart (400) for decoding a single binary decision (e.g., decoding a bin) in VVC. As shown in the flowchart (400), at (S410), input variables of a context table (e.g., ctxTable) and a context index (e.g., ctxIdx) may be received. In some embodiments, the input variables may also include a current interval range (e.g., ivlCurrRange) and an interval offset (e.g., ivlOffset). At (S420), variables of a quantization range index (e.g., qRangeIdx), a probability state (e.g., pState), a most probable symbol value (e.g., valMps), and a minimum probable symbol interval range (e.g., ivlLpsRange) may be determined, and ivlCurrRange may be updated as ivlCurrRange-ivlLpsRange. In (S430), if ivlOffset is less than ivlCurrRange, the bin value (e.g., binVal) may be determined as valMps, as shown in (S450). If ivlOffset is greater than or equal to ivlCurrRange, the bin value (e.g., binVal) may be determined as 1-valMps, ivlOffset may be decremented by ivlCurrRange, and ivlCurrRange may be set equal to ivlLpsRange, as shown in (S440). In (S460), pStateIdx0 and pStateIdx0 may be updated based on the parameters of shift0, shift1, and the bin value. In (S370), a renormalization process (e.g., RenormD) in the decoder may be performed based on the determined pStateIdx0 and pStateIdx0.

[0074] Similar to HEVC, VVC CABAC can also have a quantization parameter (QP)-dependent initialization process that is invoked at the beginning of each slice. Given an initial value of luma QP for a slice, the initial probability state of the context model, denoted as preCtxState, can be derived in equations (2)-(4) as follows: m=slopeIdx×5-45 Formula (2) n=(offsetIdx<<3)+7 Equation (3) preCtxState=Clip3(1,127,((m×(QP-32))>>4)+n) Equation (4) Here, slopeIdx and offsetIdx can be limited to 3 bits, and the total initialization value is represented with 6 bits of precision. The initial probability state of the context model, preCtxState, directly represents the probability in the linear domain. Hence, preCtxState may only require an appropriate shift operation before being input to the arithmetic coding engine (e.g., CABAC). Thus, the logarithmic to linear domain mapping and the 256-byte table can be saved (or skipped). pStateIdx0 and pStateIdx1 can further be determined based on preCtxState as per the following equations (5) and (6): pStateIdx0=preCtxState<<3 Equation (5) pStateIdx1=preCtxState<<7 Equation (6)

[0075] To further improve CABAC efficiency, an inherited context initialization method like JVET-Z0134 is proposed. For example, the context states of all context models in an inter slice (e.g., B-type slice or P-type slice) including the last CTU in the corresponding picture are stored. The stored states can be used to initialize the context models in the next inter slice with the same slice type, QP value, and temporal layer identification (ID). Furthermore, for each inter slice, a control flag can be signaled to select the context initialization scheme to be used for that slice. If the control flag is equal to 0, it indicates that the context model of the slice is initialized using one of the existing context initialization tables (indicated by the context initiation flag (e.g., sh_cabac_init_flag)). Otherwise, the context model of the slice is initialized by inheriting the context model from the previously coded picture.

[0076] Although the inherited context initialization method improves coding efficiency, the inherited context initialization method may break frame-level parallelism, such as frame parallel processing (FPP), in both the encoder and the decoder, because the inherited context initialization method introduces dependencies between frames at the same temporal level.

[0077] This disclosure may provide CABAC state inheritance at the coding tree unit (CTU) level instead of at the picture or slice level. The CABAC state may include, but is not limited to, a context and a window size. In one example, the context may include CABAC context variables indexed by ctxTable and ctxIdx.

[0078] In an embodiment, after a CTU is encoded, the CABAC states of all context models of that CTU may be stored. The stored CABAC states may further be used as a source of CABAC state inheritance for different CTUs. For example, before a current CTU of a current picture is encoded, it may be determined whether CABAC context initialization is required for the current CTU. For example, if the current CTU is the first CTU in the current picture or the first CTU in the current slice, CABAC context initialization may be required. To perform CABAC context initialization for the current CTU, (i) a predefined context model initialization lookup table or (ii) a stored CABAC state of the source CTU from the source picture may be applied.

[0079] In one example, the CABAC state of a CTU located at a particular position of a picture may be stored. In one example, only the CABAC state of the first CTU in a picture (or slice / tile / tile group) may be stored. In another example, the CABAC state for the first CTU in a picture (or slice / tile / tile group) and each of the CTUs in the first CTU's immediate spatial neighborhood may be stored.

[0080] In one embodiment, a source CTU corresponding to a current CTU may be selected based on candidate CTUs located at a certain position.

[0081] In one example, a co-located CTU of the current CTU in the source picture can be used as the source CTU that corresponds to the current CTU. The co-located CTU can be located in the source picture at the same relative position (e.g., with respect to the horizontal and vertical coordinates of the luma samples) as the current CTU in the current picture.

[0082] In one example, a source CTU corresponding to a current CTU may be selected from multiple locations including a co-located CTU of the current CTU and spatial neighbors of the co-located CTU. In one example, based on a predefined order, a first available CTU among the co-located CTU and spatial neighbors of the co-located CTU may be determined as the source CTU. In one example, a check order may be applied as follows: the co-located CTU, the left neighbor of the co-located CTU, and the top neighbor of the co-located CTU. In one example, the location of the source CTU may be signaled in the bitstream for the current CTU. In one example, the location of the source CTU may be signaled only for the current CTU that requires CABAC initialization. For example, if the current CTU is at the beginning of a picture or slice, CABAC initialization is required for the current CTU.

[0083] In one embodiment, the source of CABAC state inheritance can be determined from a CTU (or source CTU) at a predefined position in the source picture. The predefined position of the CTU in the source picture can be signaled by a higher level syntax such as at the sequence level (e.g., sequence parameter set (SPS)) or at the picture level (e.g., picture header or picture parameter set (PPS)). The CABAC context state can be stored after the CTU is encoded / decoded at the predefined position.

[0084] In one example, the predefined CTU location for applying the CABAC model inheritance (e.g., storing the CABAC state) may be at the beginning of the Nth CTU row (or the Nth row of CTUs) in the source picture. For example, N may be a positive integer, e.g., 3. For a source picture, the CABAC state may be stored only after the last CTU of the (N-1)th CTU row is finished (or decoded / encoded).

[0085] In one example, if the signaled position is outside the picture (or source picture), the CABAC model inheritance tool can be presumed to be invalid and a default initial CABAC model can be used for initialization. For example, the CABAC context initialization can be determined based on a predefined context model initialization lookup table for the current CTU.

[0086] In an embodiment, a picture is divided into multiple partitions, where the partitions include, but are not limited to, slices, tiles, sub-pictures, etc. Thus, a CTU in a picture to which the CABAC model inheritance is applied may be determined as a CTU at a predefined position in one of the partitions. For simplicity, the embodiments of the present disclosure may be provided based on slices.

[0087] In an embodiment, a source CTU at a predefined position from a corresponding slice in the source picture may be used. In one example, if the current picture is partitioned into N slices, where n represents a slice index, the nth slice in the current picture may inherit only the CABAC model state from the nth slice in the source picture that corresponds to the current picture. In one example, the CABAC model status may be inherited from a slice (or source slice) that covers the position (or co-located position) of the first CTU of the current slice in the source picture. In one example, the source slice from which the CABAC model status is inherited may be signaled. In one example, the CABAC model status may always be inherited from the first slice of the picture (or source picture). In one example, if the source picture and the current picture are partitioned into different partition sizes, CABAC inheritance may be disabled and default CABAC initialization models may be used for CABAC context initialization of the current CTU.

[0088] In one embodiment, the CABAC inheritance position within each slice (or tile / tile group) can be predefined.

[0089] In one embodiment, the CABAC inheritance position within each slice (or tile / tile group) can be signaled at a high level, such as at the sequence level in the SPS. In one example, the signaled position is available for all slices.

[0090] In an embodiment, the CABAC inheritance position within each slice can be signaled at the same level, such as in the slice header.

[0091] In an embodiment, if the predefined or signaled CTU position for CABAC inheritance is outside the range of the current slice, CABAC inheritance can be disabled, and thus only the default CABAC initialization model can be used.

[0092] In an embodiment, the selection of the source picture may be based on various parameters. The parameters may include a temporal level, a quantization parameter (QP), a reference index on a reference list, a temporal distance, etc. In one example, the source picture may be predefined according to the temporal level and / or the quantization parameter (QP) of the current picture. In one example, the closest previous picture in decoding order (or the closest previously decoded picture) having the same temporal ID as the current picture may be used as the source picture of the current picture. In one example, the closest previous picture in decoding order (or the closest previously decoded picture) having the same temporal ID and the same picture QP as the current picture may be used as the source picture. In one example, the reference picture with the smallest reference index on the reference list L0 may be used as the source picture. In one example, the reference picture on the reference list L0 with the smallest temporal distance to the current picture may be used as the source picture.

[0093] In an embodiment, if CABAC initialization is necessary (or required), it can be determined in various ways whether to use CABAC state inheritance from the source CTU. In one example, if a co-located CTU in the source picture corresponding to the current CTU in the source picture and the current picture is available, CABAC state inheritance can be used. Otherwise, a predefined CABAC initialization lookup table can be used to perform CABAC context initialization for the current CTU. In one example, whether to use CABAC state inheritance can be signaled in the bitstream at the CTU level. In an embodiment, a high level syntax can be signaled to indicate whether CTU-level CABAC state inheritance is applied. The high level syntax can be signaled at the sequence level (e.g., sequence parameter set), picture level (e.g., picture header or picture parameter set), slice level (e.g., slice header), tile level, tile group level, etc.

[0094] FIG. 5 shows a flow chart outlining a process (500) according to one embodiment of the present disclosure. The process (500) can be used in a video decoder. In various embodiments, the process (500) is performed by a processing circuit, such as a processing circuit performing the functions of the video decoder (110), a processing circuit performing the functions of the video decoder (210), etc. In some embodiments, the process (500) is implemented with software instructions, such that the processing circuit performs the process (500) when the processing circuit executes the software instructions. The process starts at (S501) and proceeds to (S510).

[0095] At (S510), a video bitstream including a current coding tree unit (CTU) in a current picture and a source CTU in a source picture is received.

[0096] At (S520), a value of a syntax element is determined. The syntax element indicates whether the current CTU is encoded using context-adaptive binary arithmetic coding (CABAC).

[0097] In (S530), context parameters of a context model of the current CTU are derived based on (i) predefined context model initialization information and (ii) CABAC state information of a source CTU corresponding to the current CTU. The CABAC state information is stored at the CTU level, not at the picture level or slice level.

[0098] In (S540), a context model for the current CTU is determined based on the derived context parameters.

[0099] In (S550), the current CTU is reconstructed based on the determined context model.

[0100] In one example, a source CTU that corresponds to a current CTU is a co-located CTU of the current CTU in the source picture. A co-located CTU is located at the same relative position in the source picture as the current CTU in the current picture.

[0101] In one example, the source CTU corresponding to the current CTU is one of the co-located CTU of the current CTU in the source picture, the left neighbor CTU of the co-located CTU, and the upper neighbor CTU of the co-located CTU. The co-located CTU is located at the same relative position in the source picture as the current CTU in the current picture.

[0102] In some embodiments, the location of the source CTU that corresponds to the current CTU is indicated by location information included in the encoding information, the location information being included at either the sequence level or the picture level.

[0103] In some embodiments, the source CTU position indicates the start position of the Nth CTU row in the source picture.

[0104] In some embodiments, context parameters of the context model of the current CTU are derived based on predefined context model initialization information and the location of the source CTU outside the source picture.

[0105] In response to the current CTU being located in an nth partition of the current picture, the source CTU is located at a predefined location within the same nth partition of the source picture.

[0106] In one example, the source CTU is determined as (i) the slice containing the co-located CTU in the source picture that corresponds to the first CTU in the current picture, and (ii) the CTU located in one of the first slices of the source picture.

[0107] In one example, the source picture is determined as one of the closest previous picture with the same temporal identification (ID) as the current picture, the closest previous picture with the same temporal ID and picture quantization parameter (QP) as the current picture, the reference picture with the smallest reference index in the reference list of the current picture, and the reference picture in the reference list of the current picture with the smallest temporal distance to the current picture.

[0108] In one example, in response to the source picture and a co-located CTU corresponding to the current CTU in the source picture being available, context parameters of a context model for the current CTU are derived based on CABAC state information of a source CTU corresponding to the current CTU. In one example, in response to the source picture and a co-located CTU corresponding to the current CTU in the source picture not being available, context parameters of a context model for the current CTU are derived based on predefined context model initialization information.

[0109] In some embodiments, whether the context parameters of the context model of the current CTU are based on the CABAC state information of the source CTU corresponding to the current CTU is determined according to CABAC inheritance information, which is included in the encoding information.

[0110] Then, the process proceeds to (S599) and ends.

[0111] The process (500) may be adapted as appropriate. Steps of the process (500) may be modified and / or omitted. Additional steps may be added. Any suitable order of implementation may be used.

[0112] 6 shows a flow chart outlining a process (600) according to one embodiment of the present disclosure. The process (600) can be used in a video encoder. In various embodiments, the process (600) is performed by a processing circuit, such as a processing circuit performing the functions of the video encoder (103), a processing circuit performing the functions of the video encoder (303), etc. In some embodiments, the process (600) is implemented with software instructions, such that the processing circuit performs the process (600) when the processing circuit executes the software instructions. The process begins at (S601) and proceeds to (S610).

[0113] In (S610), initial context parameters of a context model of a current CTU in a current picture are determined based on either (i) predefined context model initialization information and (ii) CABAC state information of a source CTU in a source picture corresponding to the current CTU.

[0114] In (S620), a context model for the current CTU is determined based on the determined initial context parameters.

[0115] In (S630), coding information of the current CTU is generated based on the determined context model, and the coding information indicates that the current CTU is CABAC coded.

[0116] Then, the process proceeds to (S699) and ends.

[0117] The process (600) may be adapted as appropriate. Steps of the process (600) may be modified and / or omitted. Additional steps may be added. Any suitable order of implementation may be used.

[0118] The techniques described above can be implemented as computer software using computer readable instructions and can be physically stored on one or more computer readable media. For example, Figure 7 illustrates a computer system (700) suitable for implementing certain embodiments of the disclosed subject matter.

[0119] Computer software may be coded using any suitable machine code or computer language and may apply assembly, compilation, linking, or similar mechanisms to create code containing instructions that can be executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc. directly, or through interpretation, microcode execution, etc.

[0120] The instructions may be executed on various types of computers or components thereof including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, and the like.

[0121] 7 for computer system (700) are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing the embodiments of the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement regarding any one or combination of components illustrated in the exemplary embodiment of computer system (700).

[0122] The computer system (700) may include certain human interface input devices that may be responsive to input by one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), and olfactory input (not shown). The human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from still image cameras), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0123] The input human interface devices may include one or more (only one of each is shown) of a keyboard (701), a mouse (702), a trackpad (703), a touch screen (710), a data glove (not shown), a joystick (705), a microphone (706), a scanner (707), and a camera (708).

[0124] The computer system (700) may also include some type of human interface output device. Such human interface output devices may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touch screen (710), data gloves (not shown), or joystick (705); although there may be haptic feedback devices that do not act as input devices), audio output devices (e.g., speakers (709), headphones (not shown)), visual output devices (e.g., screens (710) including CRT screens, LCD screens, plasma screens, OLED screens; each may or may not have touch screen input capability, each may or may not have haptic feedback capability, some of which may output two-dimensional visual output or higher than three-dimensional output through such means as stereoscopic output; virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0125] The computer system (700) may also include human accessible storage and associated media, such as optical media including CD / DVD ROM / RW (720) along with CD / DVD or similar media (721), thumb drives (722), removable hard drives or solid state drives (723), legacy magnetic media such as tapes and floppy disks (not shown), specialized ROM / ASIC / PLD based devices such as security dongles (not shown), etc.

[0126] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transitory signals.

[0127] The computer system (700) may also include an interface (754) to one or more communication networks (755). The networks may be, for example, wireless, wired, optical. The networks may further be local, wide area, metropolitan, in-vehicle and industrial, real-time, delay tolerant, etc. Examples of networks include Ethernet, WLAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable television, satellite television, terrestrial broadcast television, in-vehicle and industrial including CANBus, etc. Some networks typically require an external network interface adapter that is attached to some kind of general purpose data port or peripheral bus (749) (e.g., a USB port of the computer system (700)). Others are typically integrated into the core of the computer system (700) by attachment to a system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, computer system (700) can communicate with other entities. Such communications may be unidirectional, receive only (e.g., broadcast television), unidirectional transmit only (e.g., CANbus to certain CANbus devices), or bidirectional, for example, to other computer systems using local or wide area digital networks. With each of these networks and network interfaces as described above, certain protocols and protocol stacks may be used.

[0128] The aforementioned human interface devices, human accessible storage, and network interfaces may be attached to a core (740) of the computer system (700).

[0129] The core (740) may include one or more central processing units (CPUs) (741), graphics processing units (GPUs) (742), specialized programmable processing units in the form of field programmable gate arrays (FPGAs) (743), hardware accelerators for certain tasks (744), graphics adapters (750), etc. These devices may be connected through a system bus (748), along with read only memory (ROM) (745), random access memory (746), and internal mass storage devices (747), such as internal non-user accessible hard drives, solid state drives (SSDs), etc. In some computer systems, the system bus (748) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus (748) or through a peripheral bus (749). In one example, a screen (710) may be connected to the graphics adapter (750). Architectures for peripheral buses include PCI, USB, etc.

[0130] The CPU (741), GPU (742), FPGA (743), and accelerator (744) may execute certain instructions that may combine to constitute the above-mentioned computer code. The computer code may be stored in ROM (745) or RAM (746). Temporary data may also be stored in RAM (746), while persistent data may be stored, for example, in an internal mass storage device (747). Rapid storage and retrieval from any of the memory devices may be enabled through the use of cache memory, which may be closely associated with one or more of the CPU (741), GPU (742), mass storage device (747), ROM (745), RAM (746), etc.

[0131] The computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those having skill in the computer software arts.

[0132] By way of example and not limitation, a computer system having the architecture (700), and in particular the core (740), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer readable media. Such computer readable media can be user accessible mass storage as introduced above as well as media associated with some type of storage of the core (740) of a non-transitory nature, such as mass storage (747) internal to the core or ROM (745). Software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (740). The computer readable media can include one or more memory devices or chips, depending on the particular needs. The software can cause the core (740) and in particular the processors therein (including a CPU, GPU, FPGA, etc.) to execute certain processes or certain specific portions thereof described herein, including defining data structures stored in RAM (746) and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator (744)), which may operate in place of or together with software to perform particular processes or portions of particular processes described herein. Reference to software includes logic, and vice versa, as appropriate. Reference to a computer-readable medium may include circuitry (e.g., an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both, as appropriate. The present disclosure encompasses any suitable combination of hardware and software.

[0133] While this disclosure has described several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that fall within the scope of this disclosure. Thus, those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are thus within its spirit and scope.

Claims

1. 1. A method of video decoding performed by a video decoder, the method comprising: receiving a video bitstream including a current coding tree unit (CTU) in a current picture and a source CTU in a source picture; determining a value of a syntax element corresponding to a current CTU, the syntax element indicating whether the current CTU is encoded using context-adaptive binary arithmetic coding (CABAC); deriving context parameters of a context model of a current CTU based on (i) predefined context model initialization information and (ii) CABAC state information of a source CTU corresponding to the current CTU, where the CABAC state information is stored at a CTU level, not at a picture level or a slice level; determining a context model for the current CTU based on the derived context parameters; and reconstructing the current CTU based on the determined context model. method.

2. The method of claim 1 , wherein the source CTU is one of: (i) a first CTU of the source picture; and (ii) one or more adjacent neighboring CTUs of the first CTU of the source picture.

3. 2. The method of claim 1, wherein the source CTU corresponding to a current CTU is a co-located CTU of the current CTU in the source picture, the co-located CTU being positioned at the same relative position in the source picture as the current CTU in the current picture.

4. 3. The method of claim 2, wherein the source CTU corresponding to the current CTU is one of a co-located CTU of the current CTU in the source picture, a left neighboring CTU of the co-located CTU, and an upper neighboring CTU of the co-located CTU, and the co-located CTU is positioned at the same relative position in the source picture as the current CTU in the current picture.

5. The method of claim 1 , wherein the position of the source CTU corresponding to a current CTU is indicated by position information included in the video bitstream, the position information being included at one of a sequence level or a picture level.

6. The method of claim 5 , wherein the location of the source CTU indicates a starting location of an Nth CTU row of the source picture.

7. Deriving the context parameters includes: deriving the context parameters of the context model of the current CTU based on predefined context model initialization information and a location of the source CTU being outside the source picture; The method according to claim 5.

8. 2. The method of claim 1, wherein in response to a current CTU being located in an nth partition of a current picture, the source CTU is located at a predefined location within the same nth partition of the source picture.

9. 2. The method of claim 1, wherein the source CTU is determined as (i) a slice containing a co-located CTU in the source picture that corresponds to a first CTU of the current picture, and (ii) a CTU located in one of the first slices of the source picture.

10. The source picture is: The closest previous picture with the same temporal identity (ID) as the current picture, The closest previous picture with the same temporal ID and the same picture quantization parameter (QP) as the current picture, the reference picture with the smallest reference index in the reference list of the current picture, and The reference picture in the current picture's reference list that has the smallest temporal distance from the current picture The method of claim 1 , wherein the first and second vertices are determined as one of:

11. Deriving the context parameters includes: in response to the source picture and a co-located CTU in the source picture being available that corresponds to a current CTU, deriving the context parameters of the context model of the current CTU based on the CABAC state information of the source CTU that corresponds to the current CTU; deriving the context parameters of the context model of the current CTU based on the predefined context model initialization information in response to the source picture and the co-located CTU corresponding to the current CTU in the source picture being unavailable. The method of claim 1.

12. Deriving the context parameters includes: determining whether the context parameters of the context model of a current CTU are based on the CABAC state information of the source CTU corresponding to the current CTU according to CABAC inheritance information included in the encoding information; The method of claim 1.

13. Apparatus comprising processing circuitry configured to carry out a method according to any one of claims 1 to 12.