Temporal Initialization Points for Context-Based Arithmetic Coding

JP2025508699A5Pending Publication Date: 2026-02-10QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024547598
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-03-01
Filing Date
2023-03-02
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

In the prior art, in video encoding and decoding, it is difficult to effectively manage time initialization points, resulting in a decrease in storage limitations and encoding efficiency.

Method used

By using time identification values ​​and quantization parameter (QP) values, the time initialization points of the video encoder are determined, and a cache management strategy is adopted to determine which time initialization points should be retained or deleted based on the time ID and QP values ​​to ensure encoding efficiency and storage rationality.

Benefits of technology

Under the storage limitations, the coding efficiency is maintained to ensure the reliability of time decoding and video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The method includes determining one or more context values ​​of at least one context used to encode or decode a current slice or a current picture; determining that a buffer storing a set of temporal initialization points from two or more slices or two or more pictures for context-based arithmetic coding is full; determining a first set of temporal initialization points associated with a slice or picture from among the two or more slices or two or more pictures based on at least one of a slice type, a temporal identification value, or a quantization parameter (QP) value of the slice or picture; removing the first set of temporal initialization points associated with the slice or picture; and storing a second set of temporal initialization points associated with the current slice or the current picture, the second set of temporal initialization points being based on the determined one or more context values.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001]

[0001] This application claims priority to U.S. Application No. 18 / 176,863, filed March 1, 2023, and U.S. Provisional Application No. 63 / 268,844, filed March 3, 2022, and U.S. Provisional Application No. 63 / 362,118, filed March 29, 2022, the entire contents of which are incorporated herein by reference. U.S. Application No. 18 / 176,863, filed March 1, 2023, claims the benefit of U.S. Provisional Application No. 63 / 268,844, filed March 3, 2022, and U.S. Provisional Application No. 63 / 362,118, filed March 29, 2022.

[0002]

[0002] This disclosure relates to video encoding and decoding. [Background technology]

[0003]

[0003] Digital video capabilities may be incorporated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video gaming devices, video game consoles, cellular or satellite wireless telephones, so-called "smartphones," video teleconferencing devices, video streaming devices, and the like. Digital video devices implement video coding techniques such as those described in standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), ITU-T H.265 / High Efficiency Video Coding (HEVC), ITU-T H.266 / Versatile Video Coding (VVC), and extensions to such standards, as well as proprietary video codecs / formats such as AOMedia Video1 (AV1) developed by the Alliance for Open Media. Video devices may implement such video coding techniques to more efficiently transmit, receive, encode, decode, and / or store digital video information.

[0004]

[0004] Video coding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or remove redundancy inherent in video sequences. For block-based video coding, a video slice (e.g., a video picture or a portion of a video picture) may be partitioned into video blocks, which may also be referred to as coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are encoded using spatial prediction with respect to reference samples in neighboring blocks in the same picture. Video blocks in an inter-coded (P or B) slice of a picture may use spatial prediction with respect to reference samples in neighboring blocks in the same picture or temporal prediction with respect to reference samples in other reference pictures. A picture may be referred to as a frame, and a reference picture may be referred to as a reference frame. Summary of the Invention

[0005]

[0005] Generally, this disclosure describes techniques for determining an initialization point for one or more contexts used in context-based arithmetic coding, such as context-adaptive binary arithmetic coding (CABAC). The initialization point may be considered as a starting point for one or more contexts and may include one or more context states, window or rate adaptation parameters, and other parameters used for the arithmetic coding operation.

[0006] In some examples, an initialization point of one or more contexts of video data in a current slice or a current picture may be based on an initialization point of one or more contexts of video data in a previous picture. Such an initialization point is referred to as a temporal initialization point.

[0007]

[0007] This disclosure describes example techniques for determining a temporal initialization point for a current slice or a current picture. In some examples, a video coder (e.g., a video encoder or a video decoder) may utilize a temporal identification (ID) value and / or a quantization parameter (QP) value to determine the temporal initialization point. In some examples, the video coder may determine an initialization point for a current slice based on an initialization point of a previous slice located in a corresponding location.

[0008]

[0008] Due to memory size limitations, there may be limitations on how many temporal initialization points can be stored. When a new set of temporal initialization points is to be stored, a set of already stored temporal initialization points may be removed. This disclosure describes exemplary techniques of memory management for the insertion and removal of temporal initialization points based on temporal identification values ​​and / or QP values ​​that balance memory size limitations while ensuring that temporal initialization points that result in timely decoding remain in the buffer.

[0009]

[0009] In one example, the present disclosure provides a method for processing video data, comprising: determining one or more context values ​​of at least one context used to encode or decode a current slice or a current picture; determining that a buffer storing sets of temporal initialization points from two or more slices or two or more pictures for context-based arithmetic coding is full, where each set of temporal initialization points is associated with a slice or a picture among the two or more slices or two or more pictures, and includes one or more temporal initialization points; determining a first set of temporal initialization points associated with a slice or picture from among the two or more slices or two or more pictures based on at least one of a slice type, a temporal identification value, or a quantization parameter (QP) value of the slice or picture; removing the first set of temporal initialization points associated with the slice or picture from the buffer; and storing in the buffer a second set of temporal initialization points associated with the current slice or current picture, where the second set of temporal initialization points is based on the determined one or more context values.

[0010]

[0010] In one example, the disclosure describes a device for processing video data, the device including: a buffer configured to store a set of temporal initialization points from two or more slices or two or more pictures for context-based arithmetic coding; and a processing circuit coupled to the buffer, the processing circuit configured to determine one or more context values ​​of at least one context used to encode or decode a current slice or a current picture, and to store the set of temporal initialization points from two or more slices or two or more pictures for context-based arithmetic coding, each set of temporal initialization points associated with one slice or one picture of the two or more slices or two or more pictures. the first set of temporal initialization points associated with a slice or picture from among the two or more slices or two or more pictures based on at least one of a slice type, a temporal identification value, or a quantization parameter (QP) value of the slice or picture; removing the first set of temporal initialization points associated with the slice or picture from the buffer; and storing in the buffer a second set of temporal initialization points associated with a current slice or current picture, the second set of temporal initialization points being based on the determined one or more context values.

[0011]

[0011] In one example, the disclosure describes a computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to determine one or more context values ​​of at least one context used to encode or decode a current slice or current picture; determine that a buffer storing a set of temporal initialization points from two or more slices or two or more pictures for context-based arithmetic coding is full, where each set of temporal initialization points is associated with one slice or one picture among the two or more slices or two or more pictures, the set including one or more temporal initialization points; determine a first set of temporal initialization points associated with a slice or picture from among the two or more slices or two or more pictures based on at least one of a slice type, a temporal identification value, or a quantization parameter (QP) value of the slice or picture; remove the first set of temporal initialization points associated with the slice or picture from the buffer; and store in the buffer a second set of temporal initialization points associated with the current slice or current picture, the second set of temporal initialization points being based on the determined one or more context values.

[0012]

[0012] In one example, the disclosure describes a device for processing video data, the device including: means for determining one or more context values ​​of at least one context used to encode or decode a current slice or a current picture; means for determining that a buffer storing sets of temporal initialization points from two or more slices or two or more pictures for context-based arithmetic coding is full, where each set of temporal initialization points is associated with a slice or a picture among the two or more slices or two or more pictures, and includes one or more temporal initialization points; means for determining a first set of temporal initialization points associated with a slice or picture from among the two or more slices or two or more pictures based on at least one of a slice type, a temporal identification value, or a quantization parameter (QP) value of the slice or picture; means for removing the first set of temporal initialization points associated with the slice or picture from the buffer; and means for storing in the buffer a second set of temporal initialization points associated with the current slice or current picture, the second set of temporal initialization points being based on the determined one or more context values.

[0013]

[0013] The details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will become apparent from the description, drawings, and claims. [Brief description of the drawings]

[0014] [Figure 1]

[0014] FIG. 1 is a block diagram illustrating an example video encoding and decoding system in which techniques of this disclosure may be implemented. [Diagram 2]

[0015] 1 is a block diagram illustrating an example video encoder that may implement the techniques of this disclosure. [Diagram 3]

[0016] 1 is a block diagram illustrating an example video decoder that may implement the techniques of this disclosure. [Figure 4]

[0017] 1 is a flowchart illustrating an example method for encoding a current block, in accordance with techniques of this disclosure. [Diagram 5]

[0018] 5 is a flowchart illustrating an example method for decoding a current block, in accordance with techniques of this disclosure. [Figure 6]

[0019] 5 is a flowchart illustrating an example method for processing video data. [Figure 7]

[0020] 5 is a flowchart illustrating another exemplary method for processing video data. [Figure 8]

[0021] 5 is a flowchart illustrating another exemplary method for processing video data. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0015]

[0022] In video coding, such as context-based arithmetic coding, a video coder (e.g., a video encoder or a video decoder) utilizes an initialization point for initialization of a current picture or a current slice. For example, in the case of context-based arithmetic coding, including context-adaptive binary arithmetic coding (CABAC), an initialization point may be used for each context. The initialization point may include one or more context states, window or rate adaptation parameters, and other parameters required for the arithmetic coding operation.

[0016]

[0023] The initialization point may be predefined. However, in some examples, in addition to or instead of using a predefined initialization point, the video coder may utilize a temporal initialization point. The temporal initialization point may refer to an initialization point of a previous picture in the coding order that is utilized to determine (e.g., select) an initialization point for a current picture. For example, the temporal initialization point may be one or more context values ​​or may be derived based on one or more context values ​​of at least one context used to encode or decode the current picture, which values ​​are then used to initialize a context value of at least one context of a subsequent picture.

[0017]

[0024] There may be several problems with using temporal initialization points. For example, a buffer that stores temporal initialization points may have a limited size, and a video coder may remove a set of temporal initialization points (e.g., one or more temporal initialization points) to allow another set of temporal initialization points to be stored. This disclosure describes example techniques for determining which set of temporal initialization points to remove in a manner that ensures that the remaining set of temporal initialization points facilitates efficient encoding and decoding.

[0018]

[0025] In one or more examples, the video coder stores a respective set of temporal initialization points for a slice or picture in a buffer (e.g., after coding that slice or picture). The set of temporal initialization points may include one or more initialization points that may be context values ​​of a context, as described above, or that may be derived from context values ​​of a context used to code the current slice or current picture.

[0019]

[0026] For example, the video coder stores in the buffer a first set of temporal initialization points associated with a first slice or first picture (e.g., after coding a first slice or first picture), stores in the buffer a second set of temporal initialization points associated with a second slice or second picture (e.g., after coding a second slice or second picture), etc. In this manner, the buffer stores sets of temporal initialization points from two or more slices or two or more pictures for context-based arithmetic coding. Each set of temporal initialization points is associated with one slice or one picture of the two or more slices or two or more pictures.

[0020]

[0027] Each slice or picture may have a temporal identification (ID) value and / or a quantization parameter (QP) value. Each slice may also be associated with a slice type. The temporal ID value of a previous picture may indicate whether the previous picture can be used for inter-prediction of the current picture. For example, only pictures with a temporal ID value that is the same as or lower than the current picture may be used for inter-prediction of the current picture. In this way, based on bandwidth availability or processing power, pictures with a temporal ID value higher than a certain threshold may be dropped from the bitstream or not decoded without affecting the ability to decode pictures with temporal ID values ​​that are below the threshold. The QP value indicates the amount of quantization applied to the encoding process.

[0021]

[0028] In one or more examples, when a buffer that stores a set of temporal initialization points from two or more slices or two or more pictures for context-based arithmetic coding is full, the video coder determines (e.g., identifies) from the buffer a first set of temporal initialization points associated with a slice or picture from among the two or more slices or two or more pictures based on at least one of a slice type, a temporal identification value, or a quantization parameter (QP) value of the slice or picture. As an example, the video coder may determine (e.g., identify) from the buffer a first set of temporal initialization points associated with a slice or picture having at least one of a smallest temporal identification value or quantization parameter (QP) value from among the two or more slices or two or more pictures.

[0022]

[0029] In some examples, the first set of temporal initialization points may be associated with slices having a slice type different from the slice type of the current slice to be encoded or decoded. In some examples, multiple sets of temporal initialization points may be stored for a slice type. In such examples, the video coder may determine a group of sets of temporal initialization points associated with slices having the same slice type as the current slice. The video coder can determine a first set of temporal initialization points (e.g., temporal initialization points associated with slices or pictures having the smallest temporal identification value or QP value) from the group of sets of temporal initialization points associated with slices having the same slice type as the current slice.

[0023]

[0030] The video coder may remove from the buffer a first set of temporal initialization points associated with the slice or picture. The video coder may store in the buffer a second set of temporal initialization points associated with the current slice or current picture that are based on the determined one or more context values ​​of the current slice or current picture.

[0024]

[0031] In one or more examples, the video coder may determine (e.g., select) a set of temporal initialization points stored in a buffer and initialize one or more context values ​​of at least one context used to encode or decode a subsequent slice or subsequent picture based on the determined set of temporal initialization points. To determine the set of temporal initialization points, the video coder may determine a slice or picture having a temporal identification value and / or QP value closest to a temporal identification value and / or QP value of the subsequent slice or subsequent picture, and possibly having the same slice type. For example, if the QP value of the subsequent picture is X, the video coder may determine which set of temporal initialization points is associated with a picture having a QP value of X or a QP value closest to X. The video coder may select the determined set of temporal initialization points. The video coder may context-based arithmetic encode or decode the subsequent slice or subsequent picture.

[0025]

[0032] Possible coding gains may be obtained by removing the set of temporal initialization points associated with slices or pictures having the smallest temporal identification or QP values. In one or more examples, pictures with smaller temporal identification and / or QP values ​​tend to have more transform coefficients (e.g., due to less quantization) compared to pictures with higher temporal identification and / or QP values, which tend to have fewer transform coefficients. When there is a relatively larger number of transform coefficients (e.g., due to a lower QP value), the context value of the context may be updated relatively quickly and used for the remainder of the slice (e.g., the context may be adapted at the start of slice coding, and the remainder of the slice will be coded efficiently).

[0026]

[0033] However, when there is a relatively small number of transform coefficients (e.g., due to a higher QP value), the context values ​​of the context are updated relatively slowly, i.e., a slice with a higher QP value has fewer transform coefficients, the context adaptation is slower, and therefore fewer blocks of the slice are efficiently coded.

[0027]

[0034] In one or more examples, the set of temporal initialization points stored in the buffer may be closer to the context values ​​to be used for coding the subsequent slices or subsequent pictures (e.g., after some mapping or scaling). Thus, if the set of temporal initialization points stored in the buffer is associated with pictures having higher temporal identification values ​​or QP values, it is more likely that these sets of temporal initialization points will be usable for subsequent pictures having higher temporal identification values ​​or QP values. In other words, in some examples, higher coding efficiency gains may be realized for slices or pictures having higher temporal identification values ​​or QP values ​​using the temporal initialization points. Thus, by storing the temporal initialization points of pictures having higher temporal identification values ​​or QP values ​​and removing the temporal initialization points of pictures having lower temporal identification values ​​or QP values ​​due to buffer size limitations, the exemplary technique may promote coding efficiency without requiring a large size buffer.

[0028]

[0035] There may be additional problems using temporal initialization points. As explained above, each picture may be associated with a temporal identification (ID) value. However, if the temporal initialization point of a picture with a higher temporal ID value is to be used to determine the initialization point of the current picture, errors may exist because the temporal initialization point of the picture with the higher temporal ID value may not be available.

[0029]

[0036] As will be described in more detail, this disclosure describes example techniques in which a picture's temporal ID value is used to determine whether that picture's temporal initialization point can be utilized by a subsequent picture. In this manner, there is a reduction in the likelihood that a video decoder will rely on a temporal initialization point that is not available.

[0030]

[0037] In some examples, the temporal initialization points of a picture may vary from slice to slice. In such cases, it may be possible for the temporal initialization points of one slice of the current picture to overwrite the temporal initialization points of the slices of the previous picture when the temporal initialization points of the slices of the previous picture become available. This disclosure describes example techniques that minimize the adverse effects of overwriting temporal initialization point information and ensure that the correct temporal initialization points are stored.

[0031]

[0038] 1 is a block diagram illustrating an example video encoding and decoding system 100 that may implement techniques of this disclosure. The techniques of this disclosure are generally directed to coding (encoding and / or decoding) video data. In general, video data includes any data for processing video. Thus, video data may include raw uncoded video, coded video, decoded (e.g., reconstructed) video, and video metadata, such as signaling data.

[0032]

[0039] 1, in this example, system 100 includes a source device 102 that provides encoded video data to be decoded and displayed by a destination device 116. Specifically, source device 102 provides the video data to destination device 116 via a computer-readable medium 110. Source device 102 and destination device 116 may include any of a wide range of devices, including desktop computers, notebook (i.e., laptop) computers, mobile devices, tablet computers, set-top boxes, telephone handsets such as smartphones, televisions, cameras, display devices, digital media players, video gaming consoles, video streaming devices, broadcast receiver devices, and the like. In some cases, source device 102 and destination device 116 may be capable of wireless communication and thus may be referred to as wireless communication devices.

[0033]

[0040] In the example of FIG. 1, source device 102 includes video source 104, memory 106, video encoder 200, and output interface 108. Destination device 116 includes input interface 122, video decoder 300, memory 120, and display device 118. According to this disclosure, video encoder 200 of source device 102 and video decoder 300 of destination device 116 may be configured to apply techniques related to a temporal initialization point of one or more contexts used for context-based arithmetic coding of video data of a current picture or a current slice. Thus, source device 102 represents an example of a video encoding device, while destination device 116 represents an example of a video decoding device. In other examples, source device and destination device may include other components or configurations. For example, source device 102 may receive video data from an external video source, such as an external camera. Similarly, destination device 116 may interface with an external display device rather than including an integrated display device.

[0034]

[0041] The system 100 as shown in FIG. 1 is only an example. In general, any digital video encoding and / or decoding device may implement the techniques for one or more contexts used in context-based arithmetic coding of video data of a current picture or a current slice. The source device 102 and the destination device 116 are only examples of coding devices that generate the coded video data that the source device 102 transmits to the destination device 116. This disclosure refers to devices that perform coding (encoding and / or decoding) of data as “coding” devices. Thus, the video encoder 200 and the video decoder 300 represent examples of coding devices, specifically, video encoders and video decoders, respectively. In some examples, the source device 102 and the destination device 116 may operate in a substantially symmetrical manner, such that each of the source device 102 and the destination device 116 includes video encoding and decoding components. Thus, the system 100 may support one-way or two-way video transmission between the source device 102 and the destination device 116, for example, video streaming, video playback, video broadcasting, or video telephony.

[0035]

[0042] Generally, the video source 104 represents a source of video data (i.e., raw, unencoded video data) and provides a continuous series of pictures (also called "frames") of the video data to the video encoder 200, which encodes the picture data. The video source 104 of the source device 102 may include a video capture device such as a video camera, a video archive containing previously captured raw video, and / or a video feed interface that receives video from a video content provider. As a further alternative, the video source 104 may generate computer graphics-based data as the source video, or a combination of live video, archived video, and computer-generated video. In each case, the video encoder 200 encodes the captured video data, pre-captured video data, or computer-generated video data. The video encoder 200 may reorder the pictures from the order in which they are received (sometimes called the "display order") to a coding order for coding. The video encoder 200 may generate a bitstream that includes the encoded video data. The source device 102 may then output the encoded video data via the output interface 108 to a computer-readable medium 110, for receipt and / or retrieval by, for example, an input interface 122 of the destination device 116.

[0036]

[0043] The memory 106 of the source device 102 and the memory 120 of the destination device 116 represent general purpose memories. In some examples, the memories 106, 120 may store raw video data, e.g., raw video from the video source 104 and raw decoded video data from the video decoder 300. Additionally or alternatively, the memories 106, 120 may store software instructions executable by, e.g., the video encoder 200 and the video decoder 300, respectively. Although the memories 106 and 120 are shown in this example separately from the video encoder 200 and the video decoder 300, it should be understood that the video encoder 200 and the video decoder 300 may also include internal memories for functionally similar or equivalent purposes. Additionally, the memories 106, 120 may store, e.g., encoded video data output from the video encoder 200 and input to the video decoder 300. In some examples, portions of the memories 106, 120 may be allocated as one or more video buffers, for example, to store raw decoded video data and / or encoded video data.

[0037]

[0044] The computer-readable medium 110 may represent any type of medium or device capable of transferring encoded video data from the source device 102 to the destination device 116. In one example, the computer-readable medium 110 represents a communication medium that allows the source device 102 to transmit encoded video data directly to the destination device 116 in real time, for example, via a radio frequency network or a computer-based network. The output interface 108 may modulate a transmission signal including the encoded video data, and the input interface 122 may demodulate a received transmission signal according to a communication standard such as a wireless communication protocol. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful to facilitate communication from the source device 102 to the destination device 116.

[0038]

[0045] In some examples, source device 102 may output the encoded data from output interface 108 to storage device 112. Similarly, destination device 116 may access the encoded data from storage device 112 via input interface 122. Storage device 112 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, a Blu-ray disc, a DVD, a CD-ROM, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium that stores encoded video data.

[0039]

[0046] In some examples, source device 102 may output the encoded video data to a file server 114 or another intermediate storage device, which may store the encoded video data generated by source device 102. Destination device 116 may access the stored video data from file server 114 via streaming or download.

[0040]

[0047] The file server 114 may be any type of server device capable of storing encoded video data and transmitting the encoded video data to a destination device 116. The file server 114 may represent a web server (e.g., for a website), a server configured to provide file transfer protocol services (such as File Transfer Protocol (FTP) or File Delivery over Unidirectional Transport (FLUTE) protocol), a content delivery network (CDN) device, a hypertext transfer protocol (HTTP) server, a Multimedia Broadcast Multicast Service (MBMS) or Enhanced MBMS (eMBMS) server, and / or a network attached storage (NAS) device. The file server 114 may additionally or alternatively implement one or more HTTP streaming protocols, such as Dynamic Adaptive Streaming over HTTP (DASH), HTTP Live Streaming (HLS), Real Time Streaming Protocol (RTSP), HTTP Dynamic Streaming, etc.

[0041]

[0048] Destination device 116 may access the encoded video data from file server 114 through any standard data connection, including an Internet connection. This may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both suitable for accessing the encoded video data stored on file server 114. Input interface 122 may be configured to operate according to any one or more of the various protocols mentioned above to retrieve or receive media data from file server 114, or other such protocols to retrieve media data.

[0042]

[0049] The output interface 108 and the input interface 122 may represent a wireless transmitter / receiver, a modem, a wired network component (e.g., an Ethernet card), a wireless communication component operating according to any of the various IEEE 802.11 standards, or other physical components. In examples in which the output interface 108 and the input interface 122 include wireless components, the output interface 108 and the input interface 122 may be configured to transfer data, such as encoded video data, according to a cellular communication standard, such as 4G, 4G-LTE (Long Term Evolution), LTE Advanced, 5G, etc. In some examples in which the output interface 108 includes a wireless transmitter, the output interface 108 and the input interface 122 may be configured to transfer data, such as encoded video data, according to other wireless standards, such as the IEEE 802.11 specification, the IEEE 802.15 specification (e.g., ZigBee™), the Bluetooth™ standard, etc. In some examples, the source device 102 and / or the destination device 116 may include respective system-on-a-chip (SoC) devices. For example, the source device 102 may include a SoC device that performs functions attributed to the video encoder 200 and / or the output interface 108, and the destination device 116 may include a SoC device that performs functions attributed to the video decoder 300 and / or the input interface 122.

[0043]

[0050] The techniques of this disclosure may be applied to video coding to support any of a variety of multimedia applications, such as over-the-air television broadcast, cable television transmission, satellite television transmission, Internet streaming video transmission such as Dynamic Adaptive Streaming over HTTP (DASH), digital video encoded on a data storage medium, decoding of digital video stored on a data storage medium, or other applications.

[0044]

[0051] An input interface 122 of the destination device 116 receives an encoded video bitstream from a computer-readable medium 110 (e.g., a communications medium, a storage device 112, a file server 114, etc.). The encoded video bitstream may include signaling information defined by the video encoder 200 that is also used by the video decoder 300, such as syntax elements having values ​​that describe characteristics and / or processing of video blocks or other coded units (e.g., slices, pictures, groups of pictures, sequences, etc.). The display device 118 displays decoded pictures of the decoded video data to a user. The display device 118 may represent any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.

[0045]

[0052] Although not shown in FIG. 1, in some examples, the video encoder 200 and the video decoder 300 may each be integrated with an audio encoder and / or audio decoder and may include appropriate MUX-DEMUX units or other hardware and / or software to handle multiplexed streams containing both audio and video in a common data stream.

[0046]

[0053] The video encoder 200 and the video decoder 300 may each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the techniques are implemented partially in software, a device may store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to implement the techniques of this disclosure. Each of the video encoder 200 and the video decoder 300 may be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (CODEC) in the respective device. The device including the video encoder 200 and / or the video decoder 300 may include an integrated circuit, a microprocessor, and / or a wireless communication device such as a cellular phone.

[0047]

[0054] The video encoder 200 and the video decoder 300 may operate according to a video coding standard such as ITU-T H.265, also referred to as High Efficiency Video Coding (HEVC), or an extension standard thereof, such as multiview and / or scalable video coding extensions. Alternatively, the video encoder 200 and the video decoder 300 may operate according to other proprietary or industry standards, such as ITU-T H.266, also referred to as Generic Video Coding (VVC). In other examples, the video encoder 200 and the video decoder 300 may operate according to a proprietary video codec / format, such as AOMedia Video1 (AV1), an extension of AV1, and / or a successor version of AV1 (e.g., AV2). In other examples, the video encoder 200 and the video decoder 300 may operate according to other proprietary formats or industry standards. However, the techniques of this disclosure are not limited to any particular coding standard or format.

[0048]

[0055] In general, video encoder 200 and video decoder 300 may be configured to implement the techniques of this disclosure with any video coding technique that uses temporal initialization points of one or more contexts used for context-based arithmetic coding of the video data of a current picture or a current slice. For example, the multiple temporal initialization points are included in the video data of one or more previous slices or pictures that precede the current slice or picture in coding order.

[0049]

[0056] A temporal initialization point may be associated with a slice or a picture. Thus, there may be multiple sets of temporal initialization points for each slice or picture. For example, a first set of temporal initialization points may be associated with a first slice or a first picture, a second set of temporal initialization points may be associated with a second slice or a second picture, and so on.

[0050]

[0057] Each set of temporal initialization points may be a context value of at least one context of the slice or picture, or may be a value derived from a context value of at least one context of the slice or picture. When temporal initialization is enabled for context-based arithmetic coding, the video encoder 200 and the video decoder 300 may use a stored set of temporal initialization points (e.g., of a previously encoded or decoded slice or picture) to initialize a context value of at least one context for encoding or decoding a subsequent slice or subsequent picture. When the video encoder 200 and the video decoder 300 encode or decode a block of a subsequent slice or subsequent picture, the video encoder 200 and the video decoder 300 may update the context value after initialization using the set of initialization points.

[0051]

[0058] As will be described in more detail, in one or more examples, a buffer stores a set of temporal initialization points. However, the number of possible sets of temporal initialization points may be relatively large, and there may be size limitations on the buffer. Thus, the video encoder 200 and the video decoder 300 may be configured to implement buffer management in which the video encoder 200 and the video decoder 300 selectively determine which sets of initialization points to remove in order to make memory space for new sets of initialization points to be inserted into the buffer.

[0052]

[0059] This disclosure describes example techniques for such buffer management. For example, when a buffer is full, video encoder 200 and video decoder 300 may evaluate temporal identification (ID) values ​​and / or QP values ​​of slices or pictures, possibly along with slice types, using associated temporal initialization points stored in the buffer. For example, in addition to storing set temporal initialization points, the buffer may store information indicating the temporal ID values ​​and / or QP values ​​of slices or pictures, and possibly the slice type (e.g., I slice, B slice, P slice) associated with each set of temporal initialization points.

[0053]

[0060] The video encoder 200 and the video decoder 300 may compare the temporal ID values ​​and / or QP values ​​of the slices or pictures associated with the sets of temporal initialization points stored in the buffer. Based on this comparison, the video encoder 200 and the video decoder 300 may determine a set of temporal initialization points associated with a slice or picture based on the temporal ID value or QP value of the slice or picture. For example, the video encoder 200 and the video decoder 300 may determine a set of temporal initialization points associated with a slice or picture having at least one of the smallest temporal ID value or QP value. The video encoder 200 and the video decoder 300 may remove a first set of temporal initialization points to make memory space for a second set of temporal initialization points associated with a current slice or current picture (e.g., the slice or picture that was just encoded or decoded). The second set of temporal initialization points may be based on one or more context values ​​determined for the current slice or current picture.

[0054]

[0061] In one or more examples, the video encoder 200 and the video decoder 300 may determine the set of temporal initialization points to be removed based on the slice type as well. In some examples, if the buffer stores a set of temporal initialization points associated with a slice having the same slice type as the current slice, the video encoder 200 and the video decoder 300 may remove that set of temporal initialization points. In some cases, the video encoder 200 and the video decoder 300 may remove that set of temporal initialization points even if there is a set of temporal initialization points associated with a slice or picture with a lower temporal identification value and / or QP value.

[0055]

[0062] For example, to remove a set of temporal initialization points, the video encoder 200 and the video decoder 300 may first determine whether there exists a set of temporal initialization points associated with a slice having the same slice type as the slice type of the current slice. If there exists, the video encoder 200 and the video decoder 300 may remove that set of temporal initialization points. If there exists, the video encoder 200 and the video decoder 300 may add the set of temporal initialization points of the current slice to the buffer, assuming that the buffer is not full. In the above examples, the video encoder 200 and the video decoder 300 prioritized slice type, but the example techniques are not so limited. In some examples, to remove a set of temporal initialization points, the video encoder 200 and the video decoder 300 may first determine whether there exists a set of temporal initialization points associated with a slice having a temporal identification value or QP value that is the same as the temporal identification value or QP value of the current slice, regardless of slice type. If there is, the video encoder 200 and the video decoder 300 may remove that set of temporal initialization points.

[0056]

[0063] However, in cases where the buffer is full and a set of temporal initialization points should be removed, video encoder 200 and video decoder 300 may implement example techniques described in this disclosure, such as determining the set of temporal initialization points associated with the slice or picture having at least one of the smallest temporal ID value or QP value and then removing that set of temporal initialization points.

[0057]

[0064] As will also be described in more detail, video encoder 200 and video decoder 300 may determine (e.g., select) at least one set of temporal initialization points from the multiple sets of temporal initialization points based on respective temporal identification (ID) values ​​and / or quantization parameter (QP) values ​​of slices or pictures associated with the multiple set temporal initialization points and a temporal ID value or QP value of the current picture or current slice. In this manner, the likelihood of not having a temporal initialization point that should be available to the current picture, such as when a picture having a higher temporal ID value than the current picture's temporal ID value is removed from the bitstream or not processed, may be reduced.

[0058]

[0065] The video encoder 200 and the video decoder 300 may perform block-based coding of pictures. The term “block” generally refers to a structure that includes data to be processed (e.g., encoded, decoded, or otherwise used in an encoding and / or decoding process). For example, a block may include a two-dimensional matrix of samples of luminance and / or chrominance data. Generally, the video encoder 200 and the video decoder 300 may code video data represented in a YUV (e.g., Y, Cb, Cr) format. That is, rather than coding red, green, and blue (RGB) data for samples of a picture, the video encoder 200 and the video decoder 300 may code luminance and chrominance components, which may include both red and blue chrominance components. In some examples, the video encoder 200 converts received RGB format data to a YUV representation before encoding, and the video decoder 300 converts the YUV representation to the RGB format. Alternatively, pre-processing and post-processing units (not shown) may perform these conversions.

[0059]

[0066] This disclosure may generally refer to coding (e.g., encoding and decoding) a picture as including a process of encoding or decoding data for a picture. Similarly, this disclosure may refer to coding a block of a picture as including a process of encoding or decoding data for the block, e.g., predictive and / or residual coding. A coded video bitstream generally includes a series of values ​​of syntax elements that represent coding decisions (e.g., coding modes) and partitioning of a picture into blocks. Thus, references to coding a picture or a block should be understood generally as coding values ​​of the syntax elements that form the picture or block.

[0060]

[0067] HEVC defines various blocks, including coding units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video coder (such as the video encoder 200) partitions coding tree units (CTUs) into CUs according to a quadtree structure. That is, the video coder partitions CTUs and CUs into four equal non-overlapping squares, and each node of the quadtree has either zero or four child nodes. A node without child nodes may be called a "leaf node", and a CU of such a leaf node may include one or more PUs and / or one or more TUs. The video coder may further partition the PUs and TUs. For example, in HEVC, a residual quadtree (RQT) represents a partition of a TU. In HEVC, a PU represents inter-predicted data, and a TU represents residual data. An intra-predicted CU includes intra-prediction information, such as an intra-mode indication.

[0061]

[0068] As another example, the video encoder 200 and the video decoder 300 may be configured to operate according to VVC. According to VVC, a video coder (such as the video encoder 200) partitions a picture into multiple coding tree units (CTUs). The video encoder 200 may partition the CTUs according to a tree structure, such as a quadtree-binary tree (QTBT) structure or a Multi-Type Tree (MTT) structure. The QTBT structure eliminates the concept of multiple partition types, such as the separation between CUs, PUs, and TUs in HEVC. The QTBT structure includes two levels, a first level partitioned according to a quadtree partition, and a second level partitioned according to a binary tree partition. The root node of the QTBT structure corresponds to a CTU. The leaf nodes of the binary tree correspond to coding units (CUs).

[0062]

[0069] In the MTT partitioning structure, blocks may be partitioned using quadtree (QT) partitioning, binary tree (BT) partitioning, and one or more types of triple tree (TT) (also called ternary tree (TT)) partitioning. A triple tree partitioning or ternary tree partitioning is a partitioning in which a block is divided into three sub-blocks. In some examples, a triple tree partitioning or ternary tree partitioning divides a block into three sub-blocks without splitting the original block through the center. The partition types in MTT (e.g., QT, BT, and TT) may be symmetric or asymmetric.

[0063]

[0070] When operating according to the AV1 codec, the video encoder 200 and the video decoder 300 may be configured to code the video data in blocks. In AV1, the largest coding block that may be processed is called a superblock. In AV1, a superblock may be either 128×128 luma samples or 64×64 luma samples. However, in successor video coding formats (e.g., AV2), a superblock may be defined by a different (e.g., larger) luma sample size. In some examples, a superblock is the top level of a block quadtree. The video encoder 200 may further partition the superblock into smaller coding blocks. The video encoder 200 may partition the superblock and other coding blocks into smaller blocks using square or non-square partitioning. The non-square blocks may include N / 2×N, N×N / 2, N / 4×N, and N×N / 4 blocks. The video encoder 200 and the video decoder 300 may perform a separate prediction and transformation process for each of the coding blocks.

[0064]

[0071] AV1 also defines tiles of video data. A tile is a rectangular array of superblocks that may be coded independently of other tiles. That is, video encoder 200 and video decoder 300 may encode and decode coding blocks within a tile, respectively, without using video data from other tiles. However, video encoder 200 and video decoder 300 may perform filtering across tile boundaries. Tiles may be uniform or non-uniform in size. Tile-based coding may enable parallel processing and / or multi-threading for encoder and decoder implementations.

[0065]

[0072] In some examples, the video encoder 200 and the video decoder 300 may use a single QTBT or MTT structure to represent each of the luminance and chrominance components, while in other examples, the video encoder 200 and the video decoder 300 may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luminance component and another QTBT / MTT structure for both chrominance components (or two QTBT / MTT structures for each chrominance component).

[0066]

[0073] Video encoder 200 and video decoder 300 may be configured to use quadtree partitioning, QTBT partitioning, MTT partitioning, superblock partitioning, or other partitioning structures.

[0067]

[0074] In some examples, the CTU includes a coding tree block (CTB) of luma samples, two corresponding CTBs of chroma samples for a picture with three sample arrays, or a CTB of samples for a picture coded using three separate color planes and syntax structures used to code a monochrome picture or sample. The CTB may be an N×N block of samples for some value of N that partitions to split the components into CTBs. A component is an array or a single sample from one of the three arrays (luma and two chroma) that make up a picture in 4:2:0, 4:2:2, or 4:4:4 color format, or a single sample of an array or arrays that make up a picture in monochrome format. In some examples, the coding block is an M×N block of samples for some value of M and N that partitions to split the CTB into coding blocks.

[0068]

[0075] Blocks (e.g., CTUs or CUs) may be grouped in various ways within a picture. As an example, a brick may refer to a rectangular region of a CTU row in a particular tile in a picture. A tile can be a rectangular region of CTUs in a particular tile column and a particular tile row in a picture. A tile column refers to a rectangular region of CTUs with a height equal to the height of the picture and a width specified by a syntax element (e.g., in a picture parameter set). A tile row refers to a rectangular region of CTUs with a height specified by a syntax element (e.g., in a picture parameter set) and a width equal to the width of the picture.

[0069]

[0076] In some examples, a tile may be partitioned into multiple bricks, each of which may include one or more CTU rows within the tile. A tile that is not partitioned into multiple bricks may also be referred to as a brick. However, a brick that is a true subset of a tile may not be referred to as a tile. Bricks within a picture may also be arranged as slices. A slice may be an integer number of bricks of a picture that may be contained exclusively within a single network abstraction layer (NAL) unit. In some examples, a slice includes either several complete tiles or only a continuous sequence of complete bricks of one tile.

[0070]

[0077] This disclosure may use "NxN" and "N by N", e.g., 16x16 samples or 16 by 16 samples, interchangeably to refer to the sample dimensions of a block (such as a CU or other video block) in the vertical and horizontal dimensions. In general, a 16x16 CU has 16 samples in the vertical direction (y=16) and 16 samples in the horizontal direction (x=16). Similarly, an NxN CU generally has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non-negative integer value. Samples within a CU may be arranged in rows and columns. Moreover, a CU may not necessarily have to have the same number of samples in the horizontal direction as in the vertical direction. For example, a CU may include NxM samples, where M is not necessarily equal to N.

[0071]

[0078] The video encoder 200 encodes video data for a CU that represents prediction and / or residual information, as well as other information. The prediction information indicates how the CU will be predicted to form a predictive block for the CU. The residual information generally represents sample-by-sample differences between samples of the CU before encoding and the predictive block.

[0072]

[0079] To predict a CU, the video encoder 200 may generally form a predictive block for the CU through inter prediction or intra prediction. Inter prediction generally refers to predicting a CU from data of a previously coded picture, and intra prediction generally refers to predicting a CU from previously coded data of the same picture. To perform inter prediction, the video encoder 200 may generate a predictive block using one or more motion vectors. The video encoder 200 may generally perform a motion search to identify a reference block that closely matches the CU in terms of the difference between the CU and the reference block, for example. The video encoder 200 may calculate a difference metric using a sum of absolute difference (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared differences (MSD), or other such difference calculation to determine whether the reference block closely matches the current CU. In some examples, the video encoder 200 may predict the current CU using unidirectional prediction or bidirectional prediction.

[0073]

[0080] Some examples of VVC also provide an affine motion compensation mode, which may be considered an inter-prediction mode. In an affine motion compensation mode, video encoder 200 may determine two or more motion vectors that represent non-translational motion, such as zooming in or out, rotation, perspective movement, or other irregular motion types.

[0074]

[0081] To perform intra prediction, the video encoder 200 may determine an intra prediction mode to generate a predictive block. Some examples of VVC provide 67 intra prediction modes, including various directional modes, as well as a planar mode and a DC mode. In general, the video encoder 200 determines an intra prediction mode that describes neighboring samples for a current block (e.g., a block of a CU) from which to predict samples of the current block. Assuming that the video encoder 200 codes the CTUs and CUs in raster scan order (left to right, top to bottom), such samples may generally be above, above and to the left, or to the left of the current block in the same picture as the current block.

[0075]

[0082] The video encoder 200 encodes data representing a prediction mode for the current block. For example, in the case of an inter prediction mode, the video encoder 200 may encode data representing which of various available inter prediction modes is used as well as motion information for the corresponding mode. In the case of unidirectional or bidirectional inter prediction, for example, the video encoder 200 may encode a motion vector using an advanced motion vector prediction (AMVP) mode or a merge mode. The video encoder 200 may use a similar mode to encode a motion vector for an affine motion compensation mode.

[0076]

[0083] AV1 includes two general techniques for encoding and decoding coding blocks of video data. The two general techniques are intra-prediction (e.g., intra-frame prediction or spatial prediction) and inter-prediction (e.g., inter-frame prediction or temporal prediction). In the context of AV1, when predicting a block of a current frame of video data using an intra-prediction mode, the video encoder 200 and the video decoder 300 do not use video data from other frames of the video data. In most intra-prediction modes, the video encoder 200 encodes a block of the current frame based on a difference between a sample value in the current block and a predicted value generated from a reference sample in the same frame. The video encoder 200 determines a predicted value generated from a reference sample based on the intra-prediction mode.

[0077]

[0084] Following prediction, such as intra- or inter-prediction, of a block, the video encoder 200 may compute residual data for the block. The residual data, such as a residual block, represents sample-by-sample differences between a block and a prediction block for that block formed using a corresponding prediction mode. The video encoder 200 may apply one or more transforms to the residual block to generate transform data in a transform domain rather than the sample domain. For example, the video encoder 200 may apply a discrete cosine transform (DCT), an integer transform, a wavelet transform, or a conceptually similar transform to the residual video data. In addition, the video encoder 200 may apply a secondary transform, such as a mode-dependent non-separable secondary transform (MDNSST), a signal-dependent transform, or a Karhunen-Loeve transform (KLT), following the initial transform. The video encoder 200 generates transform coefficients following application of the one or more transforms.

[0078]

[0085] As mentioned above, following any transformation that generates transform coefficients, the video encoder 200 may perform quantization of the transform coefficients. Quantization generally refers to a process in which transform coefficients are quantized to possibly reduce the amount of data used to represent the transform coefficients, providing further compression. By performing a quantization process, the video encoder 200 may reduce the bit depth associated with some or all of the transform coefficients. For example, the video encoder 200 may truncate an n-bit value to an m-bit value during quantization, where n is greater than m. In some examples, to perform quantization, the video encoder 200 may perform a bitwise right shift of the value to be quantized.

[0079]

[0086] Following quantization, the video encoder 200 may scan the transform coefficients, generating a one-dimensional vector from the two-dimensional matrix including the quantized transform coefficients. The scan may be designed to place transform coefficients with higher energy (and therefore lower frequency) at the front of the vector and transform coefficients with lower energy (and therefore higher frequency) at the rear of the vector. In some examples, the video encoder 200 may use a predefined scan order to scan the quantized transform coefficients to generate a serialized vector and then entropy code the quantized transform coefficients of the vector. In other examples, the video encoder 200 may perform an adaptive scan. After scanning the quantized transform coefficients to form the one-dimensional vector, the video encoder 200 may entropy code the one-dimensional vector, for example, according to context-adaptive binary arithmetic coding (CABAC). The video encoder 200 may also entropy code values ​​for syntax elements that describe metadata associated with the encoded video data for use by the video decoder 300 in decoding the video data.

[0080]

[0087] To implement CABAC, the video encoder 200 may assign one or more context values ​​of a context in a context model to a symbol to be transmitted. The context value may, for example, relate to whether neighboring values ​​of the symbol are zeroed out or not. A probability decision (e.g., a context value) may be based on the context assigned to the symbol.

[0081]

[0088] Video encoder 200 may further generate syntax data, such as block-based syntax data, picture-based syntax data, and sequence-based syntax data, for example within a picture header, a block header, a slice header, or other syntax data, such as a sequence parameter set (SPS), a picture parameter set (PPS), or a video parameter set (VPS), to video decoder 300. Video decoder 300 may similarly decode such syntax data to determine how to decode corresponding video data.

[0082]

[0089] In this manner, video encoder 200 may generate a bitstream including encoded video data, e.g., syntax elements that describe partitions of a picture into blocks (e.g., CUs) and prediction and / or residual information for the blocks. Finally, video decoder 300 may receive the bitstream and decode the encoded video data.

[0083]

[0090] In general, video decoder 300 performs a reciprocal process to that performed by video encoder 200 to decode encoded video data of a bitstream. For example, video decoder 300 may decode values ​​for syntax elements of a bitstream using CABAC in a manner substantially similar to, but reciprocal to, the CABAC encoding process of video encoder 200. The syntax elements may define partition information for the partition of a picture into CTUs and the partition of each CTU according to a corresponding partition structure, such as a QTBT structure, to define CUs of the CTU. The syntax elements may further define prediction and residual information for blocks of video data (e.g., CUs).

[0084]

[0091] The residual information may be represented, for example, by quantized transform coefficients. The video decoder 300 may dequantize and inverse transform the quantized transform coefficients of the block to reconstruct a residual block for the block. The video decoder 300 uses the signaled prediction mode (intra-prediction or inter-prediction) and associated prediction information (e.g., motion information for inter-prediction) to form a predictive block for the block. The video decoder 300 may then combine (sample by sample) the predictive block and the residual block to reconstruct the original block. The video decoder 300 may perform additional processing, such as performing a deblocking process to reduce visual artifacts along block boundaries.

[0085]

[0092] This disclosure may generally refer to "signaling" some information, such as a syntax element. The term "signaling" may generally refer to communication of values ​​for syntax elements and / or other data used to decode encoded video data. That is, video encoder 200 may signal values ​​for syntax elements in a bitstream. In general, signaling refers to generating values ​​in a bitstream. As mentioned above, source device 102 may forward the bitstream to destination device 116 in substantially real-time or non-real-time, which may occur, for example, when storing syntax elements in storage device 112 for later retrieval by destination device 116.

[0086]

[0093] JVET-Y0181: "AHG12: CABAC initialization from previous inter slice", Seregin et al., Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29, 25th Teleconference, January 12-21, 2022, describes a method of using a previous encoding / decoding order CABAC (context-adaptive binary arithmetic coding) initialization point for CABAC initialization of a current picture or a current slice. Although described with respect to CABAC, the exemplary techniques are applicable to context-based arithmetic coding. For example, CABAC arithmetic coding may require a starting point for each context, where the starting point is referred to as an initialization point. This starting point (e.g., initialization point) may include one or more context states, window or rate adaptation parameters, and other parameters required for the arithmetic coding operation. For example, there may be one or more initialization points for one or more context values ​​of a context, and the one or more initialization points may be values ​​(e.g., initial values) of the one or more context values.

[0087]

[0094] In video codecs, these initialization points are typically predefined and known to both the video encoder 200 and the video decoder 300; for example, in HEVC and VVC, the initialization points are defined in initialization tables for each slice type, I slices, P slices, and B slices.

[0088]

[0095] In temporal CABAC initialization, in addition to or instead of using predefined initialization points, initialization points may be stored within a particular CTU of a picture or slice, and those stored initialization points may be used to initialize a next picture or next slice instead of or in addition to the predefined initialization points. That is, in some examples, the temporal initialization points may be for one or more contexts contained within the video data of one or more previous slices or pictures preceding the current slice or picture in coding order and used for context-based arithmetic coding of the video data of the current slice or picture.

[0089]

[0096] The CTU in which the video encoder 200 or the video decoder 300 may store the initialization point(s) may be variable, and the information indication of the CTU may be signaled. As an example, the video encoder 200 and the video decoder 300 may store an initialization point for a slice in the middle of a picture. That is, when the video encoder 200 is encoding a slice or the video decoder 300 is decoding a slice, when the video encoder 200 is encoding or the video decoder 300 is decoding a CTU in the middle of the slice, the video encoder 200 or the video decoder 300 may store the current context value or values ​​of the context as the initialization point to be used for encoding or decoding syntax elements in a subsequent slice or a subsequent picture. For example, when the video encoder 200 or the video decoder 300 begins to encode or decode a subsequent slice or a subsequent picture, the video encoder 200 or the video decoder 300 may set an initial value of the context value of the context equal to or based on (e.g., using mapping, scaling, weighting, etc.) the stored initialization point. Then, as the video encoder 200 and the video decoder 300 are encoding or decoding subsequent slices or subsequent pictures, the video encoder 200 and the video decoder 300 may update the context value from the initial value based on more recently encoded or decoded video data.

[0090]

[0097] In some examples, the video encoder 200 and the video decoder 300 may store the initialization point after coding the CABAC state, window, and other parameters of a CTU. For example, storing means storing the CABAC state, window, and other parameters for initialization for a CTU after coding. For example, parameters adapted to the coded content in a previous picture may be better at representing the starting initialization point of the video data of the current picture than a predefined initialization point.

[0091]

[0098] The storing may be performed separately for each slice type and slice quantization parameter (QP). In some examples, an initialization with the same slice type and the same QP as the current slice may be used for the CABAC initialization of the current slice.

[0092]

[0099] There may be some problems with temporal initialization points for context-based arithmetic coding. For example, because slices of the current picture may be decoded independently and pictures with a temporal identification (ID) value higher than a threshold may be removed as part of temporal scalability, techniques that utilize temporal initialization points may be improved to enable temporal scalability while ensuring that a usable temporal initialization point is available.

[0093]

[0100] In some examples, the initialization points may be stored from previous pictures in coding order, e.g., in a FIFO buffer, where the first initialization point is from a previous picture closest to the current picture. As an example, the buffer can store a first set of initialization points for one or more context values ​​of a context for a first slice or first picture, a second set of initialization points for one or more context values ​​of a context for a second slice or second picture, etc.

[0094]

[0101] An index may be introduced to indicate which initialization point is used from the FIFO buffer, and such initialization index is signaled in the bitstream, for example, in a picture or slice header. For example, the video encoder 200 may signal to the buffer, and the video decoder 300 may receive, an index indicating which set of initialization points should be used to initialize one or more context values ​​of the context of the currently decoded slice or picture. The video decoder 300 may retrieve a set of initialization points based on the index and initialize one or more context values ​​of the context. For example, the video decoder 300 may set the initial values ​​of the one or more context values ​​equal to the retrieved set of initialization points. As another example, the video decoder 300 may map, scale, weight, or perform some other operation on the retrieved set of initialization points to determine the initial values ​​of the one or more context values.

[0095]

[0102] The following describes temporal scalability and the use of temporal initialization points for temporal scalability. Every coded picture may have a temporal ID value assigned by video encoder 200. Video encoder 200 may signal the temporal ID value in a network abstraction layer unit (NALU) header. The temporal ID value is used for temporal scalability, and some pictures may be ignored (e.g., removed from the bitstream or not processed) and other pictures may be decoded without the removed or not processed pictures.

[0096]

[0103] In one example, temporal scalability may be achieved by setting a restriction that a picture with a lower temporal ID value cannot use a picture with a higher temporal ID value for inter prediction. In this case, since a picture with a higher temporal ID value cannot be used for inter prediction of a picture with a lower temporal ID value, the video decoder 300 may be able to decode a picture with a lower temporal ID value without using a picture with a higher temporal ID value. Thus, a picture with a higher temporal ID value may be removed from the bitstream or may not be processed (e.g., may be ignored). When temporal initialization for context-based arithmetic coding is applied, the video encoder 200 and the video decoder 300 may utilize the temporal ID value to identify a stored initialization point available for the current slice / picture initialization.

[0097]

[0104] In one or more examples, the video encoder 200 and the video decoder 30 may store a temporal ID value along with the initialization points. For example, each set of initialization points, including one or more initialization points, may be associated with a slice or a picture. The video encoder 200 and the video decoder 300 may store a set of initialization points and information indicating a temporal ID value of a slice or a picture associated with the set of initialization points. As an example, a first set of initialization points may be associated with a first slice or a first picture having a first temporal ID value, and a second set of initialization points may be associated with a second slice or a second picture having a second temporal ID value. The video encoder 200 and the video decoder 300 may store the first set of initialization points and the first temporal ID value, and information that the first temporal ID value is for a first slice or a first picture associated with the first set of initialization points (e.g., associating the first temporal ID value with the first set of initialization points). The video encoder 200 and the video decoder 300 may store the second set of initialization points and the second temporal ID value, and information that the second temporal ID value is for a second slice or second picture associated with the second set of initialization points (e.g., associating the second temporal ID value with the second set of initialization points).

[0098]

[0105] To determine (e.g., select) an initialization point based on the temporal ID value of the stored set of initialization points, video encoder 200 and video decoder 300 may compare the temporal ID value with the current slice / picture temporal ID. For example, video encoder 200 and video decoder 300 may select at least one set of temporal initialization points from the multiple sets of temporal initialization points based on respective temporal ID values ​​associated with the multiple sets of temporal initialization points and a temporal ID value of the current picture or current slice.

[0099]

[0106] A set of initialization points having a temporal ID less than or equal to the current temporal ID is selected as available. For example, to determine the at least one set of temporal initialization points, video encoder 200 and video decoder 300 may determine a group of temporal initialization points among a plurality of sets of temporal initialization points, where each temporal ID value for the group of temporal initialization points is less than or equal to the temporal ID value of the current picture or current slice. Video encoder 200 and video decoder 300 may select a set of temporal initialization points from the group of temporal initialization points.

[0100]

[0107] In one example, only a set of initialization points having the same time ID as the current slice / picture time ID are selected as available initialization points. For example, a smaller time ID typically has a lower QP, so an initialization point from a previous picture having a lower time ID value may not adequately represent an initialization point of a slice of the current picture. In some examples, to determine at least one set of temporal initialization points, the video encoder 200 and the video decoder 300 may determine a group of temporal initialization points among a plurality of sets of temporal initialization points, and each of the temporal initialization points in the group of temporal initialization points has a time ID value equal to the time ID value of the current picture or the current slice. The video encoder 200 and the video decoder 300 may select at least one set of temporal initialization points from the group of temporal initialization points.

[0101]

[0108] In another example, an initialization point with the same time ID is searched for, and if no such initialization point is available, a lower time ID is checked, e.g., the current time ID value minus 1, and if that is not available, the current time ID value minus 2, etc. The set of initialization points found is used for initialization.

[0102]

[0109] For example, to determine the at least one set of temporal initialization points, the video encoder 200 and the video decoder 300 may determine that none of the temporal initialization points in the multiple sets of temporal initialization points have a temporal ID value equal to the temporal ID value of the current picture or current slice. Based on a determination that none of the temporal initialization points in the multiple sets of temporal initialization points have a temporal ID value equal to the temporal ID value of the current picture or current slice, the video encoder 200 and the video decoder 300 may determine whether any of the temporal initialization points in the multiple sets of temporal initialization points have a temporal ID value that is one less than the temporal ID value of the current picture or current slice. Based on a determination that there are one or more temporal initialization points having a temporal ID value that is one less than the temporal ID value of the current picture or current slice, the video encoder 200 and the video decoder 300 may select at least one set of temporal initialization points from the one or more sets of temporal initialization points having a temporal ID value that is one less than the temporal ID value of the current picture or current slice.

[0103]

[0110] Similarly, instead of using the same QP initialization, when it is not available, QP-1 or QP+1 are searched and used if stored, if not available, QP-2 or QP+2 are checked, etc. Whatever is found is used for initialization. Time ID and QP search may be combined when the same time ID and QP initialization point is not available.

[0104]

[0111] For example, in one or more examples, the video encoder 200 and the video decoder 30 may store a QP value along with the initialization point. For example, each set of initialization points, including one or more initialization points, may be associated with a slice or a picture. The video encoder 200 and the video decoder 300 may store a set of initialization points and information indicating a QP value of a slice or a picture associated with the set of initialization points. As an example, a first set of initialization points may be associated with a first slice or a first picture having a first QP value, and a second set of initialization points may be associated with a second slice or a second picture having a second QP value. The video encoder 200 and the video decoder 300 may store the first set of initialization points and the first QP value, and information that the first QP value is for the first slice or the first picture associated with the first set of initialization points (e.g., associating the first QP value with the first set of initialization points). Video encoder 200 and video decoder 300 may store the second set of initialization points and the second QP value, and information that the second QP value is for a second slice or second picture associated with the second set of initialization points (e.g., associating the second QP value with the second set of initialization points).

[0105]

[0112] In one example, only the set of initialization points having the same QP value as the current slice / picture QP value are selected as available initialization points. In some examples, to determine the at least one set of temporal initialization points, the video encoder 200 and the video decoder 300 may determine a group of temporal initialization points among a plurality of sets of temporal initialization points, where each QP value for the temporal initialization points in the group of temporal initialization points is equal to the QP value of the current picture or the current slice. The video encoder 200 and the video decoder 300 may select the at least one set of temporal initialization points from the group of temporal initialization points.

[0106]

[0113] In another example, an initialization point with the same QP value is searched for, and if such an initialization point is not available, the closest QP value, e.g., the current QP value plus or minus 1, is checked, and if that is not available, the current QP value plus or minus 2 is checked, etc. The set of initialization points found is used for initialization.

[0107]

[0114] For example, to determine the at least one set of temporal initialization points, the video encoder 200 and the video decoder 300 may determine that none of the temporal initialization points in the multiple sets of temporal initialization points has a QP value equal to the QP value of the current picture or current slice. Based on a determination that none of the temporal initialization points in the multiple sets of temporal initialization points has a QP value equal to the QP value of the current picture or current slice, the video encoder 200 and the video decoder 300 may determine whether any of the temporal initialization points in the multiple sets of temporal initialization points has a QP value that is one less than or one greater than the QP value of the current picture or current slice. Based on a determination that there are one or more temporal initialization points that have a QP value that is one less than or one greater than the QP value of the current picture or current slice, the video encoder 200 and the video decoder 300 may select the at least one set of temporal initialization points from the one or more sets of temporal initialization points that have a QP value that is one less than or one greater than the QP value of the current picture or current slice.

[0108]

[0115] Thus, in one or more examples, the video encoder 200 and the video decoder 300 may be configured to determine (e.g., select) a set of temporal initialization points stored in a buffer and initialize one or more context values ​​of at least one context used to encode or decode a subsequent slice or a subsequent picture based on the selected set of temporal initialization points. The video encoder 200 and the video decoder 300 may context-based arithmetic encode or decode the subsequent slice or a subsequent picture. For example, the video encoder 200 and the video decoder 300 may set initial values ​​of the context values ​​based on the initialization points and use the initial values ​​to encode or decode one or more syntax values ​​of the subsequent slice or a subsequent picture. The video encoder 200 and the video decoder 300 may update the context values ​​from the initial values.

[0109]

[0116] One exemplary method of selecting a set of temporal initialization points may include the video encoder 200 and the video decoder 300 determining a temporal identification value of a subsequent slice or subsequent picture. The video encoder 200 and the video decoder 300 may determine that none of the two or more slices or two or more pictures having an associated set of temporal initialization points stored in the buffer has a temporal identification value equal to the temporal identification value of the subsequent slice or subsequent picture. In this case, the video encoder 200 and the video decoder 300 may determine, from among the two or more slices or two or more pictures, a slice or picture having a temporal identification value that is closest to and less than the temporal identification value of the subsequent slice or subsequent picture. To select a set of temporal initialization points, the video encoder 200 and the video decoder 300 may select a set of temporal initialization points associated with the determined slice or picture having a temporal identification value that is closest to and less than the temporal identification value of the subsequent slice or subsequent picture.

[0110]

[0117] Another exemplary method of selecting a set of temporal initialization points may include the video encoder 200 and the video decoder 300 determining a QP value of a subsequent slice or subsequent picture. The video encoder 200 and the video decoder 300 may determine that none of the two or more slices or two or more pictures having an associated set of temporal initialization points stored in the buffer has a QP value equal to the QP value of the subsequent slice or subsequent picture. In this case, the video encoder 200 and the video decoder 300 may determine a slice or picture from among the two or more slices or two or more pictures that has a QP value closest to the QP value of the subsequent slice or subsequent picture. In this case, the closest QP value may be greater than the QP value of the current slice or current picture. To select a set of temporal initialization points, the video encoder 200 and the video decoder 300 may select a set of initialization points associated with a slice or picture that has a QP value closest to the QP value of the subsequent slice or subsequent picture.

[0111]

[0118] In the above examples, the video encoder 200 and the video decoder 300 may determine a set of initialization points associated with a slice or picture having a temporal identification value that is closest to the temporal identification value of the subsequent slice or picture without being greater than the temporal identification value of the current slice and / or associated with a slice or picture having a QP value that is closest to the QP value of the subsequent slice or picture. In some examples, the video encoder 200 and the video decoder 300 may also consider the slice type in determining which set of temporal initialization points to use. For example, the set of temporal initialization points may be associated with a slice type that is the same as the slice type of the subsequent slice.

[0112]

[0119] The following describes the storage of initialization points. As described above, slices of the same picture can be decoded independently (i.e., the next slice of the same picture cannot depend on the previous slice of the same picture). Therefore, the initialization point stored in the previous slice of the same picture cannot be used for any slice of the same picture.

[0113]

[0120] To achieve this, in one example, the initialization points (e.g., in a set of initialization points, which may include one or more initialization points) are temporarily stored in a temporary buffer and added to the storage buffer from the temporary buffer only after all slices of the same picture have been processed (coded, decoded, parsed). In this manner, the initialization points in the temporary buffer may be updated until all slices of the same picture have been processed, and then the storage buffer receives the initialization points from the temporary buffer.

[0114]

[0121] When an initialization point is added, the previously stored initialization point may be removed / replaced so that the buffer has limitations, and if there is an update while a slice of the same picture is being processed, the preferred initialization point of the previous picture may be replaced by the previous slice of the same picture. That is, if a temporary buffer is not utilized, the initialization point in the storage buffer that should be retained may be overwritten. By utilizing a temporary buffer, it may be possible to avoid overwriting the initialization point in the storage buffer that should be retained, and to overwrite the storage buffer with the initialization point in the temporary buffer only after a slice of the same picture has been processed.

[0115]

[0122] Therefore, it may be beneficial to store the initialization point separately in a temporary buffer and update the storage buffer with the temporary buffer only after all slices of the picture have been processed.To check the end of a picture, the CTU address may be checked to see if it is equal to the last CTU, or the slice index may be checked to see if it is the last slice of the picture.

[0116]

[0123] For example, the video encoder 200 and the video decoder 300 may store in a first buffer one or more temporal initialization points of one or more contexts used for context-based arithmetic coding of the video data of one or more previous pictures. The video encoder 200 and the video decoder 300 may store in a second buffer temporal initialization points (e.g., a set of temporal initialization points) of one or more contexts used for context-based arithmetic coding of the video data of a slice of a current picture, where storing in the second buffer includes storing in the second buffer during coding of the video data of the current picture. The video encoder 200 and the video decoder 300 may store in the first buffer the temporal initialization points stored in the second buffer following processing a last coding tree unit (CTU) or slice of the current picture.

[0117]

[0124] In another example, the initialization point may be signaled within adaptation parameter sets (APS), e.g., similar to an adaptive loop filter parameter set that carries filter coefficients. APS handles temporal ID values ​​and may already have a defined restriction not to use an APS derived from the current picture that is applied to the same picture coding.

[0118]

[0125] In one or more examples, the initialization points may be stored in various picture locations, such as the center or end of the picture, although other locations within the picture may be used as well.

[0119]

[0126] The selection of which storage location is used may be signaled within the bitstream. In one example, video encoder 200 may signal such an indication, and video decoder 300 may parse such an indication in a picture header or slice header, or any other parameter set, or elsewhere. Because the picture header is shared by all slices of a picture, the picture header may be an exemplary signaling location, and it may not be necessary to signal the storage selection within each slice to achieve picture-level adaptation.

[0120]

[0127] There may be multiple syntax elements signaled in the picture or slice header for temporal CABAC, such as whether temporal initialization is applied to a picture or slice, storage location signaling, or how to remove entries from the storage initialization buffer. Such syntax element signaling in the picture or slice header may be conditioned by higher level indications, for example, by syntax signaled in the SPS or PPS that indicates whether temporal CABAC initialization is used.

[0121]

[0128] The syntax element signaling in the picture or slice header may also depend on whether an initialization entry exists in the buffer that can be used for the slice. If no such entry exists in the buffer, temporal initialization may not be applied and any syntax elements associated with temporal initialization may not be signaled in the slice or picture header.

[0122]

[0129] The entry identification may be performed by comparing the temporal ID and / or QP value of the slice to be coded with the values ​​of the entries in the buffer. That is, in one or more examples, the video decoder 300 may compare the temporal identification value and / or QP value of the slice or picture being decoded with the temporal identification value and / or QP value associated with each set of temporal initialization points to select a set of initialization points to be used to initialize context values ​​of the context used to encode or decode the slice or picture.

[0123]

[0130] In some examples, storing it (e.g., initialization points, although other information for storage may be possible) for every time ID and QP may be expensive, such as requiring a large buffer size, so the initialization storage may be limited by a certain number of entries for implementation purposes. For example, the buffer size may be limited to N per slice type. In such a case, when the buffer becomes full, i.e., when all N entries have been added to the buffer, one entry should be removed before adding the next entry. In one example, the number N may be set equal to 5, since it represents a typical GOP32 coding.

[0124]

[0131] Entry removal may be performed by video encoder 200 and video decoder 300 according to certain rules based on the buffer entry's temporal ID value and / or QP value. For example, the rules may be that the entry with the smallest temporal ID is removed, and / or the entry with the smallest temporal ID and smallest QP is removed.

[0125]

[0132] In some examples, if there is an entry in the set of temporal initialization points associated with a slice having the same slice type as the subsequent slice, video encoder 200 and video decoder 300 may remove that entry in the set of temporal initialization points even if there is an entry in the set of initialization points associated with a slice or picture with a lower temporal ID or lower QP in the buffer. In some examples, the entry with the smallest temporal ID and / or smallest QP may be for a set of temporal initialization points associated with a slice having a slice type different from the slice type of the subsequent slice.

[0126]

[0133] In some examples, there may be multiple entries for one slice type. For example, for slice type I slices, the buffer may store up to five sets of temporal initialization points, for slice type P slices, the buffer may store up to five sets of temporal initialization points, and for slice type B slices, the buffer may store up to five sets of temporal initialization points. In such examples, the video encoder 200 and the video decoder 300 may first determine a set of temporal initialization points associated with a slice type that is the same as the slice type of the subsequent slice. The video encoder 200 and the video decoder 300 may then determine, within this determined set of temporal initialization points, a set of temporal initialization points with a minimum temporal identification value and / or QP value, and determine that the set of temporal initialization points with a minimum temporal identification value and / or QP value from within the set of temporal initialization points with the same slice type as the subsequent slice should be removed.

[0127]

[0134] The removal of the entry with the lowest QP may be due to the lowest QP slice having more transform coefficients (less quantization), so the context may be adapted at the beginning of the slice coding and the remainder of the slice will be coded efficiently. Slices with higher QP have fewer transform coefficients and the context adaptation may be slower, so fewer blocks of the slice will be coded efficiently.

[0128]

[0135] In other words, having a temporal initialization point allows for the initialization of context values, which may then be adapted as part of encoding or decoding. In cases where the context values ​​of a slice can be adapted relatively quickly (e.g., slices having lower QP values), there may be some benefit in having a temporal initialization point. However, such benefit may be reduced because the context values ​​of slices with lower QP values ​​may be adapted relatively quickly, even if not properly initialized. For slices whose context values ​​do not adapt relatively quickly (e.g., slices with higher QP values), there may be more benefit in properly initializing the context values.

[0129]

[0136] As an example, assume that the context values ​​of a first slice with a lower QP value tend to adapt quickly, and the context values ​​of a second slice with a higher QP value tend not to adapt quickly. If the time initialization point of the first slice is available, there may be some benefit in speeding up the adaptation of the context values. However, such benefit may not be large because the context values ​​of the first slice tend to adapt quickly even if not properly initialized.

[0130]

[0137] If a temporal initialization point for the second slice is available, there may be a greater benefit in speeding up the adaptation of the context values ​​compared to the first slice. For example, by initializing the context values ​​for the second slice, the adaptation of the context values ​​starts from a value closer to the final value, and therefore there is better compression of the blocks in the second slice compared to if the temporal initialization point was not available.

[0131]

[0138] As explained above, slices with higher QP values ​​tend to be slices whose context values ​​adapt more slowly. Thus, if the temporal initialization points of slices with higher QP values ​​are stored in a buffer, such temporal initialization points will be available for future slices that have more benefit from having temporal initialization points. As a result, retaining initialization entries with higher QPs may provide efficient coding for future pictures.

[0132]

[0139] Similarly, lower temporal ID slices are typically coded with smaller QPs, so that lower temporal ID entries may be removed. For example, similar to above, having a temporal initialization point allows for the initialization of context values ​​that may then be adapted as part of the encoding or decoding. If the context values ​​of a slice can be adapted relatively quickly (e.g., the slice has a lower temporal identification value), there may be some benefit in having a temporal initialization point. However, such benefit may be reduced because the context values ​​of slices with lower temporal identification values ​​may be adapted relatively quickly, even if not properly initialized. For slices whose context values ​​are not adapted relatively quickly (e.g., slices with higher temporal identification values), there may be more benefit in properly initializing the context values.

[0133]

[0140] As an example, assume that the context value of a first slice having a lower temporal identification value tends to adapt quickly, and the context value of a second slice having a higher temporal identification value tends not to adapt quickly. If the temporal initialization point of the first slice is available, there may be some benefit in speeding up the adaptation of the context value. However, such benefit may not be large because the context value of the first slice tends to adapt quickly even if it is not properly initialized.

[0134]

[0141] If a temporal initialization point for the second slice is available, there may be a greater benefit in speeding up the adaptation of the context values ​​compared to the first slice. For example, by initializing the context values ​​for the second slice, the adaptation of the context values ​​starts from a value closer to the final value, and therefore there is better compression of the blocks in the second slice compared to if the temporal initialization point was not available.

[0135]

[0142] As explained above, slices with higher temporal identification values ​​tend to be slices whose context values ​​adapt more slowly. Therefore, if the temporal initialization points of slices with higher temporal identification values ​​are stored in the buffer, such temporal initialization points will be available for future slices that have more benefits from having temporal initialization points. As a result, retaining initialization entries with higher temporal IDs may provide efficient coding for future pictures.

[0136]

[0143] Other rules based on temporal ID values, QP values, or slice types are possible, and this disclosure is not limited to the example rules for temporal ID values ​​and QP values, or slice types. The choice of rule may be signaled in the bitstream, in the picture header or slice header, or any other parameter set, or elsewhere.

[0137]

[0144] Thus, in one or more examples, the video encoder 200 and the video decoder 300 may be configured to process video data. To process the video data, the video encoder 200 and the video decoder 300 may be configured to determine one or more context values ​​of at least one context used to encode or decode a current slice or a current picture.

[0138]

[0145] Video encoder 200 and video decoder 300 may determine that a buffer that stores a set of temporal initialization points from two or more slices or two or more pictures for context-based arithmetic coding is full (e.g., there are N entries in a buffer that can store N entries). As described above, each set of temporal initialization points is associated with one slice or one picture of the two or more slices or two or more pictures and includes one or more temporal initialization points.

[0139]

[0146] The video encoder 200 and the video decoder 300 may determine (e.g., identify) a first set of temporal initialization points associated with a slice or picture from among the two or more slices or two or more pictures based on at least one of a slice type, a temporal identification value, or a quantization parameter (QP) value of the slice or picture. The video encoder 200 and the video decoder 300 may remove the first set of temporal initialization points associated with the slice or picture from the buffer and store a second set of temporal initialization points associated with the current slice or current picture in the buffer. The second set of temporal initialization points may be based on the determined one or more context values ​​(e.g., the second set of temporal initialization points are equal to or generated from one or more context values ​​of the context).

[0140]

[0147] As an example, to determine the first set of temporal initialization points (e.g., the set of temporal initialization points to be removed), the video encoder 200 and the video decoder 300 may determine a first set of temporal initialization points associated with a slice or picture having at least one of a minimum temporal identification value or a quantization parameter (QP) value from among two or more slices or two or more pictures. For example, the video encoder 200 and the video decoder 300 may determine a slice or picture having a minimum temporal identification value from among the temporal identification values ​​of the two or more slices or two or more pictures, or determine a slice or picture having a minimum QP value. In some cases, a slice associated with the first set of temporal initialization points may have a slice type different from the slice type of the current slice from among the QP values ​​of the two or more slices or two or more pictures.

[0141]

[0148] In some examples, if the buffer is full, or in some cases even if the buffer is not full, video encoder 200 and video decoder 300 may first remove duplicate entries before removing a set of initialization points associated with a slice or picture having the smallest temporal identification value or QP value. For example, if the buffer is full, or in some cases even if the buffer is not full, before storing a set of initialization points associated with a current slice or current picture in the buffer, video encoder 200 and video decoder 300 may determine whether any set of initialization points exists in the buffer that is associated with a slice or picture having the same temporal identification value or QP value as the current slice or current picture and / or has the same slice type.

[0142]

[0149] If a set of initialization points associated with a slice or picture having the same temporal identification value, QP value, or slice type as the current slice or current picture exists in the buffer, video encoder 200 and video decoder 300 may remove that set of initialization points even if there are other sets of initialization points associated with slices or pictures having lower temporal identification values ​​or QP values. If a set of initialization points associated with a slice or picture having the same temporal identification value, QP value, or slice type as the current slice or current picture does not exist in the buffer, video encoder 200 and video decoder 300 may remove the set of initialization points with the lowest temporal identification value or QP value.

[0143]

[0150] In some examples, the set of initialization points with the smallest temporal identification value or QP value to be removed may have a different slice type than the current slice. In some examples, such as when multiple sets of temporal initialization points may be stored for a slice type, the video encoder 200 and the video decoder 300 may determine a group of sets of temporal initialization points associated with slices having the same slice type as the current slice. The video encoder 200 and the video decoder 300 may then determine the set of temporal initialization points in this group with the smallest temporal identification value or QP value as the set of temporal initialization points to be removed.

[0144]

[0151] Thus, the video encoder 200 and the video decoder 300 may determine that at least one of the temporal identification values ​​or QP values ​​of the current slice or the current picture is different from the temporal identification values ​​or QP values ​​of each of two or more slices or pictures having an associated set of temporal initialization points stored in the buffer. In such an example, the video encoder 200 and the video decoder 300 may remove the first set of temporal initialization points based on a determination that at least one of the temporal identification values ​​or QP values ​​of the current slice or the current picture is different from the temporal identification values ​​or QP values ​​of each of the two or more slices or pictures. For example, the temporal identification values ​​or QP values ​​of the slices or pictures associated with the first set of temporal initialization points are different from the temporal identification values ​​or QP values ​​of the current slice or the current picture.

[0145]

[0152] For example, assume that the current slice or picture is a first slice or picture. In this example, the video encoder 200 and the video decoder 300 may determine one or more context values ​​of at least one context used to encode or decode the second slice or picture. The video encoder 200 and the video decoder 300 may determine a third slice or picture from the two or more slices or pictures that has at least one of the temporal identification value or QP value that is the same as the temporal identification value or QP value of the second slice or picture. In this example, the video encoder 200 and the video decoder 300 may remove a third set of temporal initialization points associated with the third slice or picture from the buffer and store in the buffer a fourth set of temporal initialization points associated with the second slice or picture that is based on the determined one or more context values ​​of the at least one context used to encode or decode the second slice or picture.

[0146]

[0153] The initialization point may include several parameters, such as multiple context states and multiple adaptation rates or adaptation windows (to indicate how quickly the context states can be adapted after each binning). The initialization storage buffer may be reduced by storing quantized values ​​of the parameters to reduce the dynamic range of possible values, which may require fewer bits to store them. In one example, the video encoder 200 and the video decoder 300 may store a sum of states and a sum of adaptation rates. The video encoder 200 and the video decoder 300 may assign the sum of states (average state) divided by the number of states to a state value. The video encoder 200 and the video decoder 300 may assign the sum of adaptation rates (average adaptation rate) divided by the number of adaptation rates to an adaptation rate (adaptation window) value.

[0147]

[0154] In another example, the video encoder 200 and the video decoder 300 may store only certain parameters (e.g., only some of the initialization points). In such an example, the video encoder 200 and the video decoder 300 may initialize other parameters from default values. For example, only one state and one adaptation rate are stored. When the initialization is performed, the stored values ​​are assigned to the first state and the first adaptation rate, respectively, and the second state and the second adaptation rate are assigned from default initialization values ​​that may already be stored in the codec (e.g., the video encoder 200 and the video decoder 300), such as the initialization points stored for each I-slice, P-slice, B-slice described above. Other value packing mechanisms applied to the state and adaptation rate values ​​may be applied as well.

[0148]

[0155] The following describes multiple initialization points. Multiple initialization points may be stored per picture. That is, for a picture or slice, there may be a set of initialization points, which includes one initialization point or multiple initialization points. In one example, multiple initialization points may be used when more than one slice is used in a picture. For example, an initialization point may be stored for each slice.

[0149]

[0156] Slices have relative positions within a picture, so initialization points are stored from the previous picture that correspond to the current slice position. In one example, initialization points are stored in the center of each slice in the previous picture and used to initialize the corresponding current slice.

[0150]

[0157] The slice boundaries of the current and previous pictures do not have to be aligned. In one example, the location of storing the initialization point is guided by the current picture slice boundary.

[0151]

[0158] In another example, an initialization point(s) is stored for each slice of the current picture at a certain location, and subsequent pictures can determine which initialization to use by certain rules. Such rules may, for example, be to check the coordinates where the initialization was stored and compare whether such location belongs to the current slice, and if so, such initialization may be used. In another example, initialization points for each slice may be stored in a buffer, and an index is signaled for the current slice to identify which initialization point should be used.

[0152]

[0159] For example, video encoder 200 and video decoder 300 may store a previous temporal initialization point of one or more contexts used for context-based arithmetic coding of the video data of a previous slice in a previous picture. For a current slice, video encoder 200 and video decoder 300 may determine that the current slice in the current picture has a location in the current picture that corresponds to a location of the previous slice in the previous picture. Based on the current slice having a location in the current picture that corresponds to a location of the previous slice in the previous picture, video encoder 200 and video decoder 300 may determine a current temporal initialization point of the current slice based on the previous temporal initialization point.

[0153]

[0160] 2 is a block diagram illustrating an example video encoder 200 that may implement techniques of this disclosure. FIG. 2 is provided for illustrative purposes and should not be considered as limiting the techniques broadly illustrated and described in this disclosure. For illustrative purposes, this disclosure describes a video encoder 200 according to VVC (ITU-T H.266 under development) and HEVC (ITU-T H.265) techniques. However, the techniques of this disclosure may be implemented by video encoding devices configured for other video coding standards and video coding formats, such as AV1 and successors of the AV1 video coding format.

[0154]

[0161] 2, the video encoder 200 includes a video data memory 230, a mode selection unit 202, a residual generation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a filter unit 216, a decoded picture buffer (DPB) 218, and an entropy coding unit 220. Any or all of the video data memory 230, the mode selection unit 202, the residual generation unit 204, the transform processing unit 206, the quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a filter unit 216, a DPB 218, and an entropy coding unit 220 may be implemented in one or more processors or processing circuits. For example, the units of the video encoder 200 may be implemented as one or more circuits or logic elements as part of a hardware circuit, or as part of a processor, an ASIC, or an FPGA. Moreover, video encoder 200 may include additional or alternative processors or processing circuitry that perform these and other functions.

[0155]

[0162] The video data memory 230 may store video data to be encoded by the components of the video encoder 200. The video encoder 200 may receive the video data stored in the video data memory 230, for example, from the video source 104 (FIG. 1). The DPB 218 may function as a reference picture memory that stores reference video data for use in predicting subsequent video data by the video encoder 200. The video data memory 230 and the DPB 218 may be formed by any of a variety of memory devices, such as DRAM, including synchronous dynamic random access memory (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The video data memory 230 and the DPB 218 may be provided by the same memory device or separate memory devices. In various examples, the video data memory 230 may be on-chip with the other components of the video encoder 200, as shown, or may be off-chip relative to those components.

[0156]

[0163] In this disclosure, references to video data memory 230 should not be construed as limited to memory internal to video encoder 200, unless specifically stated as such, or to memory external to video encoder 200, unless specifically stated as such. Rather, references to video data memory 230 should be understood as a reference memory that stores video data that video encoder 200 receives for encoding (e.g., video data for a current block to be encoded). Memory 106 of FIG. 1 may also provide temporary storage of outputs from various units of video encoder 200.

[0157]

[0164] The various units in FIG. 2 are illustrated to aid in understanding the operations performed by the video encoder 200. The units may be implemented as fixed-function circuits, programmable circuits, or a combination thereof. A fixed-function circuit refers to a circuit that provides a specific function, and the operations that may be performed are predefined. A programmable circuit refers to a circuit that may be programmed to perform various tasks, and provides flexible functionality in the operations that may be performed. For example, a programmable circuit may execute software or firmware that causes the programmable circuit to operate in a manner defined by the software or firmware instructions. Although a fixed-function circuit may execute software instructions (e.g., receive a parameter or output a parameter), the type of operation that the fixed-function circuit performs is generally unchanged. In some examples, one or more of the units may be different circuit blocks (fixed function or programmable), and in some examples, one or more of the units may be an integrated circuit.

[0158]

[0165] Video encoder 200 may include arithmetic logic units (ALUs), elementary function units (EFUs), digital circuits, analog circuits, and / or a programmable core formed from programmable circuits. In examples in which the operations of video encoder 200 are implemented using software executed by programmable circuits, memory 106 (FIG. 1) may store software instructions (e.g., object code) that video encoder 200 receives and executes, or another memory (not shown) within video encoder 200 may store such instructions.

[0159]

[0166] The video data memory 230 is configured to store the received video data. The video encoder 200 may retrieve pictures of the video data from the video data memory 230 and provide the video data to the residual generation unit 204 and the mode selection unit 202. The video data in the video data memory 230 may be raw video data to be encoded.

[0160]

[0167] The mode selection unit 202 includes a motion estimation unit 222, a motion compensation unit 224, and an intra prediction unit 226. The mode selection unit 202 may include additional functional units that perform video prediction according to other prediction modes. By way of example, the mode selection unit 202 may include a palette unit, an intra block copy unit (which may be part of the motion estimation unit 222 and / or the motion compensation unit 224), an affine unit, a linear model (LM) unit, etc.

[0161]

[0168] The mode selection unit 202 generally coordinates multiple encoding passes to test combinations of encoding parameters and the resulting rate-distortion values ​​for such combinations. The encoding parameters may include partitioning of the CTU into CUs, prediction modes for the CUs, transform types for the residual data of the CUs, quantization parameters for the residual data of the CUs, etc. The mode selection unit 202 may ultimately select a combination of encoding parameters that has a better rate-distortion value than the other tested combinations.

[0162]

[0169] Video encoder 200 may partition a picture retrieved from video data memory 230 into a series of CTUs and encapsulate one or more CTUs within a slice. Mode selection unit 202 may partition the CTUs of a picture according to a tree structure, such as the MTT structure, QTBT structure, superblock structure, or quadtree structure described above. As described above, video encoder 200 may form one or more CUs from partitioning the CTUs according to the tree structure. Such CUs may also be generally referred to as "video blocks" or "blocks."

[0163]

[0170] In general, the mode selection unit 202 also controls its components (e.g., the motion estimation unit 222, the motion compensation unit 224, and the intra prediction unit 226) to generate a prediction block for a current block (e.g., a current CU, or in HEVC, an overlapping portion of a PU and a TU). In the case of inter prediction of a current block, the motion estimation unit 222 may perform a motion search to identify one or more closely matching reference blocks among one or more reference pictures (e.g., one or more previously coded pictures stored in the DPB 218). In particular, the motion estimation unit 222 may calculate a value representing how similar a potential reference block is to the current block according to, for example, a sum of absolute differences (SAD), a sum of squared differences (SSD), a mean absolute difference (MAD), a mean squared difference (MSD), etc. The motion estimation unit 222 may generally perform these calculations using a sample-by-sample difference between the current block and the reference block under consideration. Motion estimation unit 222 may identify the reference block having the lowest value resulting from these calculations, which indicates the reference block that most closely matches the current block.

[0164]

[0171] The motion estimation unit 222 may form one or more motion vectors (MVs) that define the position of a reference block in a reference picture relative to the position of a current block in a current picture. The motion estimation unit 222 may then provide the motion vectors to the motion compensation unit 224. For example, in the case of unidirectional inter prediction, the motion estimation unit 222 may provide a single motion vector, while in the case of bidirectional inter prediction, the motion estimation unit 222 may provide two motion vectors. The motion compensation unit 224 may then generate a predictive block using the motion vectors. For example, the motion compensation unit 224 may use the motion vectors to retrieve data of a reference block. As another example, if the motion vectors have fractional sample precision, the motion compensation unit 224 may interpolate values ​​for the predictive block according to one or more interpolation filters. Moreover, in the case of bidirectional inter prediction, the motion compensation unit 224 may retrieve data for two reference blocks identified by the respective motion vectors and combine the retrieved data, for example, through a sample-wise average or weighted average.

[0165]

[0172] The motion estimation unit 222 and the motion compensation unit 224, when operating according to the AV1 video coding format, may be configured to encode coding blocks of video data (e.g., both luma coding blocks and chroma coding blocks) using translational motion compensation, affine motion compensation, overlapped block motion compensation (OBMC), and / or synthetic inter-intra prediction.

[0166]

[0173] As another example, in the case of intra prediction or intra predictive coding, intra prediction unit 226 may generate a predictive block from samples neighboring the current block. For example, in the case of a directional mode, intra prediction unit 226 may generally mathematically combine values ​​of neighboring samples and populate these calculated values ​​in a defined direction across the current block to generate a predictive block. As another example, in the case of a DC mode, intra prediction unit 226 may calculate an average of neighboring samples for the current block and generate a predictive block to include this resulting average for each sample of the predictive block.

[0167]

[0174] When operating according to the AV1 video coding format, the intra prediction unit 226 may be configured to encode coding blocks of video data (e.g., both luma coding blocks and chroma coding blocks) using directional intra prediction, non-directional intra prediction, recursive filter intra prediction, chroma-from-luma (CFL) prediction, intra block copy (IBC), and / or color palette modes. The mode selection unit 202 may include additional functional units that perform video prediction according to other prediction modes.

[0168]

[0175] The mode select unit 202 provides the prediction block to a residual generation unit 204. The residual generation unit 204 receives a raw, uncoded version of the current block from the video data memory 230 and the prediction block from the mode select unit 202. The residual generation unit 204 calculates sample-by-sample differences between the current block and the prediction block. The resulting sample-by-sample differences define a residual block for the current block. In some examples, the residual generation unit 204 may also determine differences between sample values ​​in the residual block to generate the residual block using residual differential pulse code modulation (RDPCM). In some examples, the residual generation unit 204 may be formed using one or more subtractor circuits that perform binary subtraction.

[0169]

[0176] In an example where the mode selection unit 202 partitions a CU into PUs, each PU may be associated with a luma prediction unit and a corresponding chroma prediction unit. The video encoder 200 and the video decoder 300 may support PUs having various sizes. As mentioned above, the size of a CU may refer to the size of the luma coding block of the CU, and the size of a PU may refer to the size of the luma prediction unit of the PU. Assuming that the size of a particular CU is 2N×2N, the video encoder 200 may support a PU size of 2N×2N or N×N for intra prediction, and a symmetric PU size of 2N×2N, 2N×N, N×2N, N×N, or similar for inter prediction. The video encoder 200 and the video decoder 300 may also support asymmetric partitioning of PU sizes of 2N×nU, 2N×nD, nL×2N, and nR×2N for inter prediction.

[0170]

[0177] In examples where the mode select unit 202 does not further partition the CUs into PUs, each CU may be associated with a luma coding block and a corresponding chroma coding block. As noted above, the size of a CU may refer to the size of the luma coding block of the CU. The video encoder 200 and the video decoder 300 may support CU sizes of 2N×2N, 2N×N, or N×2N.

[0171]

[0178] For other video coding techniques, such as intra block copy mode coding, affine mode coding, and linear model (LM) mode coding, as some examples, the mode select unit 202 generates a predictive block for the current block being coded via a respective unit associated with the coding technique. In some examples, such as palette mode coding, the mode select unit 202 may not generate a predictive block, but instead generate syntax elements that indicate how to reconstruct the block based on a selected palette. In such modes, the mode select unit 202 may provide these syntax elements to the entropy coding unit 220 to be coded.

[0172]

[0179] As described above, the residual generation unit 204 receives video data for a current block and a corresponding predictive block. The residual generation unit 204 then generates a residual block for the current block. To generate the residual block, the residual generation unit 204 calculates sample-by-sample differences between the predictive block and the current block.

[0173]

[0180] Transform processing unit 206 applies one or more transforms to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). Transform processing unit 206 may apply various transforms to the residual block to form the transform coefficient block. For example, transform processing unit 206 may apply a discrete cosine transform (DCT), a directional transform, a Karhunen-Loeve transform (KLT), or a conceptually similar transform to the residual block. In some examples, transform processing unit 206 may perform multiple transforms on the residual block, e.g., a linear transform and a secondary transform, such as a rotation transform. In some examples, transform processing unit 206 does not apply a transform to the residual block.

[0174]

[0181] When transform processing unit 206 operates according to AV1, it may apply one or more transforms to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). Transform processing unit 206 may apply various transforms to the residual block to form the transform coefficient block. For example, transform processing unit 206 may apply a horizontal / vertical transform combination, which may include a discrete cosine transform (DCT), an asymmetric discrete sine transform (ADST), an inverse ADST (e.g., ADST in reverse order), and an identity transform (IDTX). When using an identity transform, the transform is skipped in one of the vertical or horizontal directions. In some examples, the transform process may be skipped.

[0175]

[0182] The quantization unit 208 may quantize the transform coefficients in the transform coefficient block to generate a quantized transform coefficient block. The quantization unit 208 may quantize the transform coefficients of the transform coefficient block according to a quantization parameter (QP) value associated with the current block. The video encoder 200 (e.g., via the mode selection unit 202) may adjust the degree of quantization applied to the transform coefficient block associated with the current block by adjusting the QP value associated with the CU. Quantization may result in loss of information, and thus the quantized transform coefficients may be less accurate than the original transform coefficients generated by transform processing unit 206.

[0176]

[0183] Inverse quantization unit 210 and inverse transform processing unit 212 may apply inverse quantization and inverse transform, respectively, to the quantized transform coefficient block to reconstruct a residual block from the transform coefficient block. Reconstruction unit 214 may generate a reconstructed block that corresponds to the current block (possibly with some distortion) based on the reconstructed residual block and the predictive block generated by mode selection unit 202. For example, reconstruction unit 214 may add samples of the reconstructed residual block to corresponding samples from the predictive block generated by mode selection unit 202 to generate the reconstructed block.

[0177]

[0184] Filter unit 216 may perform one or more filter operations on the reconstructed blocks. For example, filter unit 216 may perform a deblocking operation to reduce blockiness artifacts along edges of a CU. The operations of filter unit 216 may be skipped in some examples.

[0178]

[0185] When operating according to AV1, filter unit 216 may perform one or more filter operations on the reconstructed blocks. For example, filter unit 216 may perform a deblocking operation to reduce blockiness artifacts along the edges of a CU. In other examples, filter unit 216 may apply a constrained directional enhancement filter (CDEF), which may be applied after deblocking and may include application of a non-separable, non-linear, low-pass directional filter based on the estimated edge direction. Filter unit 216 may also include a loop restoration filter, which may be applied after CDEF and may include a separable symmetric normalized Wiener filter or a dual auto-induced filter.

[0179]

[0186] Video encoder 200 stores the reconstructed blocks in DPB 218. For example, in examples where the operations of filter unit 216 are not performed, reconstruction unit 214 may store the reconstructed blocks in DPB 218. In examples where the operations of filter unit 216 are performed, filter unit 216 may store the filtered reconstructed blocks in DPB 218. Motion estimation unit 222 and motion compensation unit 224 may retrieve reference pictures formed from the reconstructed (and possibly filtered) blocks from DPB 218 to inter predict blocks of a later-encoded picture. In addition, intra prediction unit 226 may use the reconstructed blocks of the current picture in DPB 218 to intra predict other blocks in the current picture.

[0180]

[0187] In general, the entropy encoding unit 220 may entropy encode syntax elements received from other functional components of the video encoder 200. For example, the entropy encoding unit 220 may entropy encode quantized transform coefficient blocks from the quantization unit 208. As another example, the entropy encoding unit 220 may entropy encode predictive syntax elements (e.g., motion information for inter prediction or intra mode information for intra prediction) from the mode selection unit 202. The entropy encoding unit 220 may perform one or more entropy encoding operations on syntax elements, which are another example of video data, to generate entropy encoded data. For example, entropy encoding unit 220 may perform a context-adaptive variable length coding (CAVLC) operation, a CABAC operation, a variable-to-variable (V2V) coding operation, a syntax-based context-adaptive binary arithmetic coding (SBAC) operation, a Probability Interval Partitioning Entropy (PIPE) coding operation, an Exponential-Golomb coding operation, or another type of entropy coding operation on the data. In some examples, entropy encoding unit 220 may operate in a bypass mode in which syntax elements are not entropy coded.

[0181]

[0188] The video encoder 200 may output a bitstream that includes entropy coding syntax elements used to reconstruct blocks of a slice or picture. In particular, the entropy coding unit 220 may output the bitstream.

[0182]

[0189] The entropy encoding unit 220 may be configured as a symbol-to-symbol adaptive multi-symbol arithmetic coder in accordance with AV1. The syntax elements in AV1 include an alphabet of N elements, and the context (e.g., a probability model) includes a set of N probabilities. The entropy encoding unit 220 may store the probabilities as n-bit (e.g., 15-bit) cumulative distribution functions (CDFs). The entropy encoding unit 22 may perform recursive scaling with an update factor based on the alphabet size to update the context.

[0183]

[0190] The operations described above are described with respect to blocks. Such descriptions should be understood as operations on luma coding blocks and / or chroma coding blocks. As described above, in some examples, the luma coding blocks and chroma coding blocks are luma and chroma components of a CU. In some examples, the luma coding blocks and chroma coding blocks are luma and chroma components of a PU.

[0184]

[0191] In some examples, operations performed with respect to luma coding blocks may not be repeated for chroma coding blocks. As an example, operations of identifying motion vectors (MVs) and reference pictures for luma coding blocks may not be repeated to identify MVs and reference pictures for chroma blocks. Rather, the MVs of luma coding blocks may be scaled to determine the MVs of chroma blocks, and the reference pictures may be the same. As another example, the intra prediction process may be the same for luma coding blocks and chroma coding blocks.

[0185]

[0192] In one or more examples, in the case of entropy coding, such as context-based arithmetic coding, the entropy coding unit 220 may determine context values ​​for one or more contexts. During the coding of a slice or picture, the entropy coding unit 220 may update the context values ​​(e.g., probability values). However, at the start of a slice or picture, the context values ​​may be undefined. Rather than starting with undefined context values, the entropy coding unit 220 may use a predefined initialization point (e.g., an initial value stored in the DPB 218, the video data memory 230, or some other memory) to initialize one or more context values.

[0186]

[0193] In one or more examples, rather than or in addition to using a predefined initialization point, entropy encoding unit 220 may utilize one or more context values ​​of a previously encoded slice or picture, or a mapped, scaled, weighted version, etc. of one or more context values, of a previously encoded slice or picture, as the initialization point. For example, entropy encoding unit 220 may store one or more context values ​​of a context of a coded slice or picture (e.g., after coding the last CTU of a slice or picture) as a set of initialization points in a buffer (e.g., DPB 218, video data memory 230, or some other memory). Entropy encoding unit 220 may utilize the set of initialization points (e.g., context values, or values ​​based on context values ​​of a previously encoded slice or picture) to initialize context values ​​of a context used to code a subsequent slice or subsequent picture.

[0187]

[0194] Entropy encoding unit 220 may store sets of initialization points for multiple previously encoded slices or pictures. For example, a buffer may store a first set of initialization points associated with a first slice or first picture, a second set of initialization points associated with a second slice or second picture, etc. In addition, to determine (e.g., select) which set of initialization points entropy encoding unit 220 may use for a subsequent slice or subsequent picture, entropy encoding unit 220 may also store information of temporal identification values ​​and / or QP values ​​of the slices or pictures associated with each of the respective sets of initialization points.

[0188]

[0195] The video encoder 200 represents an example of a device configured to encode video data, including a memory configured to store video data and one or more processing units implemented in a circuit and configured to perform the example techniques described in this disclosure. For example, in one or more examples, when a buffer is full, the entropy encoding unit 220 may determine which set of initialization points to remove to make space for the just determined set of initialization points. For example, the entropy encoding unit 220 may determine one or more context values ​​of at least one context used to encode the current slice or current picture. The entropy encoding unit 220 may determine that a buffer storing a set of temporal initialization points from two or more slices or two or more pictures for context-based arithmetic coding is full. As described, each set of temporal initialization points is associated with one slice or one picture of the two or more slices or two or more pictures and includes one or more temporal initialization points.

[0189]

[0196] The entropy encoding unit 220 may determine a first set of temporal initialization points associated with a slice or picture from among the two or more slices or two or more pictures based on a temporal identification value or a quantization parameter (QP) value of the slice or picture. The entropy encoding unit 220 may remove the first set of temporal initialization points associated with the slice or picture from the buffer and store a second set of temporal initialization points associated with the current slice or current picture in the buffer. The second set of temporal initialization points is based on the determined one or more context values ​​(e.g., equal to one or more context values ​​determined for at least one context of the current slice or derived from one or more context values ​​of the picture or at least one context of the current slice or picture).

[0190]

[0197] Thus, instead of using when a slice or picture was coded as the sole factor for determining (e.g., selecting) which set of initialization points should be removed from the buffer, entropy coding unit 220 may utilize a temporal identification value and / or a QP value to determine which set of initialization points to remove, possibly including slice type. For example, to determine a first set of temporal initialization points, entropy coding unit 220 may determine a first set of temporal initialization points associated with a slice or picture having at least one of a minimum temporal identification value or a quantization parameter (QP) value from among two or more slices or two or more pictures. For example, entropy coding unit 220 may determine a slice or picture having a minimum temporal identification value from among the temporal identification values ​​of two or more slices or two or more pictures, or determine a slice or picture having a minimum QP value from among the QP values ​​of two or more slices or two or more pictures. Entropy coding unit 220 may then remove the set of initialization points associated with the determined slice or picture (e.g., the one having the minimum temporal identification value or QP value) from the buffer.

[0191]

[0198] In one or more examples, the entropy encoding unit 220 may remove a set of initialization points associated with a slice or picture having a minimum temporal identification value or QP value if none of the sets of initialization points are associated with a slice or picture having a temporal identification value or QP value that is the same as the temporal identification value or QP value of the current slice or current picture. For example, the entropy encoding unit 220 may determine that at least one of the temporal identification values ​​or QP values ​​of the current slice or current picture is different from the temporal identification value or QP value of each of the two or more slices or two or more pictures. In this example, the entropy encoding unit 220 may remove a first set of temporal initialization points based on a determination that at least one of the temporal identification values ​​or QP values ​​of the current slice or current picture is different from the temporal identification value or QP value of each of the two or more slices or two or more pictures. For example, the temporal identification value or QP value of the slice or picture associated with the first set of temporal initialization points is different from the temporal identification value or QP value of the current slice or current picture.

[0192]

[0199] As an example, assume that the current slice or picture is a first slice or picture. In this example, the entropy encoding unit 220 may determine one or more context values ​​of at least one context used to encode the second slice or picture, and determine a third slice or picture from the two or more slices or pictures that has at least one of the temporal identification value or QP value that is the same as the temporal identification value or QP value of the second slice or picture. In this example, the entropy encoding unit 220 may remove a third set of temporal initialization points associated with the third slice or picture from the buffer, and store in the buffer a fourth set of temporal initialization points associated with the second slice or picture that is based on the determined one or more context values ​​of the at least one context used to encode the second slice or picture.

[0193]

[0200] Then, to encode a subsequent picture, the entropy encoding unit 220 may determine (e.g., select) a set of temporal initialization points stored in the buffer and initialize one or more context values ​​of at least one context used to encode the subsequent slice or subsequent picture based on the selected set of temporal initialization points. For example, the entropy encoding unit 220 may assign initial values ​​for the one or more context values ​​equal to the selected set of temporal initialization points or derive initial values ​​for the one or more context values ​​based on the selected set of temporal initialization points (e.g., using mapping, scaling, weighting, etc.). The entropy encoding unit 220 may context-based arithmetic encode the subsequent slice or subsequent picture. For example, the entropy encoding unit 220 may utilize the initial values ​​for the context values ​​to encode syntax elements for the subsequent slice or subsequent picture and update the context values ​​during encoding of the subsequent slice or subsequent picture.

[0194]

[0201] In one or more examples, the video encoder 200 may also be configured to store a plurality of temporal initialization points of one or more contexts used for context-based arithmetic coding of the video data of the current picture or the current slice, the temporal initialization points being included within the video data of one or more previous pictures preceding the current picture in coding order, store a respective temporal identification (ID) value associated with each of the plurality of temporal initialization points, select at least one temporal initialization point of the plurality of temporal initialization points based on the respective temporal ID value associated with the plurality of temporal initialization points and the temporal ID value of the current picture or the current slice, and context-based arithmetic code the video data of the current picture or the current slice based on the selected at least one temporal initialization point.

[0195]

[0202] The video encoder 200 may be configured to store in a first buffer one or more temporal initialization points of one or more contexts used for context-based arithmetic coding of video data of one or more previous pictures and store in a second buffer one or more temporal initialization points of one or more contexts used for context-based arithmetic coding of video data of a slice of a current picture, where storing in the second buffer includes storing in the second buffer during coding of the video data of the current picture and storing the temporal initialization points stored in the second buffer in the first buffer following processing a last coding tree unit (CTU) or slice of the current picture.

[0196]

[0203] The video encoder 200 may be configured to store previous temporal initialization points of one or more contexts used for context-based arithmetic coding of video data of a previous slice in a previous picture, determine for a current slice that the current slice in the current picture has a location in the current picture that corresponds to a location of the previous slice in the previous picture based on the current slice having a location in the current picture that corresponds to a location of the previous slice in the previous picture, and determine a current temporal initialization point for the current slice based on the previous temporal initialization point.

[0197]

[0204] 3 is a block diagram illustrating an example video decoder 300 that may implement techniques of this disclosure. FIG. 3 is provided for purposes of illustration and not to limit the techniques broadly illustrated and described in this disclosure. For purposes of illustration, this disclosure describes a video decoder 300 in accordance with VVC (ITU-T H.266 under development) and HEVC (ITU-T H.265) techniques. However, the techniques of this disclosure may be implemented by video coding devices configured for other video coding standards.

[0198]

[0205] In the example of FIG. 3, the video decoder 300 includes a coded picture buffer (CPB) memory 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transform processing unit 308, a reconstruction unit 310, a filter unit 312, and a decoded picture buffer (DPB) 314. Any or all of the CPB memory 320, the entropy decoding unit 302, the prediction processing unit 304, the inverse quantization unit 306, the inverse transform processing unit 308, the reconstruction unit 310, the filter unit 312, and the DPB 314 may be implemented in one or more processors or processing circuits. For example, the units of the video decoder 300 may be implemented as one or more circuits or logic elements as part of a hardware circuit, or as part of a processor, ASIC, or FPGA. Moreover, the video decoder 300 may include additional or alternative processors or processing circuits that perform these and other functions.

[0199]

[0206] Prediction processing unit 304 includes a motion compensation unit 316 and an intra prediction unit 318. Prediction processing unit 304 may include additional units to perform prediction according to other prediction modes. By way of example, prediction processing unit 304 may include a palette unit, an intra block copy unit (which may form part of motion compensation unit 316), an affine unit, a linear model (LM) unit, etc. In other examples, video decoder 300 may include more, fewer, or different functional components.

[0200]

[0207] When operating according to AV1, the motion compensation unit 316 may be configured to decode coding blocks of the video data (e.g., both luma coding blocks and chroma coding blocks) using translational motion compensation, affine motion compensation, OBMC, and / or synthetic inter-intra prediction, as described above. The intra prediction unit 318 may be configured to decode coding blocks of the video data (e.g., both luma coding blocks and chroma coding blocks) using directional intra prediction, non-directional intra prediction, recursive filter intra prediction, CFL, intra block copy (IBC), and / or color palette mode, as described above.

[0201]

[0208] The CPB memory 320 may store video data, such as an encoded video bitstream, to be decoded by components of the video decoder 300. The video data stored in the CPB memory 320 may be obtained, for example, from the computer-readable medium 110 (FIG. 1). The CPB memory 320 may include a CPB that stores encoded video data (e.g., syntax elements) from the encoded video bitstream. The CPB memory 320 may also store video data other than syntax elements of coded pictures, such as temporary data representing output from various units of the video decoder 300. The DPB 314 generally stores decoded pictures that the video decoder 300 may output and / or use as reference video data when decoding subsequent data or pictures of the encoded video bitstream. The CPB memory 320 and the DPB 314 may be formed by any of a variety of memory devices, such as DRAM, including SDRAM, MRAM, RRAM, or other types of memory devices. The CPB memory 320 and the DPB 314 may be provided by the same memory device or separate memory devices. In various examples, the CPB memory 320 may be on-chip with other components of the video decoder 300, or may be off-chip relative to those components.

[0202]

[0209] Additionally or alternatively, in some examples, video decoder 300 may retrieve coded video data from memory 120 (FIG. 1). That is, memory 120 may store data as discussed above for CPB memory 320. Similarly, memory 120 may store instructions to be executed by video decoder 300 when some or all of the functionality of video decoder 300 is implemented in software to be executed by processing circuitry of video decoder 300.

[0203]

[0210] The various units shown in FIG. 3 are presented to aid in understanding the operations performed by the video decoder 300. The units may be implemented as fixed function circuits, programmable circuits, or a combination thereof. As with FIG. 2, fixed function circuits refer to circuits that provide a specific function and are predefined in the operations that may be performed. Programmable circuits refer to circuits that may be programmed to perform various tasks and provide flexible functionality in the operations that may be performed. For example, a programmable circuit may execute software or firmware that causes the programmable circuit to operate in a manner defined by the software or firmware instructions. Although a fixed function circuit may execute software instructions (e.g., receive a parameter or output a parameter), the type of operation that the fixed function circuit performs is generally immutable. In some examples, one or more of the units may be different circuit blocks (fixed function or programmable), and in some examples, one or more of the units may be an integrated circuit.

[0204]

[0211] The video decoder 300 may include a programmable core formed from ALUs, EFUs, digital circuits, analog circuits, and / or programmable circuits. In examples where the operations of the video decoder 300 are performed by software executing on programmable circuits, on-chip or off-chip memory may store instructions (e.g., object code) of the software that the video decoder 300 receives and executes.

[0205]

[0212] The entropy decoding unit 302 may receive the encoded video data from the CPB and entropy decode the video data to recover the syntax elements. The prediction processing unit 304, the inverse quantization unit 306, the inverse transform processing unit 308, the reconstruction unit 310, and the filter unit 312 may generate decoded video data based on the syntax elements extracted from the bitstream.

[0206]

[0213] In general, the video decoder 300 reconstructs a picture on a block-by-block basis. The video decoder 300 may perform a reconstruction operation on each block individually (the block currently being reconstructed, i.e., decoded, may be referred to as the “current block”).

[0207]

[0214] The entropy decoding unit 302 may entropy decode syntax elements that define the quantized transform coefficients of the quantized transform coefficient block as well as transform information, such as a quantization parameter (QP) and / or a transform mode indication(s). The inverse quantization unit 306 may use the QP associated with the quantized transform coefficient block to determine the degree of quantization, and likewise the degree of inverse quantization that the inverse quantization unit 306 should apply. The inverse quantization unit 306 may, for example, perform a bitwise left shift operation to inverse quantize the quantized transform coefficients. The inverse quantization unit 306 may thereby form a transform coefficient block including the transform coefficients.

[0208]

[0215] After the inverse quantization unit 306 forms the transform coefficient block, the inverse transform processing unit 308 may apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, the inverse transform processing unit 308 may apply an inverse DCT, an inverse integer transform, an inverse Karhunen-Loeve transform (KLT), an inverse rotational transform, an inverse transform, or another inverse transform to the transform coefficient block.

[0209]

[0216] Further, prediction processing unit 304 generates a prediction block according to the prediction information syntax element entropy decoded by entropy decoding unit 302. For example, if the prediction information syntax element indicates that the current block is inter predicted, motion compensation unit 316 may generate a prediction block. In this case, the prediction information syntax element may indicate a reference picture in DPB 314 from which to retrieve a reference block, as well as a motion vector that identifies the location of the reference block in the reference picture relative to the location of the current block in the current picture. Motion compensation unit 316 may generally perform an inter prediction process in a manner substantially similar to that described with respect to motion compensation unit 224 (FIG. 2).

[0210]

[0217] As another example, if the prediction information syntax element indicates that the current block is intra predicted, the intra prediction unit 318 may generate a prediction block according to the intra prediction mode indicated by the prediction information syntax element. Again, the intra prediction unit 318 may generally perform an intra prediction process in a manner substantially similar to that described with respect to the intra prediction unit 226 (FIG. 2). The intra prediction unit 318 may retrieve data of neighboring samples for the current block from the DPB 314.

[0211]

[0218] The reconstruction unit 310 may reconstruct the current block using the predictive block and the residual block. For example, the reconstruction unit 310 may add samples of the residual block to corresponding samples of the predictive block to reconstruct the current block.

[0212]

[0219] Filter unit 312 may perform one or more filter operations on the reconstructed blocks. For example, filter unit 312 may perform a deblocking operation to reduce blockiness artifacts along edges of the reconstructed blocks. The operations of filter unit 312 are not necessarily performed in all instances.

[0213]

[0220] The video decoder 300 may store the reconstructed blocks in the DPB 314. For example, in examples where the operations of the filter unit 312 are not performed, the reconstruction unit 310 may store the reconstructed blocks in the DPB 314. In examples where the operations of the filter unit 312 are performed, the filter unit 312 may store the filtered reconstructed blocks in the DPB 314. As described above, the DPB 314 may provide reference information to the prediction processing unit 304, such as samples of the current picture for intra prediction and previously decoded pictures for subsequent motion compensation. Additionally, the video decoder 300 may output the decoded pictures (e.g., decoded video) from the DPB 314 for later display on a display device, such as the display device 118 of FIG. 1.

[0214]

[0221] In this manner, the video decoder 300 represents an example of a video decoding device including a memory configured to store video data and one or more processing units implemented in a circuit and configured to perform the example techniques described in this disclosure. For example, in the case of entropy decoding, such as context-based arithmetic decoding, the entropy decoding unit 302 may determine a context value of one or more contexts. During the decoding of a slice or picture, the entropy decoding unit 302 may update the context value (e.g., a probability value). However, at the start of a slice or picture, the context value may be undefined. Rather than starting with an undefined context value, the entropy decoding unit 302 may use a predefined initialization point (e.g., an initial value stored in the DPB 314, the CPB memory 320, or some other memory) to initialize one or more context values.

[0215]

[0222] In one or more examples, rather than or in addition to using a predefined initialization point, the entropy decoding unit 302 may utilize one or more context values ​​of a previously decoded slice or picture, or a mapped, scaled, weighted version of one or more context values, etc., of the previously decoded slice or picture as the initialization point. For example, the entropy decoding unit 302 may store one or more context values ​​of a context of a decoded slice or picture (e.g., after decoding the last CTU of the slice or picture) as a set of initialization points in a buffer (e.g., the DPB 314, the CPB memory 320, or some other memory). The entropy decoding unit 302 may utilize the set of initialization points (e.g., the context values, or values ​​based on the context values ​​of a previously decoded slice or picture) to initialize the context values ​​of a context used to decode a subsequent slice or subsequent picture.

[0216]

[0223] The entropy decoding unit 302 may store sets of initialization points for multiple previously decoded slices or pictures. For example, a buffer may store a first set of initialization points associated with a first slice or first picture, a second set of initialization points associated with a second slice or second picture, etc. In addition, to determine (e.g., select) which set of initialization points the entropy decoding unit 302 may use for a subsequent slice or subsequent picture, the entropy decoding unit 302 may also store information of temporal identification values ​​and / or QP values ​​of the slices or pictures associated with each of the respective sets of initialization points.

[0217]

[0224] In one or more examples, when a buffer is full, the entropy decoding unit 302 may determine which set of initialization points to remove to make room for the just determined set of initialization points. For example, the entropy decoding unit 302 may determine one or more context values ​​of at least one context used to decode a current slice or a current picture. The entropy decoding unit 302 may determine that a buffer that stores a set of temporal initialization points from two or more slices or two or more pictures for context-based arithmetic coding is full. As described, each set of temporal initialization points is associated with one slice or one picture of the two or more slices or two or more pictures and includes one or more temporal initialization points.

[0218]

[0225] The entropy decoding unit 302 may determine a first set of temporal initialization points associated with a slice or picture from among the two or more slices or two or more pictures based on a temporal identification value or a quantization parameter (QP) value of the slice or picture. The entropy decoding unit 302 may remove the first set of temporal initialization points associated with the slice or picture from the buffer and store a second set of temporal initialization points associated with the current slice or current picture in the buffer. The second set of temporal initialization points is based on the determined one or more context values ​​(e.g., equal to one or more context values ​​determined for at least one context of the current slice or derived from one or more context values ​​of the picture or at least one context of the current slice or picture).

[0219]

[0226] Thus, instead of using when a slice or picture was decoded as the sole factor for determining (e.g., selecting) which set of initialization points should be removed from the buffer, the entropy decoding unit 302 may utilize a temporal identification value and / or a QP value to determine which set of initialization points to remove. For example, to determine a first set of temporal initialization points, the entropy decoding unit 302 may determine a first set of temporal initialization points associated with a slice or picture having at least one of a minimum temporal identification value or a quantization parameter (QP) value from among two or more slices or two or more pictures. For example, the entropy decoding unit 302 may determine a slice or picture having a minimum temporal identification value from among the temporal identification values ​​of the two or more slices or two or more pictures, or may determine a slice or picture having a minimum QP value from among the QP values ​​of the two or more slices or two or more pictures. The entropy decoding unit 302 may then remove the set of initialization points associated with the determined slice or picture (e.g., the one having the minimum temporal identification value or QP value) from the buffer.

[0220]

[0227] In one or more examples, the entropy decoding unit 302 may remove a set of initialization points associated with a slice or picture having a minimum temporal identification value or QP value if none of the sets of initialization points are associated with a slice or picture having a temporal identification value or QP value that is the same as the temporal identification value or QP value of the current slice or current picture. For example, the entropy decoding unit 302 may determine that at least one of the temporal identification values ​​or QP values ​​of the current slice or current picture is different from the temporal identification value or QP value of each of the two or more slices or two or more pictures. In this example, the entropy decoding unit 302 may remove a first set of temporal initialization points based on a determination that at least one of the temporal identification values ​​or QP values ​​of the current slice or current picture is different from the temporal identification value or QP value of each of the two or more slices or two or more pictures. For example, the temporal identification value or QP value of the slice or picture associated with the first set of temporal initialization points is different from the temporal identification value or QP value of the current slice or current picture.

[0221]

[0228] As an example, assume that the current slice or picture is a first slice or picture. In this example, the entropy decoding unit 302 may determine one or more context values ​​of at least one context used to decode the second slice or picture, and determine a third slice or picture from the two or more slices or pictures that has at least one of the temporal identification value or QP value that is the same as the temporal identification value or QP value of the second slice or picture. In this example, the entropy decoding unit 302 may remove a third set of temporal initialization points associated with the third slice or picture from the buffer, and store in the buffer a fourth set of temporal initialization points associated with the second slice or picture that is based on the determined one or more context values ​​of the at least one context used to decode the second slice or picture.

[0222]

[0229] Then, to decode a subsequent picture, the entropy decoding unit 302 may determine (e.g., select) a set of temporal initialization points stored in the buffer and initialize one or more context values ​​of at least one context used to decode the subsequent slice or subsequent picture based on the selected set of temporal initialization points. For example, the entropy decoding unit 302 may assign initial values ​​for the one or more context values ​​equal to the selected set of temporal initialization points or derive initial values ​​for the one or more context values ​​based on the selected set of temporal initialization points (e.g., using mapping, scaling, weighting, etc.). The entropy decoding unit 302 may context-based arithmetic decode the subsequent slice or subsequent picture. For example, the entropy decoding unit 302 may utilize the initial values ​​for the context values ​​to decode syntax elements of the subsequent slice or subsequent picture and update the context values ​​during decoding of the subsequent slice or subsequent picture.

[0223]

[0230] In one or more examples, the video decoder 300 may also be configured to store a plurality of temporal initialization points of one or more contexts used for context-based arithmetic coding of the video data of the current picture or the current slice, the context being included within the video data of one or more previous pictures preceding the current picture in coding order, store a respective temporal identification (ID) value associated with each of the plurality of temporal initialization points, select at least one temporal initialization point of the plurality of temporal initialization points based on the respective temporal ID value associated with the plurality of temporal initialization points and the temporal ID value of the current picture or the current slice, and context-based arithmetic decode the video data of the current picture or the current slice based on the selected at least one temporal initialization point.

[0224]

[0231] The video decoder 300 may be configured to store in a first buffer one or more temporal initialization points of one or more contexts used for context-based arithmetic coding of video data of one or more previous pictures and store in a second buffer one or more temporal initialization points of one or more contexts used for context-based arithmetic coding of video data of a slice of a current picture, where storing in the second buffer includes storing in the second buffer during coding of the video data of the current picture and storing the temporal initialization points stored in the second buffer in the first buffer following processing a last coding tree unit (CTU) or slice of the current picture.

[0225]

[0232] The video decoder 300 may be configured to store previous temporal initialization points of one or more contexts used for context-based arithmetic coding of video data of a previous slice in a previous picture, determine for a current slice that a current slice in the current picture has a location in the current picture that corresponds to a location of a previous slice in the previous picture based on the current slice having a location in the current picture that corresponds to a location of the previous slice in the previous picture, and determine a current temporal initialization point for the current slice based on the previous temporal initialization point.

[0226]

[0233] 4 is a flowchart illustrating an example method for encoding a current block according to the techniques of this disclosure. The current block may include a current CU. Although described with respect to video encoder 200 (FIGS. 1 and 2), it should be understood that other devices may be configured to implement a method similar to that of FIG.

[0227]

[0234] In this example, video encoder 200 first predicts the current block (400). For example, video encoder 200 may form a predictive block for the current block. Video encoder 200 may then calculate a residual block for the current block (402). To calculate the residual block, video encoder 200 may calculate a difference between an original uncoded block and a predictive block for the current block. Video encoder 200 may then transform the residual block and quantize transform coefficients of the residual block (404). Video encoder 200 may then scan the quantized transform coefficients of the residual block (406). During or following the scan, video encoder 200 may entropy code the transform coefficients (408). For example, video encoder 200 may code the transform coefficients using CAVLC or CABAC. According to one or more examples, video encoder 200 may code the transform coefficients using context values ​​determined using techniques described in this disclosure. The video encoder 200 may then output the entropy coded data for the block (410).

[0228]

[0235] 5 is a flowchart illustrating an example method for decoding a current block of video data in accordance with the techniques of this disclosure. The current block may include a current CU. Although described with respect to video decoder 300 (FIGS. 1 and 3), it should be understood that other devices may be configured to implement a method similar to that of FIG.

[0229]

[0236] The video decoder 300 may receive entropy coded data for a current block, such as entropy coded prediction information and entropy coded data for transform coefficients of a residual block corresponding to the current block (500). The video decoder 300 may entropy decode the entropy coded data to determine prediction information for the current block and reconstruct transform coefficients of the residual block (502). According to one or more examples, the video decoder 300 may decode the coded data using context values ​​determined using techniques described in this disclosure. The video decoder 300 may predict the current block, e.g., using an intra prediction mode or an inter prediction mode indicated by the prediction information for the current block, to compute a predictive block for the current block (504). The video decoder 300 may then inverse scan the reconstructed transform coefficients to generate a block of quantized transform coefficients (506). The video decoder 300 may then inverse quantize the transform coefficients and apply an inverse transform to the transform coefficients to generate the residual block (508). Video decoder 300 may finally decode the current block by combining the predictive block and the residual block (510).

[0230]

[0237] 6 is a flow chart illustrating an example method of processing video data. For ease of explanation, the example of FIG. 6 is described with respect to processing circuitry, examples of which include the processing circuitry of video encoder 200 and video decoder 300, and buffers, examples of which include memory 106, memory 120, video data memory 230, DPB 218, CPB memory 320, DPB 314, or any other memory of video encoder 200 or video decoder 300.

[0231]

[0238] The processing circuit may be configured to determine one or more context values ​​of at least one context used to encode or decode a current slice or a current picture (600). For example, when the video encoder 200 and the video decoder 300 are encoding or decoding a current slice or a current picture, the video encoder 200 and the video decoder 300 may be updating the context value relative to a context base on recently encoded or decoded video data of the slice or picture.

[0232]

[0239] The processing circuit may determine that a buffer storing a set of temporal initialization points from two or more slices or two or more pictures for context-based arithmetic coding is full (602). As described, each set of temporal initialization points is associated with one slice or one picture of the two or more slices or two or more pictures and includes one or more temporal initialization points. For example, to keep the size of the buffer practical, there may be a limit on the number of sets of temporal initialization points that the buffer can store. As an example, the buffer may store up to five sets of temporal initialization points.

[0233]

[0240] The processing circuit may determine (e.g., identify) a first set of temporal initialization points associated with a slice or picture from among the two or more slices or two or more pictures based on at least one of a slice type, a temporal identification value, or a quantization parameter (QP) value of the slice or picture (604). For example, the buffer may also store temporal identification values, QP values, and / or slice type information for slices and pictures associated with respective sets of initialization points stored in the buffer.

[0234]

[0241] As an example, to determine the first set of temporal initialization points, the processing circuit may be configured to determine, from among two or more slices or two or more pictures, a first set of temporal initialization points associated with a slice or picture having at least one of a minimum temporal identification value or a quantization parameter (QP) value. For example, the processing circuit may determine a slice or picture having a minimum temporal identification value from among the temporal identification values ​​of the two or more slices or two or more pictures, and / or determine a slice or picture having a minimum QP value from among the QP values ​​of the two or more slices or two or more pictures.

[0235]

[0242] In some examples, the first set of initialization points with the smallest temporal identification value or QP value may have a different slice type than the current slice. In some examples, such as when multiple sets of temporal initialization points may be stored for a slice type, the video encoder 200 and the video decoder 300 may determine a group of sets of temporal initialization points associated with slices having the same slice type as the current slice. The video encoder 200 and the video decoder 300 may then determine the set of temporal initialization points in this group with the smallest temporal identification value or QP value as the first set of temporal initialization points.

[0236]

[0243] The processing circuit may remove from the buffer a first set of temporal initialization points associated with the slice or picture (606). The processing circuit may store in the buffer a second set of temporal initialization points associated with the current slice or current picture (608). The second set of temporal initialization points is based on the determined one or more context values ​​(e.g., equal to or derived from one or more context values ​​determined for the context of the current slice or current picture). In some examples, following processing the last coding tree unit (CTU) of the current slice or current picture, the processing circuit may store the second set of temporal initialization points.

[0237]

[0244] In one or more examples, for a subsequent slice or subsequent picture in the coding order, the processing circuit may determine (e.g., select) a set of temporal initialization points stored in the buffer. The processing circuit may initialize one or more context values ​​of at least one context used to encode or decode the subsequent slice or subsequent picture based on the selected set of temporal initialization points (e.g., set initial values ​​for the context values ​​equal to the selected set of temporal initialization points or derive initial values ​​for the context values ​​based on the selected set of temporal initialization points). The processing circuit may context-based arithmetic encode or decode the subsequent slice or subsequent picture.

[0238]

[0245] 7 is a flow chart illustrating another exemplary method of processing video data. For ease of explanation, the example of FIG. 7 is described with respect to processing circuitry, examples of which include processing circuitry of video encoder 200 and video decoder 300, and buffers, examples of which include memory 106, memory 120, video data memory 230, DPB 218, CPB memory 320, DPB 314, or any other memory of video encoder 200 or video decoder 300.

[0239]

[0246] As described above, the processing circuit may remove the set of temporal initialization points associated with the slice or picture having the smallest temporal identification value or QP value. However, in some examples, the processing circuit may perform such a removal process if there is no set of temporal initialization points associated with a slice or picture having the same temporal identification value, QP value, or slice type as the current slice or current picture.

[0240]

[0247] For example, assume that the current slice or current picture in Figure 6 is a first slice or first picture. In this example, the processing circuit may determine one or more context values ​​of at least one context used to encode or decode a second slice or second picture (700). That is, the processing circuit may perform similar encoding or decoding operations on the second slice or second picture and update the context values, as described above.

[0241]

[0248] The processing circuit may determine a third slice or third picture from the two or more slices or pictures that has at least one of a temporal identification value, QP value, or slice type that is the same as the temporal identification value, QP value, or slice type of the second slice or second picture (702). In this example, rather than removing the set of initialization points associated with the slice or picture with the smallest temporal identification value or QP value, the processing circuit may remove the third set of temporal initialization points associated with the third slice or third picture from the buffer (704). The processing circuit may store in the buffer a fourth set of temporal initialization points associated with the second slice or second picture that is based on the determined one or more context values ​​of at least one context used to encode or decode the second slice or second picture (706).

[0242]

[0249] The example of FIG. 7 is provided for illustrative purposes only and should not be considered limiting. In some examples, the processing circuit may not implement the method of FIG. 7. Rather, the processing circuit may remove a set of initialization points based on a temporal identification value or QP value (e.g., a set of temporal initialization points associated with a slice or picture having a minimum temporal identification value or QP value). Also, slice type may not be considered as one of the factors. That is, the processing circuit may remove entries of a set of temporal initialization points having the same slice type and / or the same temporal identification value and / or the same QP value as the current slice or current picture.

[0243]

[0250] 8 is a flow chart illustrating another exemplary method of processing video data. For ease of explanation, the example of FIG. 8 is described with respect to processing circuitry, examples of which include the processing circuitry of video encoder 200 and video decoder 300, and buffers, examples of which include memory 106, memory 120, video data memory 230, DPB 218, CPB memory 320, DPB 314, or any other memory for video encoder 200 or video decoder 300.

[0244]

[0251] 6 and 7, a processing circuit may determine one or more context values ​​of at least one context used to encode / decode a current slice or a current picture (800). In one or more examples, the processing circuit may determine whether a buffer already stores a set of temporal initialization points for a slice or picture that has the same temporal identification value or QP value as the current slice or the current picture (802).

[0245]

[0252] If the buffer already stores a set of temporal initialization points for a slice or picture having the same temporal identification value or QP value as the current slice or current picture (yes in 802), the processing circuit may overwrite the stored set of temporal initialization points with the temporal initialization points (e.g., the determined context value or values, or derived from the context value or values) of the current slice or current picture (804). The processing circuit may then set the next slice or next picture as the current slice or current picture and return to determining one or more context values ​​of the at least one context used to encode and decode the current slice or current picture (800).

[0246]

[0253] If the buffer does not already store a set of temporal initialization points for a slice or picture having the same temporal identification value or QP value as the current slice or current picture (no in 802), the processing circuit may determine whether the buffer is full (806). If the buffer is not full (no in 806), the processing circuit may store a temporal initialization point (e.g., a context value for the current slice or current picture, or a value derived from a context value for the current slice or current picture) in the buffer (808). The processing circuit may then set the next slice or next picture as the current slice or current picture and return to determining one or more context values ​​of at least one context used to encode and decode the current slice or current picture (800).

[0247]

[0254] If the buffer is full (yes at 806), the processing circuit may determine (e.g., identify) from among the two or more slices or two or more pictures a first set of temporal initialization points associated with a slice or picture having at least one of a minimum temporal identification value or quantization parameter (QP) value (810). Similar to FIG. 6, the processing circuit may remove the first set of temporal initialization points associated with the slice or picture from the buffer (812) and store in the buffer a second set of temporal initialization points associated with the current slice or current picture, the second set of temporal initialization points being based on the determined one or more context values ​​(814).

[0248]

[0255] The exemplary order of operations shown in FIG. 8 and performed by the processing circuit should not be considered limiting. For example, the processing circuit may first determine whether the buffer is full (806), and if the buffer is not full, the processing circuit may store initialization points (808) even if the buffer already stores temporal initialization points for a slice or picture with the same temporal identification value or QP value as the current slice or current picture. As another example, if the buffer is full, the processing circuit may remove the set of temporal initialization points associated with the slice or picture with the smallest temporal identification value or QP value, even if the buffer already stores temporal initialization points for a slice or picture with the same temporal identification value or QP value as the current slice or current picture. Other modifications to the order of operations are possible.

[0249]

[0256] The techniques indicated by reference numbers 804 and 812 include examples of removing a set of initialization points. In general, the processing circuitry may first determine whether the buffer stores a temporal initialization point for the same temporal ID or QP value as the current slice or current picture, and if so, remove that set of temporal initialization points and write the temporal initialization points of the current slice or current picture. If not, the processing circuitry may remove the determined (e.g., identified) set of initialization points associated with the slice or picture having the smallest temporal identification value or QP value.

[0250]

[0257] The following describes several example techniques, which can be implemented together or separately.

[0251]

[0258] Clause 1. A method for coding video data, comprising: storing a plurality of temporal initialization points of one or more contexts used for context-based arithmetic coding of video data of a current picture or a current slice, the context initialization points being included within video data of one or more previous pictures preceding a current picture in coding order; storing a respective temporal identification (ID) value associated with each of the plurality of temporal initialization points; selecting at least one (e.g., if available) of the plurality of temporal initialization points based on the respective temporal ID values ​​associated with the plurality of temporal initialization points and the temporal ID value of the current picture or current slice; and context-based arithmetic coding of the video data of the current picture or current slice based on the selected at least one temporal initialization point.

[0252]

[0259] Clause 2. The method of clause 1, wherein selecting at least one temporal initialization point includes determining a set of temporal initialization points from a plurality of temporal initialization points, wherein each temporal initialization point in the set of temporal initialization points has a temporal ID value less than or equal to a temporal ID value of a current picture or a current slice, and selecting at least one temporal initialization point from the set of temporal initialization points.

[0253]

[0260] Clause 3. The method of clause 1, wherein selecting at least one temporal initialization point includes determining a set of temporal initialization points from a plurality of temporal initialization points, wherein a temporal ID value of each of the temporal initialization points in the set of temporal initialization points is equal to a temporal ID value of a current picture or a current slice, and selecting at least one temporal initialization point from the set of temporal initialization points.

[0254]

[0261] Clause 4. The method of clause 1, wherein selecting at least one temporal initialization point includes: determining that none of the plurality of temporal initialization points has a temporal ID value equal to the temporal ID value of the current picture or current slice; determining whether any of the plurality of temporal initialization points has a temporal ID value that is one less than the temporal ID value of the current picture or current slice based on a determination that none of the plurality of temporal initialization points has a temporal ID value equal to the temporal ID value of the current picture or current slice; and selecting at least one temporal initialization point from the one or more temporal initialization points having a temporal ID value that is one less than the temporal ID value of the current picture or current slice based on a determination that there are one or more temporal initialization points having a temporal ID value that is one less than the temporal ID value of the current picture or current slice.

[0255]

[0262] Clause 5. A method for coding video data, comprising: storing in a first buffer one or more temporal initialization points of one or more contexts used for context-based arithmetic coding of video data of one or more previous pictures; storing in a second buffer temporal initialization points of one or more contexts used for context-based arithmetic coding of video data of a slice of a current picture, wherein storing in the second buffer includes storing in the second buffer during coding of the video data of the current picture; and storing in the first buffer the temporal initialization points stored in the second buffer following processing a last coding tree unit (CTU) or slice of the current picture.

[0256]

[0263] Clause 6. The method of any one of clauses 1 to 4, wherein storing the plurality of time initialization points comprises storing the plurality of time initialization points according to the method of clause 5.

[0257]

[0264] Clause 7. A method for coding video data, comprising: storing previous temporal initialization points of one or more contexts used for context-based arithmetic coding of video data of a previous slice in a previous picture; determining, for a current slice, that the current slice in the current picture has a location in the current picture that corresponds to a location of the previous slice in the previous picture; and determining a current temporal initialization point for the current slice based on the previous temporal initialization point based on the current slice having a location in the current picture that corresponds to a location of the previous slice in the previous picture.

[0258]

[0265] Clause 8. A method for coding video data comprising any combination of clauses 1 to 7.

[0259]

[0266] Clause 9. The method of any one of clauses 1 to 8, wherein the context-based arithmetic coding comprises context-adaptive binary arithmetic coding (CABAC).

[0260]

[0267] Clause 10. The method of any one of clauses 1 to 9, wherein the context-based arithmetic coding comprises context-based arithmetic decoding.

[0261]

[0268] Clause 11. The method of any one of clauses 1 to 9, wherein the context-based arithmetic coding comprises context-based arithmetic coding.

[0262]

[0269] Clause 12. A device for coding video data, comprising a memory configured to store the video data and a processing circuit configured to carry out a method according to any one of clauses 1 to 11 or a combination thereof.

[0263]

[0270] Clause 13. The device of clause 12, wherein the device includes a video decoder.

[0264]

[0271] Clause 14. A device according to clause 12 or 13, wherein the device comprises a video encoder.

[0265]

[0272] Clause 15. The device of any one of clauses 12 to 14, further comprising a display configured to display the decoded video data.

[0266]

[0273] Clause 16. A device according to any one of clauses 12 to 15, wherein the device comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

[0267]

[0274] Clause 17. A computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to perform the method of any one of clauses 1 to 11.

[0268]

[0275] Clause 18. A device for coding video data, comprising means for implementing the method according to any one of clauses 1 to 11 or a combination thereof.

[0269]

[0276] Clause 1A. A method for processing video data, comprising: determining one or more context values ​​of at least one context used to encode or decode a current slice or a current picture; determining that a buffer that stores sets of temporal initialization points from two or more slices or two or more pictures for context-based arithmetic coding is full, each set of temporal initialization points being associated with a slice or a picture among the two or more slices or two or more pictures, the set including the one or more temporal initialization points; determining a first set of temporal initialization points associated with a slice or picture from among the two or more slices or two or more pictures based on at least one of a slice type, a temporal identification value, or a quantization parameter (QP) value of the slice or picture; removing the first set of temporal initialization points associated with the slice or picture from the buffer; and storing in the buffer a second set of temporal initialization points associated with the current slice or current picture, the second set of temporal initialization points being based on the determined one or more context values.

[0270]

[0277] Clause 2A. The method of clause 1A, wherein determining a first set of temporal initialization points includes determining from among two or more slices or two or more pictures a first set of temporal initialization points associated with a slice or picture having at least one of the smallest temporal identification values ​​or quantization parameter (QP) values ​​from among the temporal identification values ​​or QP values ​​of the two or more slices or two or more pictures.

[0271]

[0278] Clause 3A. The method of clause 2A, wherein determining a slice or picture having at least one of a minimum temporal identification value or a minimum QP value from among two or more slices or two or more pictures includes determining a slice or picture having a minimum temporal identification value from among the temporal identification values ​​of the two or more slices or two or more pictures.

[0272]

[0279] Clause 4A. The method of clause 2A, wherein determining a slice or picture having at least one of a minimum temporal identification value or a minimum QP value from among two or more slices or two or more pictures includes determining a slice or picture having a minimum QP value from among the QP values ​​of the two or more slices or two or more pictures.

[0273]

[0280] Clause 5A. The method of any one of clauses 1A to 4A, further comprising determining that at least one of the temporal identification values ​​or QP values ​​of the current slice or current picture is different from the temporal identification values ​​or QP values ​​of each of two or more slices or two or more pictures, wherein removing the first set of temporal initialization points comprises removing the first set of temporal initialization points based on a determination that at least one of the temporal identification values ​​or QP values ​​of the current slice or current picture is different from the temporal identification values ​​or QP values ​​of each of two or more slices or two or more pictures.

[0274]

[0281] Clause 6A. The method of any one of clauses 1A-5A, wherein the current slice or current picture is a first slice or first picture, and the method includes determining one or more context values ​​of at least one context used to encode or decode a second slice or second picture; determining a third slice or third picture from the two or more slices or pictures, the third slice or third picture having at least one of a temporal identification value, a QP value, or a slice type that is the same as a temporal identification value, a QP value, or a slice type of the second slice or second picture; removing a third set of temporal initialization points associated with the third slice or third picture from the buffer; and storing in the buffer a fourth set of temporal initialization points associated with the second slice or second picture that are based on the determined one or more context values ​​of the at least one context used to encode or decode the second slice or second picture.

[0275]

[0282] Clause 7A. A method according to any one of clauses 1A to 6A, further comprising: determining a set of temporal initialization points stored in a buffer; initializing one or more context values ​​of at least one context used to encode or decode a subsequent slice or subsequent picture based on the determined set of temporal initialization points; and context-based arithmetic encoding or decoding the subsequent slice or subsequent picture.

[0276]

[0283] Clause 8A. The method of clause 7A, further comprising: determining a temporal identification value of a subsequent slice or subsequent picture; determining that none of two or more slices or two or more pictures having an associated set of temporal initialization points stored in the buffer have a temporal identification value equal to the temporal identification value of the subsequent slice or subsequent picture; and determining from among the two or more slices or two or more pictures a second slice or second picture having a temporal identification value that is closest to and less than the temporal identification value of the subsequent slice or subsequent picture, wherein determining the set of temporal initialization points includes selecting the set of temporal initialization points associated with the second slice or second picture.

[0277]

[0284] Clause 9A. The method of clause 7A, further comprising: determining a QP value of a subsequent slice or subsequent picture; determining that none of two or more slices or two or more pictures having an associated set of temporal initialization points stored in the buffer have a QP value equal to the QP value of the subsequent slice or subsequent picture; and determining from among the two or more slices or two or more pictures, a second slice or second picture having a QP value closest to the QP value of the subsequent slice or subsequent picture, wherein determining the set of temporal initialization points includes selecting the set of initialization points associated with the second slice or second picture.

[0278]

[0285] Clause 10A. A method according to any one of clauses 1A to 9A, wherein storing the second set of initialization points includes storing the second set of temporal initialization points following processing the last coding tree unit (CTU) of the current slice or current picture.

[0279]

[0286] Clause 11A. A device for processing video data, comprising: a buffer configured to store a set of temporal initialization points from two or more slices or two or more pictures for context-based arithmetic coding; and a processing circuit coupled to the buffer, wherein the processing circuit determines one or more context values ​​of at least one context used to encode or decode a current slice or a current picture, and stores the set of temporal initialization points from two or more slices or two or more pictures for context-based arithmetic coding, each set of temporal initialization points being associated with one slice or one picture of the two or more slices or two or more pictures, and wherein one or more context values ​​of at least one context used to encode or decode a current slice or a current picture are stored in the buffer. 11. A device configured to: determine that a buffer that stores a set of temporal initialization points is full, the set including a plurality of temporal initialization points; determine a first set of temporal initialization points associated with a slice or picture from among two or more slices or two or more pictures based on at least one of a slice type, a temporal identification value, or a quantization parameter (QP) value of the slice or picture; remove the first set of temporal initialization points associated with the slice or picture from the buffer; and store in the buffer a second set of temporal initialization points associated with a current slice or current picture, the second set of temporal initialization points being based on the determined one or more context values.

[0280]

[0287] Clause 12A. The device of clause 11A, wherein to determine the first set of temporal initialization points, the processing circuitry is configured to determine from among two or more slices or two or more pictures a first set of temporal initialization points associated with a slice or picture having at least one of the smallest temporal identification values ​​or quantization parameter (QP) values ​​among the temporal identification values ​​or QP values ​​of the two or more slices or two or more pictures.

[0281]

[0288] Clause 13A. The device described in Clause 12A, wherein the processing circuit is configured to determine a slice or picture having at least one of the minimum temporal identification value or QP value from among two or more slices or two or more pictures, the slice or picture having the minimum temporal identification value from among the temporal identification values ​​of the two or more slices or two or more pictures.

[0282]

[0289] Clause 14A. The device of clause 12A, wherein the processing circuitry is configured to determine a slice or picture having at least one of a minimum temporal identification value or QP value from among two or more slices or two or more pictures, the slice or picture having a minimum QP value from the QP values ​​of the two or more slices or two or more pictures.

[0283]

[0290] Clause 15A. A device as described in any one of clauses 11A to 14A, wherein the processing circuit is configured to determine that at least one of the temporal identification values ​​or QP values ​​of the current slice or current picture is different from the temporal identification values ​​or QP values ​​of each of two or more slices or two or more pictures, and the processing circuit is configured to remove the first set of temporal initialization points based on a determination that at least one of the temporal identification values ​​or QP values ​​of the current slice or current picture is different from the temporal identification values ​​or QP values ​​of each of two or more slices or two or more pictures, to remove the first set of temporal initialization points.

[0284]

[0291] Clause 16A. The device of any one of clauses 11A to 15A, wherein the current slice or current picture is a first slice or first picture, and the processing circuit is configured to determine one or more context values ​​of at least one context used to encode or decode the second slice or second picture, determine a third slice or third picture from the two or more slices or pictures having at least one of a temporal identification value, a QP value, or a slice type that is the same as the temporal identification value, a QP value, or a slice type of the second slice or second picture, remove a third set of temporal initialization points associated with the third slice or third picture from the buffer, and store in the buffer a fourth set of temporal initialization points associated with the second slice or second picture that are based on the determined one or more context values ​​of the at least one context used to encode or decode the second slice or second picture.

[0285]

[0292] Clause 17A. A device described in any one of clauses 11A to 16A, wherein the processing circuitry is configured to determine a set of temporal initialization points stored in the buffer, initialize one or more context values ​​of at least one context used to encode or decode a subsequent slice or subsequent picture based on the determined set of temporal initialization points, and perform context-based arithmetic encoding or decoding of the subsequent slice or subsequent picture.

[0286]

[0293] Clause 18A. The device described in Clause 17A, wherein the processing circuit is configured to determine a temporal identification value of a subsequent slice or subsequent picture, determine that none of two or more slices or two or more pictures having an associated set of temporal initialization points stored in the buffer have a temporal identification value equal to the temporal identification value of the subsequent slice or subsequent picture, and determine from among the two or more slices or two or more pictures a second slice or second picture having a temporal identification value that is closest to and less than the temporal identification value of the subsequent slice or subsequent picture, and wherein the processing circuit is configured to select the set of temporal initialization points associated with the second slice or second picture to determine the set of temporal initialization points.

[0287]

[0294] Clause 19A. The device of clause 17A, wherein the processing circuit is configured to determine a QP value of a subsequent slice or subsequent picture, determine that none of two or more slices or two or more pictures having an associated set of temporal initialization points stored in the buffer have a QP value equal to the QP value of the subsequent slice or subsequent picture, and determine from among the two or more slices or two or more pictures a second slice or second picture having a QP value closest to the QP value of the subsequent slice or subsequent picture, and wherein the processing circuit is configured to select the set of initialization points associated with the second slice or second picture to determine the set of temporal initialization points.

[0288]

[0295] Clause 20A. A device described in any one of clauses 11A to 19A, wherein the processing circuitry is configured to store the second set of temporal initialization points following processing the last coding tree unit (CTU) of the current slice or current picture to store the second set of initialization points.

[0289]

[0296] Clause 21A. The device of any one of clauses 11A to 20A, wherein the device includes one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

[0290]

[0297] Clause 22A. A computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to determine one or more context values ​​of at least one context used to encode or decode a current slice or a current picture; determine that a buffer storing a set of temporal initialization points from two or more slices or two or more pictures for context-based arithmetic coding is full, where each set of temporal initialization points is associated with a slice or a picture among the two or more slices or two or more pictures, the set of temporal initialization points including the one or more temporal initialization points; determine a first set of temporal initialization points associated with a slice or picture from among the two or more slices or two or more pictures based on at least one of a slice type, a temporal identification value, or a quantization parameter (QP) value of the slice or picture; remove the first set of temporal initialization points associated with the slice or picture from the buffer; and store in the buffer a second set of temporal initialization points associated with the current slice or the current picture, the second set of temporal initialization points being based on the determined one or more context values.

[0291]

[0298] Clause 23A. A computer-readable storage medium according to clause 22A, further comprising instructions for causing one or more processors to perform a method according to any one of clauses 1A to 10A.

[0292]

[0299] Clause 24A. A device for processing video data, comprising: means for determining one or more context values ​​of at least one context used to encode or decode a current slice or a current picture; means for determining that a buffer storing sets of temporal initialization points from two or more slices or two or more pictures for context-based arithmetic coding is full, each set of temporal initialization points being associated with a slice or a picture among the two or more slices or two or more pictures, the set including the one or more temporal initialization points; means for determining a first set of temporal initialization points associated with a slice or picture from among the two or more slices or two or more pictures based on at least one of a slice type, a temporal identification value, or a quantization parameter (QP) value of the slice or picture; means for removing the first set of temporal initialization points associated with the slice or picture from the buffer; and means for storing in the buffer a second set of temporal initialization points associated with the current slice or the current picture, the second set of temporal initialization points being based on the determined one or more context values.

[0293]

[0300] Clause 25A. A device as described in clause 24A, further comprising instructions to cause one or more processors to perform a method as described in any one of clauses 1A to 10A.

[0294]

[0301] It should be appreciated that in some examples, some acts or events of any of the techniques described herein may be performed in a different order, or may be added, merged, or omitted entirely (e.g., not all acts or events described may be required to practice the techniques). Moreover, in some examples, acts or events may be performed in parallel rather than sequentially, for example, through multithreading, interrupt processing, or multiple processors.

[0295]

[0302] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium, such as a data storage medium, or a communication medium including any medium that facilitates transfer of a computer program from one place to another, for example according to a communication protocol. As such, a computer-readable medium may generally correspond to (1) a tangible computer-readable storage medium that is non-transitory, or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures to implement the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0296]

[0303] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included within the definition of media. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory media, but instead cover non-transitory tangible storage media. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc, where disks typically reproduce data magnetically, while discs reproduce data optically using lasers. Combinations of the above are also intended to be included within the scope of computer readable media.

[0297]

[0304] The instructions may be executed by one or more processors, such as one or more DSPs, general-purpose microprocessors, ASICs, FPGAs, or other equivalent integrated or discrete logic circuits. Thus, the terms "processor" and "processing circuitry" as used herein may refer to any of the above structures, or any other structures suitable for implementing the techniques described herein. In addition, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or may be incorporated into a combined codec. The techniques may also be fully implemented in one or more circuits or logic elements.

[0298]

[0305] The techniques of the present disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC), or a set of ICs (e.g., a chipset). Various components, modules, or units have been described in this disclosure to highlight functional aspects of devices configured to implement the disclosed techniques, but they do not necessarily require realization by different hardware units. Rather, as described above, the various units may be combined in a codec hardware unit or may be provided by a collection of interoperable hardware units, including one or more processors as described above, in conjunction with suitable software and / or firmware.

[0299]

[0306] Various examples have been described. These and other examples are within the scope of the following claims.

Claims

1. 1. A method for processing video data, comprising: determining one or more context values ​​of at least one context used to encode or decode a current slice or a current picture; determining that a buffer storing sets of temporal initialization points from two or more slices or two or more pictures for context-based arithmetic coding, each set of temporal initialization points being associated with one slice or one picture among the two or more slices or two or more pictures, the sets including one or more temporal initialization points, is full; determining a first set of temporal initialization points associated with the slices or pictures from among the two or more slices or the two or more pictures based on at least one of a slice type, a temporal identification value, or a quantization parameter (QP) value of the slices or pictures; removing the first set of temporal initialization points associated with the slice or the picture from the buffer; storing in the buffer a second set of temporal initialization points associated with the current slice or the current picture, the second set of temporal initialization points being based on the determined one or more context values; A method comprising:

2. 2. The method of claim 1, wherein determining the first set of temporal initialization points comprises determining, from among the two or more slices or the two or more pictures, the first set of temporal initialization points associated with the slice or the picture having at least one of a minimum temporal identification value or a minimum quantization parameter (QP) value among the temporal identification values ​​or the minimum QP value of the two or more slices or the two or more pictures.

3. 3. The method of claim 2, wherein determining the slice or the picture having at least one of the smallest temporal identification value or the smallest QP value from among the two or more slices or the two or more pictures comprises determining the slice or the picture having the smallest temporal identification value from among the temporal identification values ​​of the two or more slices or the two or more pictures.

4. 3. The method of claim 2, wherein determining the slice or the picture having the smallest temporal identification value or the smallest QP value from among the two or more slices or the two or more pictures comprises determining the slice or the picture having the smallest QP value from among QP values ​​of the two or more slices or the two or more pictures.

5. 2. The method of claim 1, further comprising: determining that at least one of a temporal identification value or a QP value of the current slice or the current picture is different from a temporal identification value or a QP value of each of the two or more slices or the two or more pictures; the temporal identification value or the QP value of the slice or the picture associated with the first set of temporal initialization points is different from the temporal identification value or the QP value of the current slice or the current picture; removing the first set of temporal initialization points includes removing the first set of temporal initialization points based on the determination that at least one of the temporal identification value or the QP value of the current slice or the current picture is different from the temporal identification value or the QP value of each of the two or more slices or the two or more pictures. The method of claim 1.

6. The current slice or the current picture is a first slice or a first picture, and the method comprises: determining one or more context values ​​of at least one context used to encode or decode the second slice or the second picture; determining a third slice or a third picture from the two or more slices or pictures, the third slice or the third picture having at least one of a temporal identification value, a QP value, or a slice type that is the same as a temporal identification value, a QP value, or a slice type of the second slice or the second picture; removing a third set of temporal initialization points associated with the third slice or the third picture from the buffer; storing in the buffer a fourth set of temporal initialization points associated with the second slice or the second picture based on the determined one or more context values ​​of at least one context used to encode or decode the second slice or the second picture; The method of claim 1 further comprising:

7. determining a set of time initialization points stored in the buffer; initializing one or more context values ​​of at least one context used for encoding or decoding a subsequent slice or a subsequent picture based on the determined set of temporal initialization points; context-based arithmetic coding or decoding the subsequent slice or the subsequent picture; The method of claim 1 further comprising:

8. 1. A device for processing video data, comprising: a buffer configured to store a set of temporal initialization points from two or more slices or two or more pictures for context-based arithmetic coding; a processing circuit coupled to the buffer; wherein the processing circuitry comprises: determining one or more context values ​​of at least one context used to encode or decode the current slice or the current picture; determining that the buffer storing sets of temporal initialization points from two or more slices or two or more pictures for context-based arithmetic coding, each set of temporal initialization points being associated with one slice or one picture among the two or more slices or two or more pictures, the sets including one or more temporal initialization points, is full; determining a first set of temporal initialization points associated with the slices or pictures from among the two or more slices or the two or more pictures based on at least one of a slice type, a temporal identification value, or a quantization parameter (QP) value of the slices or pictures; removing the first set of temporal initialization points associated with the slice or the picture from the buffer; storing in the buffer a second set of temporal initialization points associated with the current slice or the current picture, the second set of temporal initialization points being based on the determined one or more context values. The device is configured as follows:

9. 9. The device of claim 8, wherein, to determine the first set of temporal initialization points, the processing circuitry is configured to determine, from among the two or more slices or the two or more pictures, the first set of temporal initialization points associated with the slice or the picture having at least one of the smallest temporal identification value or quantization parameter (QP) value among the temporal identification values ​​or QP values ​​of the two or more slices or the two or more pictures.

10. 10. The device of claim 9, wherein the processing circuit is configured to determine the slice or the picture having the smallest temporal identification value from among the two or more slices or the two or more pictures, the slice or the picture having the smallest temporal identification value from among the temporal identification values ​​of the two or more slices or the two or more pictures.

11. 10. The device of claim 9, wherein the processing circuitry is configured to determine the slice or the picture having the smallest QP value from QP values ​​of the two or more slices or the two or more pictures to determine the slice or the picture having the smallest temporal identification value or the smallest QP value from among the two or more slices or the two or more pictures.

12. the processing circuitry determining that at least one of the temporal identification value or QP value of the current slice or the current picture is different from the temporal identification value or QP value of each of the two or more slices or the two or more pictures; It is structured as follows: the temporal identification value or the QP value of the slice or the picture associated with the first set of temporal initialization points is different from the temporal identification value or the QP value of the current slice or the current picture; 9. The device of claim 8, wherein, to remove the first set of temporal initialization points, the processing circuitry is configured to remove the first set of temporal initialization points based on the determination that at least one of the temporal identification value or the QP value of the current slice or the current picture is different from the temporal identification value or the QP value of each of the two or more slices or the two or more pictures.

13. the current slice or the current picture is a first slice or a first picture, and the processing circuitry: determining one or more context values ​​of at least one context used to encode or decode the second slice or the second picture; determining a third slice or a third picture from the two or more slices or pictures, the third slice or the third picture having at least one of a temporal identification value, a QP value, or a slice type that is the same as a temporal identification value, a QP value, or a slice type of the second slice or the second picture; removing a third set of temporal initialization points associated with the third slice or the third picture from the buffer; 10. The device of claim 8, further configured to store in the buffer a fourth set of temporal initialization points associated with the second slice or the second picture that are based on the determined one or more context values ​​of at least one context used to encode or decode the second slice or the second picture.

14. A computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to: determining one or more context values ​​of at least one context used to encode or decode the current slice or the current picture; determining that a buffer storing sets of temporal initialization points from two or more slices or two or more pictures for context-based arithmetic coding, each set of temporal initialization points being associated with one slice or one picture among the two or more slices or two or more pictures, the sets including one or more temporal initialization points, is full; determining a first set of temporal initialization points associated with the slices or pictures from among the two or more slices or the two or more pictures based on at least one of a slice type, a temporal identification value, or a quantization parameter (QP) value of the slices or pictures; removing the first set of temporal initialization points associated with the slice or the picture from the buffer; storing in the buffer a second set of temporal initialization points associated with the current slice or the current picture, the second set of temporal initialization points being based on the determined one or more context values. A computer-readable storage medium.

15. 1. A device for processing video data, comprising: means for determining one or more context values ​​of at least one context used to encode or decode a current slice or a current picture; means for determining that a buffer storing sets of temporal initialization points from two or more slices or two or more pictures for context-based arithmetic coding, each set of temporal initialization points being associated with one slice or one picture among the two or more slices or two or more pictures, the sets including one or more temporal initialization points, is full; means for determining a first set of temporal initialization points associated with a slice or a picture from among the two or more slices or the two or more pictures based on at least one of a slice type, a temporal identification value, or a quantization parameter (QP) value of the slice or picture; means for removing the first set of temporal initialization points associated with the slice or the picture from the buffer; means for storing in the buffer a second set of temporal initialization points associated with the current slice or the current picture, the second set of temporal initialization points being based on the determined one or more context values; Including, the device.