Reference picture list and collocated picture signaling in video coding

By deriving co-position images from reference image list 0 through inference of temporal motion vector predictions during video coding, the inefficiency of signaling notification in VVC draft 8 is resolved, enabling a more efficient video decoding process.

CN115191117BActive Publication Date: 2026-04-10QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
QUALCOMM INC
Filing Date
2021-02-23
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In video coding, the reference picture list signaling notification in VVC draft 8 has inefficiencies and potential problems, especially since the co-position picture indicating temporal motion vector prediction in the P-strip is derived from reference picture list 1, which limits the parsing and decoding process of the video decoder.

Method used

The video decoder is configured to infer, without actually receiving syntax elements, that the co-bit image indicating that it will be derived from reference image list 0, covering the PH-level syntax of the P-strip, thus achieving efficient decoding.

Benefits of technology

By inferring syntax element values, the video decoder can parse and decode bitstreams without restricting the encoding method of the video encoder, thus improving the efficiency and flexibility of video decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115191117B_ABST
    Figure CN115191117B_ABST
Patent Text Reader

Abstract

The video decoder can be configured to, in response to receiving a first syntax element indicating that reference picture list information is included in the picture header syntax structure, receive, in the picture header syntax structure, a second syntax element indicating whether a collocated picture for temporal motion vector prediction is to be derived from the first reference picture list or the second reference picture list, receive a slice of video data referring to the picture header syntax structure, and in response to the slice being a P slice, set a value for a third syntax element associated with the slice to a first value for the third syntax element, the first value for the third syntax element indicating that the collocated picture for temporal motion vector prediction is to be derived from the first reference picture list.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-referencing

[0002] This application claims priority to U.S. Application No. 17 / 181,876, filed February 22, 2021, and U.S. Provisional Patent Application No. 62 / 980,695, filed February 24, 2020, the entire contents of each of which are incorporated herein by reference. U.S. Application No. 17 / 181,876, filed February 22, 2021, claims the benefit of U.S. Provisional Patent Application No. 62 / 980,695, filed February 24, 2020. Technical Field

[0003] This disclosure relates to video encoding and video decoding. Background Technology

[0004] Digital video functionality can be integrated into a wide variety of devices, including digital televisions, digital direct broadcasting systems, wireless broadcasting systems, personal digital assistants (PDAs), laptops or desktop computers, tablets, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radio phones, so-called "smartphones," video conferencing equipment, video streaming devices, and more. Digital video devices implement video decoding technologies, such as those described in standards like MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10, Advanced Video Decoding (AVC), ITU-T H.265 / High-Efficiency Video Decoding (HEVC), and extensions to these standards. By implementing such video decoding technologies, video devices can more efficiently transmit, receive, encode, decode, and / or store digital video information.

[0005] Video decoding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or eliminate inherent redundancy in video sequences. For block-based video decoding, video strips (e.g., video pictures or portions of video pictures) can be partitioned into video blocks, which may also be referred to as decoding tree units (CTUs), decoding units (CUs), and / or decoding nodes. Video blocks in intra-frame decoding (I) strips of a picture are encoded using spatial predictions of reference samples in adjacent blocks within the same picture. Video blocks in inter-frame decoding (P or B) strips of a picture can be encoded using spatial predictions relative to reference samples in adjacent blocks within the same picture, or temporal predictions against reference samples in other reference pictures. A picture may be referred to as a frame, and a reference picture may be referred to as a reference frame. Summary of the Invention

[0006] The techniques of this disclosure relate to inter prediction, and more specifically, to signaling of reference pictures used in inter prediction. As discussed in more detail below, the reference picture list signaling used in VVC Draft 8 can have some inefficiencies and other potential issues. For example, even though a PH can have an associated P slice that uses only reference picture list 0, the signaling technique of VVC Draft 8 allows PH level signaling that indicates that a collocated picture to be used for temporal motion vector prediction is to be derived from reference picture list 1. According to the techniques described by this disclosure, a video decoder can be configured to, in response to a slice being a P slice, and setting a value of a syntax element associated with the slice to a value of a syntax element that indicates that a collocated picture to be used for temporal motion vector prediction is to be derived from reference picture list 0. As described in more detail below, the video decoder can be configured to infer the value that indicates that the collocated picture to be used for temporal motion vector prediction is to be derived from reference picture list 0 without actually receiving an instance of the syntax element, effectively overriding the PH level syntax for the P slice. Thus, the techniques of this disclosure can advantageously enable a video decoder to parse and decode such bitstreams without unduly restricting the manner in which a video encoder encodes the bitstream.

[0007] According to one example of the disclosure, a method of decoding video data includes receiving a first syntax element, in response to the first syntax element indicating that reference picture list information is included in a picture header syntax structure, receiving a second syntax element in the picture header syntax structure, wherein a first value for the second syntax element indicates that a collocate picture to be used for temporal motion vector prediction is to be derived from a first reference picture list and a second value for the second syntax element indicates that the collocate picture to be used for temporal motion vector prediction is to be derived from a second reference picture list, receiving a slice of video data referring to the picture header syntax structure, and in response to the slice being a P slice, setting a value for a third syntax element associated with the slice to the first value of the third syntax element, wherein the first value of the third syntax element indicates that the collocate picture to be used for temporal motion vector prediction is to be derived from the first reference picture list and a second value for the third syntax element indicates that the collocate picture to be used for temporal motion vector prediction is to be derived from the second reference picture list.

[0008] According to another example of the disclosure, a device for decoding video data includes a memory configured to store video data, and one or more processors implemented in circuitry and configured to receive a first syntax element, in response to the first syntax element indicating that reference picture list information is included in a picture header syntax structure, receive a second syntax element in the picture header syntax structure, wherein a first value for the second syntax element indicates that a co-located picture used for temporal motion vector prediction is to be derived from a first reference picture list, and a second value for the second syntax element indicates that the co-located picture used for temporal motion vector prediction is to be derived from a second reference picture list, receive a slice of video data referring to the picture header syntax structure, and in response to the slice being a P slice, set a value for a third syntax element associated with the slice to the first value for the third syntax element, wherein the first value for the third syntax element indicates that the co-located picture used for temporal motion vector prediction is to be derived from the first reference picture list, and a second value for the third syntax element indicates that the co-located picture used for temporal motion vector prediction is to be derived from the second reference picture list.

[0009] According to another example of the disclosure, a computer-readable storage medium stores instructions that, when executed by one or more processors, cause the one or more processors to receive a first syntax element, in response to the first syntax element indicating that reference picture list information is included in a picture header syntax structure, receive a second syntax element in the picture header syntax structure, wherein a first value for the second syntax element indicates that a co-located picture used for temporal motion vector prediction is to be derived from a first reference picture list, and a second value for the second syntax element indicates that the co-located picture used for temporal motion vector prediction is to be derived from a second reference picture list, receive a slice of video data referring to the picture header syntax structure, and in response to the slice being a P slice, set a value for a third syntax element associated with the slice to the first value for the third syntax element, wherein the first value for the third syntax element indicates that the co-located picture used for temporal motion vector prediction is to be derived from the first reference picture list, and a second value for the third syntax element indicates that the co-located picture used for temporal motion vector prediction is to be derived from the second reference picture list.

[0010] According to another example of the disclosure, an apparatus for decoding video data includes means for receiving a first syntax element; in response to the first syntax element indicating that reference picture list information is included in a picture header syntax structure, means for receiving a second syntax element in the picture header syntax structure, wherein a first value for the second syntax element indicates that a co-located picture used for temporal motion vector prediction is to be derived from a first reference picture list and a second value for the second syntax element indicates that the co-located picture used for temporal motion vector prediction is to be derived from a second reference picture list; means for receiving a slice of video data referring to the picture header syntax structure; and in response to the slice being a P slice, means for setting a value for a third syntax element associated with the slice to the first value for the third syntax element, wherein the first value for the third syntax element indicates that the co-located picture used for temporal motion vector prediction is to be derived from the first reference picture list and a second value for the third syntax element indicates that the co-located picture used for temporal motion vector prediction is to be derived from the second reference picture list.

[0011] According to another example of the disclosure, a method for encoding video data includes, in response to determining that reference picture list information is contained in a picture header syntax structure, generating a first syntax element indicating that the reference picture list information is included in the picture header syntax structure; generating a second syntax element for inclusion in the picture header syntax structure, wherein a first value for the second syntax element indicates that a co-located picture used for temporal motion vector prediction is to be derived from a first reference picture list and a second value for the second syntax element indicates that the co-located picture used for temporal motion vector prediction is to be derived from a second reference picture list; determining that a slice of video data referring to the picture header syntax structure is a P slice; in response to the slice being a P slice, determining that a value for a third syntax element associated with the slice is equal to the first value for the third syntax element, wherein the first value for the third syntax element indicates that the co-located picture used for temporal motion vector prediction is to be derived from the first reference picture list and a second value for the third syntax element indicates that the co-located picture used for temporal motion vector prediction is to be derived from the second reference picture list; and outputting a bitstream of encoded video data, wherein the encoded video data contains the first syntax element and the picture header syntax structure.

[0012] According to another example of the disclosure, an apparatus for encoding video data includes a memory configured to store video data, and one or more processors implemented in circuitry and configured to, in response to determining that reference picture list information is included in a picture header syntax structure, generate a first syntax element indicating that the reference picture list information is included in the picture header syntax structure, generate a second syntax element for inclusion in the picture header syntax structure, wherein a first value for the second syntax element indicates that a collocated picture used for temporal motion vector prediction is to be derived from a first reference picture list and a second value for the second syntax element indicates that the collocated picture used for temporal motion vector prediction is to be derived from a second reference picture list, determine that a slice of video data referring to the picture header syntax structure is a P slice, in response to the slice being a P slice, determine that a value for a third syntax element associated with the slice is equal to the first value for the third syntax element, wherein the first value for the third syntax element indicates that the collocated picture used for temporal motion vector prediction is to be derived from the first reference picture list and a second value for the third syntax element indicates that the collocated picture used for temporal motion vector prediction is to be derived from the second reference picture list, and output a bitstream of encoded video data, wherein the encoded video data includes the first syntax element and the picture header syntax structure.

[0013] According to another example of the disclosure, a computer-readable storage medium stores instructions that, when executed by one or more processors, cause the one or more processors to: in response to determining that reference picture list information is included in a picture header syntax structure, generate a first syntax element indicating that the reference picture list information is included in the picture header syntax structure, generate a second syntax element for inclusion in the picture header syntax structure, wherein a first value for the second syntax element indicates that a collocated picture used for temporal motion vector prediction is to be derived from a first reference picture list and a second value for the second syntax element indicates that the collocated picture used for temporal motion vector prediction is to be derived from a second reference picture list, determine that a slice of video data referring to the picture header syntax structure is a P slice, in response to the slice being a P slice, determine that a value for a third syntax element associated with the slice is equal to the first value for the third syntax element, wherein the first value for the third syntax element indicates that the collocated picture used for temporal motion vector prediction is to be derived from the first reference picture list and a second value for the third syntax element indicates that the collocated picture used for temporal motion vector prediction is to be derived from the second reference picture list, and output a bitstream of encoded video data, wherein the encoded video data includes the first syntax element and the picture header syntax structure.

[0014] According to another example of the disclosure, an apparatus for encoding video data includes means for generating, in response to determining that reference picture list information is included in a picture header syntax structure, a first syntax element that indicates that the reference picture list information is included in the picture header syntax structure, means for generating a second syntax element for inclusion in the picture header syntax structure, wherein a first value for the second syntax element indicates that a collocated picture used for temporal motion vector prediction is to be derived from a first reference picture list and a second value for the second syntax element indicates that the collocated picture used for temporal motion vector prediction is to be derived from a second reference picture list, means for determining that a slice of video data referred to by the picture header syntax structure is a P slice, means for determining, in response to the slice being a P slice, a value for a third syntax element associated with the slice is equal to the first value for the third syntax element, wherein the first value for the third syntax element indicates that the collocated picture used for temporal motion vector prediction is to be derived from the first reference picture list and a second value for the third syntax element indicates that the collocated picture used for temporal motion vector prediction is to be derived from the second reference picture list, and means for outputting a bitstream of encoded video data, wherein the encoded video data includes the first syntax element and the picture header syntax structure. BRIEF DESCRIPTION OF DRAWINGS

[0015] FIG. 1 A block diagram of an example video encoding and decoding system that can perform the techniques of this disclosure is shown.

[0016] FIG. 2A And FIG. 2B A conceptual diagram showing an example quad-tree binary-tree (QTBT) structure and corresponding coding tree unit (CTU) is shown.

[0017] FIG. 3 A block diagram of an example video encoder that can perform the techniques of this disclosure is shown.

[0018] FIG. 4 A block diagram of an example video decoder that can perform the techniques of this disclosure is shown.

[0019] FIG. 5 A flowchart of an example video encoding process is shown.

[0020] FIG. 6 A flowchart of an example video decoding process is shown.

[0021] FIG. 7 A flowchart of an example video encoding process is shown.

[0022] FIG. 8 A flowchart of an example video decoding process is shown. DETAILED DESCRIPTION

[0023] Video coding (e.g., video encoding and / or video decoding) often involves predicting a block of video data from already coded blocks of video data in the same picture (e.g., intra prediction) or from already coded blocks of video data in different pictures (e.g., inter prediction). In some cases, a video encoder also computes residual data by comparing a predicted block to an original block. Thus, the residual data represents the difference between the predicted block and the original block. To reduce the number of bits needed to signal the residual data, the video encoder transforms and quantizes the residual data and signals the transformed and quantized residual data in an encoded bitstream. The compression achieved by the transform and quantization process can be lossy, meaning that the transform and quantization process can introduce distortion into the decoded video data.

[0024] A video decoder decodes the residual data and adds it to the predicted block to produce a reconstructed video block that more closely matches the original video block than the predicted block alone. Due to the loss introduced by the transform and quantization of the residual data, the first reconstructed block can have distortion or artifacts. One common type of artifact or distortion is referred to as blocking artifacts, in which the boundaries of the blocks of video data used for coding are visible.

[0025] To further improve the quality of the decoded video, a video decoder can perform one or more filtering operations on the reconstructed video block. Examples of these filtering operations include a deblocking filter, a sample adaptive offset (SAO) filter, and an adaptive loop filter (ALF). Parameters for these filtering operations can be determined by a video encoder and explicitly signaled in an encoded video bitstream, or can be implicitly determined by the video decoder without needing to explicitly signal the parameters in the encoded video bitstream.

[0026] The techniques of this disclosure relate to inter prediction and, more specifically, to signaling of reference pictures used in inter prediction. Inter prediction in ITU-T H.266, also referred to as Versatile Video Coding (VVC), utilizes reference picture lists, which generally refer to lists of reference pictures used for inter prediction of P or B slices. Inter prediction in VVC utilizes two reference picture lists, where reference picture list 0 refers to a reference picture list used for inter prediction of P slices or a first reference picture list used for inter prediction of B slices, and reference picture list 1 refers to a second reference picture list used for inter prediction of B slices. For a decoding process of a predictive (P) slice, only reference picture list 0 is used for inter prediction. For a decoding process of a bi-predictive (B) slice, both reference picture list 0 and reference picture list 1 are used for inter prediction. For a decoding of slice data of an intra (I) slice, no reference picture list is used for inter prediction.

[0027] A P slice is a slice that is decoded using intra prediction or inter prediction with at most one motion vector and reference index to predict sample values of each block. A B slice is a slice that is decoded using intra prediction or inter prediction with at most two motion vectors and reference indexes to predict sample values of each block. An I slice is a slice that is decoded using only intra prediction.

[0028] Information for generating and maintaining reference picture lists is signaled in various syntax structures, including a sequence parameter set (SPS) syntax structure, a picture parameter set (PPS) syntax structure, a picture header (PH) syntax structure, and a slice header (SH) syntax structure. In VVC, SPS refers to a syntax structure that includes syntax elements that apply to zero or more entire coded layer video sequences (CLVSs) as determined by the content of syntax elements found in PPSs, which are referred to as syntax structures that include syntax elements that apply to zero or more complete coded pictures as determined by syntax elements found in each PH. In VVC, PH refers to a syntax structure that includes syntax elements that apply to all slices of a coded picture, and SH refers to a portion of a coded slice that includes data elements related to all tiles or coded tree unit (CTU) rows within tiles represented in a slice.

[0029] As discussed in more detail below, the reference picture list signaling used in VVC Draft 8 can have some inefficiencies and other potential issues. For example, even though a PH can have an associated P slice that uses only reference picture list 0, the signaling technique of VVC Draft 8 allows PH-level signaling that indicates that a co-located picture to be used for temporal motion vector prediction is to be derived from reference picture list 1. According to the techniques described in this disclosure, a video decoder can be configured to, in response to a slice being a P slice, and a value of a syntax element associated with the slice being set to a value of a syntax element that indicates that a co-located picture to be used for temporal motion vector prediction is to be derived from reference picture list 0. As described in more detail below, a video decoder can be configured to infer the value that indicates that a co-located picture to be used for temporal motion vector prediction is to be derived from reference picture list 0 without actually receiving an instance of the syntax element, effectively overriding the PH-level syntax for a P slice. Thus, the techniques of this disclosure can advantageously enable a video decoder to parse and decode such bitstreams without unduly restricting the manner in which a video encoder encodes the bitstream.

[0030] FIG. 1A block diagram of an example video encoding and decoding system 100 that can perform the techniques of this disclosure is shown. The techniques of this disclosure are generally directed toward coding (encoding and / or decoding) video data. In general, video data includes any data used for processing video. Thus, video data can include raw, unencoded video, encoded video, decoded (e.g., reconstructed) video, and video metadata (such as signaling data).

[0031] As FIG. 1 shown, in this example, system 100 includes a source device 102 that provides encoded video data to be decoded and displayed by a destination device 116. In particular, source device 102 provides the video data to destination device 116 via a computer-readable medium 110. Source device 102 and destination device 116 can comprise any of a wide variety of devices, including desktop computers, notebook (i.e., laptop) computers, tablet computers, set-top boxes, telephone handsets such as smartphones, televisions, cameras, display devices, digital media players, video gaming consoles, video streaming devices, or the like. In some cases, source device 102 and destination device 116 can be equipped for wireless communication, and thus can be referred to as wireless communication devices.

[0032] In FIG. 1 the example of FIG. 1, source device 102 includes a video source 104, a memory 106, a video encoder 200, and an output interface 108. Destination device 116 includes an input interface 122, a video decoder 300, a memory 120, and a display device 118. In accordance with this disclosure, video encoder 200 of source device 102 and video decoder 300 of destination device 116 can be configured to apply the techniques for reference picture list and collocated picture signaling described herein. Thus, source device 102 represents an example of a video encoding device, while destination device 116 represents an example of a video decoding device. In other examples, a source device and a destination device can include other components or

[0033] As FIG. 1The illustrated system 100 is merely one example. In general, any digital video encoding and / or decoding device can perform the techniques for reference picture list and collocated picture signaling described herein. The source device 102 and the destination device 116 are merely examples of such coding devices in which the source device 102 generates coded video data for transmission to the destination device 116. This disclosure refers to a "coding device" as a device that performs coding (encoding and / or decoding) of data. Thus, video encoder 200 and video decoder 300 represent examples of coding devices, in particular examples of a video encoder and a video decoder, respectively. In some examples, the source device 102 and the destination device 116 can operate in a substantially symmetrical manner, such that each of the source device 102 and the destination device 116 includes video encoding and decoding components. Hence, the system 100 can support one-way or two-way video transmission between the source device 102 and the destination device 116, e.g., for video streaming, video playback, video broadcasting, or video telephony.

[0034] In general, the video source 104 represents a source of video data (i.e., raw, unencoded video) and provides a sequence of pictures (also referred to as "frames") of the video data to the video encoder 200, which encodes data for the pictures. The video source 104 of the source device 102 can include a video capture device, such as a video camera, a video archive containing previously captured raw video, and / or a video feed interface to receive video from a video content provider. As a further alternative, the video source 104 can generate computer graphics-based data as the source video, or a combination of live video, archived video, and computer-generated video. In each case, the video encoder 200 encodes the captured, pre-captured, or computer-generated video data. The video encoder 200 can rearrange the pictures from the received order (sometimes referred to as "display order") into the coding order for coding. The video encoder 200 can generate a bitstream including encoded video data. The source device 102 can then output the encoded video data via the output interface 108 onto the computer- readable medium 110 for reception and / or retrieval by an input interface 122 of, for example, the destination device 116.

[0035] Memory 106 of source device 102 and memory 120 of destination device 116 represent general storage memory. In some examples, memories 106, 120 can store raw video data, e.g., raw video from video source 104 and raw decoded video data from video decoder 300. Additionally or alternatively, memories 106, 120 can store software instructions executable by, e.g., video encoder 200 and video decoder 300, respectively. While memories 106 and 120 are shown as separate from video encoder 200 and video decoder 300 in this example, it should be understood that video encoder 200 and video decoder 300 can also include internal memories for similar or equivalent purposes. Furthermore, memories 106, 120 can store encoded video data, e.g., encoded video data output from video encoder 200 and input to video decoder 300. In some examples, portions of memories 106, 120 can be allocated as one or more video buffers, e.g., to store raw, decoded, and / or encoded video data.

[0036] Computer-readable medium 110 can represent any type of medium or device capable of transporting the encoded video data from source device 102 to destination device 116. In one example, computer-readable medium 110 represents a communication medium to enable source device 102 to transmit encoded video data directly to destination device 116 in real-time, e.g., via a radio frequency network or computer-based network. Output interface 108 can modulate transmission signals including the encoded video data, and input interface 122 can demodulate received transmission signals, according to a communication standard, such as a wireless communication protocol. The communication medium can comprise any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network, such as a local area network, a wide-area network, or a global network such as the Internet. The communication medium can include routers, switches, base stations, or any other equipment that can be useful to facilitate communication from source device 102 to destination device 116.

[0037] In some examples, source device 102 can output encoded data from output interface 108 to storage device 112. Similarly, destination device 116 can access encoded data from storage device 112 via input interface 122. Storage device 112 can include any of a variety of distributed or locally accessed data storage media such as a hard drive, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data.

[0038] In some examples, source device 102 can output encoded video data to a file server 114 or another intermediate storage device that can store the encoded video generated by source device 102. Destination device 116 can access stored video data from file server 114 through streaming or download. File server 114 can be any type of server device capable of storing encoded video data and transmitting that encoded video data to the destination device 116. File server 114 can represent a web server (e.g., for a website), a file transfer protocol (FTP) server, a content delivery network device, or a network attached storage (NAS) device. Destination device 116 can access encoded video data from file server 114 through any standard data connection, including an Internet connection. This can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both that is suitable for accessing encoded video data stored on file server 114. File server 114 and input interface 122 can be configured to operate according to a streaming protocol, a download transmission protocol, or a combination thereof.

[0039] Output interface 108 and input interface 122 can represent wireless transmitters / receivers, modems, wired networking components (e.g., Ethernet cards), wireless communication components operating according to any of a variety of IEEE 802.11 standards, or other physical components. In examples where output interface 108 and input interface 122 comprise wireless components, output interface 108 and input interface 122 can be configured to transmit data, such as encoded video data, according to a cellular communication standard, such as 4G, 4G-LTE (Long-Term Evolution), LTE Advanced, 5G, or the like. In some examples where output interface 108 includes a wireless transmitter, output interface 108 and input interface 122 can be configured to transmit data, such as encoded video data, according to other wireless standards, such as an IEEE 802.11 specification, an IEEE 802.15 specification (e.g., ZigBee), a Bluetooth Standard, or the like. In some examples, source device 102 and / or destination device 116 can include respective system on a chip (SoC) devices. For example, source device 102 can include a SoC device that performs the functions attributable to video encoder 200 and / or output interface 108, and destination device 116 can include a SoC device that performs the functions attributable to video decoder 300 and / or input interface 122. TM TM In some examples, source device 102 and / or destination device 116 can include respective system on a chip (SoC) devices. For example, source device 102 can include a SoC device that performs the functions attributable to video encoder 200 and / or output interface 108, and destination device 116 can include a SoC device that performs the functions attributable to video decoder 300 and / or input interface 122.

[0040] ​The techniques of this disclosure can be applied to video coding in support of any of a variety of multimedia applications, such as over-the-air television broadcasts, cable television transmissions, satellite television transmissions, Internet streaming video transmissions, such as dynamic adaptive streaming over HTTP (DASH), digital video that is encoded onto a data storage medium, decoding of digital video stored on a data storage medium, or other applications.

[0041] Input interface 122 of destination device 116 receives an encoded video bitstream from computer- readable medium 110 (e.g., a communication medium, storage device 112, file server 114, or the like). The encoded video bitstream can include signaling information defined by video encoder 200, which is also used by video decoder 300, such as syntax elements having values that describe properties and / or processing of video blocks or other coded units (e.g., slices, pictures, groups of pictures, sequences, or the like). Display device 118 displays decoded pictures of the decoded video data to a user. Display device 118 can represent any of a variety of display devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.

[0042] Although not shown in FIG. 1, in some examples, video encoder 200 and video decoder 300 can each be integrated with an audio encoder and / or audio decoder. The video and audio encoders can be integrated as a combined video and audio encoder, and the video and audio decoders can be integrated as a combined video and audio decoder. In such examples, the combined video and audio encoder can include MUX-DEMUX units, or other hardware and / or software, to handle multiplexed streams of audio and video, if applicable. If applicable, MUX-DEMUX units can be configured to FIG. 1

[0043] Video encoder 200 and video decoder 300 each can be implemented as any of a variety of suitable encoder and / or decoder circuitry, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic circuitry, software, hardware, firmware, or any combinations thereof. When the techniques are implemented in software, adevice can store instructions for the software in a suitable, non- transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Each of video encoder 200 and video decoder 300 can be included in one or more encoders or decoders, either of which can be integrated as part of a combined video encoder / decoder (CODEC) in a respective device. A device including video encoder 200 and / or video decoder 300 can comprise an integrated circuit, a microprocessor, and / or a wireless communication device, such as a cellular telephone.

[0044] Video encoder 200 and video decoder 300 can operate according to a video coding standard, such as ITU-T H.265, also referred to as High Efficiency Video Coding (HEVC), or extensions thereto, such as multi-view and / or scalable video coding extensions. Alternatively, video encoder 200 and video decoder 300 can operate according to other proprietary or industry standards, such as VVC. A recent draft of the VVC standard is described in Bross, et al., “Versatile Video Coding (Draft 8),” JVET-Q2001-v13, 17th Meeting of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, Joint Video Team (JVET) of Brussels, Belgium, Jan. 7-17, 2020 (hereinafter “VVC Draft 8”). The techniques of this disclosure, however, are not limited to any particular coding standard.

[0045] In general, video encoder 200 and video decoder 300 can perform block-based picture coding. The term “block” generally refers to a structure comprising data to be processed (e.g., encoded, decoded, or otherwise used in the encoding and / or decoding process). For example, a block can comprise a two-dimensional matrix of samples of luma and / or chroma data. In general, video encoder 200 and video decoder 300 can code video data represented in a YUV (e.g., Y, Cb, Cr) format. That is, video encoder 200 and video decoder 300 can code luma and chroma components as opposed to coding picture samples of Red, Green, and Blue (RGB) data. In some examples, video encoder 200 converts received data in an RGB format to a YUV representation prior to encoding and video decoder 300 converts the YUV representation to the RGB format. Alternatively, pre- and post-processing units (not shown) can perform these conversions.

[0046] This disclosure can generally relate to coding (e.g., encoding and decoding) of pictures, to include processes that encode or decode data of a picture. Similarly, this disclosure can relate to coding of blocks of pictures, to include processes that encode or decode data of a block, such as prediction and / or residual coding. An encoded video bitstream generally includes a series of values representing coding decisions (e.g., coding modes) and syntax elements partitioning a picture into blocks. Accordingly, a reference to coding a picture or block should generally be understood to code values of syntax elements forming the picture or block.

[0047] HEVC defines various blocks, including coding units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video coder, such as video encoder 200, partitions a coding tree unit (CTU) into CUs according to a quad tree structure. That is, the video coder partitions a CTU and a CU into four equal, non overlapping squares, and each node of the quad tree has either zero or four child nodes. Nodes with zero child nodes can be referred to as“leaf nodes,” and CUs of such leaf nodes can include one or more PUs and / or one or more TUs. The video coder can further partition PUs and TUs. For example, in HEVC, a residual quad tree (RQT) represents partitioning of TUs. In HEVC, PUs represent inter prediction data, while TUs represent residual data. Intra predicted CUs include intra prediction information, such as an intra mode indication.

[0048] As another example, video encoder 200 and video decoder 300 can be configured to operate according to VVC. According to VVC, a video coder, such as video encoder 200, partitions a picture into coding tree units (CTUs) according to a tree structure. Video encoder 200 can partition a CTU according to a quad tree binary tree (QTBT) structure or a multi-type tree (MTT) structure. The QTBT structure removes the concepts of multiple partition types, such as the separation between CUs, PUs, and TUs of HEVC. The QTBT structure includes two layers: a first layer partitioned according to quad tree partitioning, and a second layer partitioned according to binary tree partitioning. A root node of the QTBT structure corresponds to a CTU. Leaf nodes of the binary trees correspond to coding units (CUs).

[0049] In the MTT partitioning structure, blocks can be partitioned using quad tree (QT) partitioning, binary tree (BT) partitioning, and one or more types of ternary tree (TT) (also referred to as tri-tree (TT)) partitioning. Ternary or tri-tree partitioning is a partitioning that splits a block into three sub-blocks. In some examples, ternary or tri-tree partitioning divides a block into three sub-blocks without dividing the original block through a center. The partition types (e.g., QT, BT, and TT) in the MTT can be symmetric or asymmetric.

[0050] In some examples, video encoder 200 and video decoder 300 can use a single QTBT or MTT structure to represent each of luma and chroma samples, while in other examples, video encoder 200 and video decoder 300 can use two or more QTBT or MTT structures, such as one QTBT / MTT structure for luma samples and another QTBT / MTT structure for two chroma samples (or two QTBT / MTT structures for respective chroma samples).

[0051] Video encoder 200 and video decoder 300 can be configured to use quadtree partitioning, QTBT partitioning, MTT partitioning, or other partition structures in accordance with HEVC. For purposes of illustration, the description of the techniques of this disclosure is presented with respect to QTBT partitioning. However, it should be understood that the techniques of this disclosure can also be applied to video coders configured to use quadtree partitioning or other types of partitioning.

[0052] Blocks (e.g., CTUs or CUs) can be grouped in various ways in a picture. As one example, a brick can refer to a rectangular region of CTU rows within a particular tile in a picture. A tile can be a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. A tile column refers to a rectangular region of CTUs having a height equal to a height of the picture and a width specified by a syntax element (e.g., in a picture parameter set). A tile row refers to a rectangular region of CTUs having a height specified by a syntax element (e.g., in a picture parameter set) and a width equal to a width of the picture.

[0053] In some examples, a tile can be partitioned into multiple bricks, each of which can include one or more CTU rows within the tile. A tile that is not partitioned into multiple bricks can also be referred to as a brick. However, a brick that is a true subset of a tile cannot be referred to as a tile.

[0054] Bricks in a picture can also be arranged in slices. A slice can be an integer number of bricks of a picture that can be contained exclusively in a single network abstraction layer (NAL) unit. In some examples, a slice includes multiple complete tiles or only a contiguous sequence of all bricks of one tile.

[0055] This disclosure can use “NxN” and “N by N” interchangeably to refer to the sample dimensions of a block (such as a CU or other video block) in terms of vertical and horizontal dimensions, e.g., 16x16 samples or 16 by 16 samples. In general, a 16x16 CU will have 16 samples in a vertical direction (y = 16) and 16 samples in a horizontal direction (x = 16). Likewise, an NxN CU generally has N samples in a vertical direction and N samples in a horizontal direction, where N represents a nonnegative integer value. The samples in a CU can be arranged in rows and columns. Moreover, a CU need not necessarily have the same number of samples in a horizontal direction as in a vertical direction. For example, a CU can comprise NxM samples, where M is not necessarily equal to N.

[0056] Video encoder 200 encodes video data of CUs that represent prediction and / or residual information, among other information. Prediction information indicates how to predict the CU in order to form a prediction block for the CU. Residual information generally represents a sample-by-sample difference between the CU prior to encoding and the prediction block.

[0057] To code a CU, video encoder 200 generally can form a prediction block for the CU through inter prediction or intra prediction. Inter prediction generally refers to predicting the CU from data of a previously coded picture, while intra prediction generally refers to predicting the CU from previously coded data of the same picture. To perform inter prediction, video encoder 200 can use one or more motion vectors to generate the prediction block. Video encoder 200 can generally perform a motion search to identify a reference block that closely matches the CU, e.g., in terms of differences between the CU and the reference block. Video encoder 200 can calculate a difference metric using a sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared differences (MSD), or other such difference calculations to determine whether a reference block closely matches a current CU. In some examples, video encoder 200 can use uni -prediction or bi-prediction to predict a current CU.

[0058] Some examples of VVC also provide an affine motion compensation mode, which can be considered an inter prediction mode. In the affine motion compensation mode, video encoder 200 can determine two or more motion vectors that represent non-translational motion, such as zooming or scaling, rotation, perspective motion, or other irregular types of motion.

[0059] To perform intra prediction, video encoder 200 can select an intra prediction mode to generate the prediction block. Some examples of VVC provide sixty-seven intra prediction modes, including various directional modes, as well as a planar mode and a DC mode. Generally, video encoder 200 selects an intra prediction mode that describes neighboring samples of a current block (e.g., a block of a CU) from which to predict samples of the current block. Assuming video encoder 200 is coding CTUs and CUs in a raster scan order (left to right, top to bottom), such samples can generally be above, above and to the left, or to the left of the current block in the same picture as the current block.

[0060] Video encoder 200 encodes data representing the prediction mode for the current block. For example, for inter prediction modes, video encoder 200 can encode data representing which of various available inter prediction modes to use, as well as motion information for the corresponding mode. For example, for uni - or bi-prediction, video encoder 200 can use advanced motion vector prediction (AMVP) or merge mode to encode motion vectors. Video encoder 200 can use similar modes to encode motion vectors for affine motion compensation modes.

[0061] Following prediction, such as intra prediction or inter prediction of a block, video encoder 200 can calculate residual data for the block. The residual data, such as a residual block, represents sample-by-sample differences between the block and a prediction block formed using a corresponding prediction mode for the block. Video encoder 200 can apply one or more transforms to the residual block to produce transform data in a transform domain instead of the sample domain. For example, video encoder 200 can apply a discrete cosine transform (DCT), an integer transform, a wavelet transform, or a conceptually similar transform. Additionally, video encoder 200 can apply a secondary transform, such as a mode-dependent non-separable secondary transform (MDNSST), a signal-dependent transform, a Karhunen-Loeve transform (KLT), and so on, following the first transform. Video encoder 200 produces transform coefficients following application of the one or more transforms.

[0062] As described above, following any transform to produce transform coefficients, video encoder 200 can perform quantization of the transform coefficients. Quantization generally refers to a process in which transform coefficients are quantized to reduce the number of bits used to represent the transform coefficients, thereby providing further compression. By performing the quantization process, video encoder 200 can reduce the bit depth of some or all of the transform coefficients associated with the transform coefficients. For example, video encoder 200 can round n-bit values down to m-bit values during quantization, where n is greater than m. In some examples, to perform quantization, video encoder 200 can perform a bitwise right-shift of the values to be quantized.

[0063] Following quantization, video encoder 200 can scan the transform coefficients, producing a one-dimensional vector from the two-dimensional matrix including the quantized transform coefficients. The scan can be designed to place higher energy (and thus lower frequency) transform coefficients at a front of the vector, and lower energy (and thus higher frequency) transform coefficients at a back of the vector. In some examples, video encoder 200 can utilize a pre-defined scan order to scan the quantized transform coefficients to produce a serialized vector, and then entropy encode the quantized transform coefficients of the vector. In other examples, video encoder 200 can perform an adaptive scan. Following scanning of the quantized transform coefficients to form a one-dimensional vector, video encoder 200 can entropy encode the one-dimensional vector, e.g., according to context adaptive binary arithmetic coding (CABAC). Video encoder 200 can also entropy encode values for syntax elements describing metadata associated with the encoded video data for use by video decoder 300 when decoding the video data.

[0064] To perform CABAC, video encoder 200 can assign a context within a context model to a symbol to be transmitted. The context can relate to, for example, whether neighboring values of the symbol are zero-valued or not. Probability determination can be based on the context assigned to the symbol.

[0065] Video encoder 200 can also generate syntax data, such as block-based syntax data, picture-based syntax data, and sequence-based syntax data, or other syntax data, such as a sequence parameter set (SPS), a picture parameter set (PPS), or a video parameter set (VPS), to video decoder 300, e.g., in a PH, a block header, a SH. Likewise, video decoder 300 can decode such syntax data to determine how to decode corresponding video data.

[0066] In this way, video encoder 200 can generate a bitstream that includes encoded video data (e.g., syntax elements that describe partitioning of a picture into blocks (e.g., CUs) and prediction and / or residual information for the blocks). Ultimately, video decoder 300 can receive the bitstream and decode the encoded video data.

[0067] In general, video decoder 300 performs a reciprocal process as that performed by video encoder 200 to decode the encoded video data of the bitstream. For example, video decoder 300 can decode values for syntax elements of the bitstream using CABAC in a manner substantially similar to, but reciprocal to, the CABAC encoding process of video encoder 200. The syntax elements can define partitioning information of a picture into CTUs, and partitioning of each CTU according to a corresponding partition structure, such as a QTBT structure, to define CUs of the CTU. The syntax elements can further define prediction and residual information for blocks (e.g., CUs) of the video data.

[0068] The residual information can be represented by, for example, quantized transform coefficients. Video decoder 300 can inverse quantize and inverse transform the quantized transform coefficients of a block to reproduce a residual block for the block. Video decoder 300 forms a prediction block for the block using the signaled prediction mode (intra- or inter-prediction) and related prediction information (e.g., motion information for inter-prediction). Video decoder 300 can then combine the prediction block and the residual block (on a sample-by-sample basis) to reproduce the original block. Video decoder 300 can perform additional processing, such as performing a deblocking process to reduce visual artifacts along boundaries of the blocks.

[0069] This disclosure can generally relate to “signaling” certain information, such as syntax elements. The term “signaling” can generally relate to the communication of values for syntax elements and / or other data used to decode encoded video data. That is, video encoder 200 can signal values for syntax elements in a bitstream. In general, signaling involves generating the values in the bitstream. As described above, source device 102 can transmit the bitstream to destination device 116 in substantially real time or non-real time, such as can occur when syntax elements are stored to storage device 112 for later retrieval by destination device 116.

[0070] FIG. 2A and 2B is a conceptual diagram illustrating an example quad-tree binary-tree (QTBT) structure 130 and a corresponding coding tree unit (CTU) 132. Solid lines represent quad-tree splitting, and dashed lines indicate binary-tree splitting. In each split (i.e., non-leaf) node of the binary tree, a flag is signaled to indicate which type of splitting is used (i.e., horizontal or vertical), where 0 indicates horizontal splitting and 1 indicates vertical splitting in this example. For quad-tree splitting, the split type does not need to be indicated because a quad-tree node splits a block horizontally and vertically into 4 sub-blocks with equal size. Thus, video encoder 200 can encode and video decoder 300 can decode syntax elements (such as splitting information) for a region tree layer of QTBT structure 130 (i.e., solid lines) and syntax elements (such as splitting information) for a prediction tree layer of QTBT structure 130 (i.e., dashed lines). Video encoder 200 can encode and video decoder 300 can decode video data, such as prediction and transform data, for CUs represented by terminal leaf nodes of QTBT structure 130.

[0071] In general, FIG. 2B CTU 132 of FIG. 13A can be associated with parameters defining sizes of blocks corresponding to nodes of the first and second layers of QTBT structure 130. These parameters can include a CTU size (representing a size of CTU 132 in samples), a minimum quad-tree size (MinQTSize representing a minimum allowed quad-tree leaf node size), a maximum binary-tree size (MaxBTSize representing a maximum allowed binary-tree root node size), a maximum binary-tree depth (MaxBTDepth representing a maximum allowed binary-tree depth), and a minimum binary-tree size (MinBTSize representing a minimum allowed binary-tree leaf node size).

[0072] A root node of a QTBT structure corresponding to a CTU can have four child nodes at a first level of the QTBT structure, where each child node can be partitioned according to quadtree partitioning. That is, a node of the first level is either a leaf node (having no child nodes) or has four child nodes. The example of QTBT structure 130 represents such a node as including a parent node and child nodes with solid lines for branches. If a node of the first level is not larger than a maximum allowed binary tree root node size (MaxBTSize), the node can be further partitioned by a corresponding binary tree. Binary tree partitioning of a node can be iterated until the partitioning of the node results in nodes that are either at a minimum allowed binary tree leaf node size (MinBTSize) or at a maximum allowed binary tree depth (MaxBTDepth). The example of QTBT structure 130 represents such a node as having a dashed line for a branch. Binary tree leaf nodes are referred to as coding units (CUs), which are used for prediction (e.g., intra- or inter-prediction) and transform without any further partitioning. As noted above, a CU can also be referred to as a “video block” or “block.”

[0073] In one example of a QTBT partitioning structure, the CTU size is set to 128x128 (luma samples and two corresponding 64x64 chroma samples), the MinQTSize is set to 16x16, the MaxBTSize is set to 64x64, the MinBTSize (of both width and height) is set to 4, and the MaxBTDepth is set to 4. Quadtree partitioning is first applied to the CTU to generate quadtree leaf nodes. A quadtree leaf node can have a size from 16x16 (i.e., MinQTSize) to 128x128 (i.e., CTU size). If a quadtree leaf node is 128x128, the leaf quadtree node is not further partitioned by a binary tree because its size exceeds the MaxBTSize (i.e., 64x64 in this example). Otherwise, the quadtree leaf node is further partitioned by a binary tree. Thus, a quadtree leaf node is also a root node of a binary tree and the binary tree depth is 0. When the binary tree depth reaches the MaxBTDepth (4 in this example), no further partitioning is allowed. A binary tree node having a width equal to the MinBTSize (4 in this example) means that no further vertical partitioning is allowed. Similarly, a binary tree node having a height equal to the MinBTSize means that no further horizontal partitioning of the binary tree node is allowed. As noted above, leaf nodes of a binary tree are referred to as CUs and are further processed according to prediction and transform without further partitioning.

[0074] In VVC Draft 8, reference picture list (RPL) can be signaled in PH or SH. However, temporal motion vector prediction (TMVP) is signaled in PH only. The number of sent active reference pictures (pictures used for inter prediction) is signaled in SH only.

[0075] RPL is a list of reference pictures used for inter prediction of P or B slices. For each slice of a non-IDR picture, two reference picture lists, i.e., reference picture list 0 and reference picture list 1, are generated. The unique set of pictures referred to by all entries in the two reference picture lists associated with a picture includes all reference pictures that can be used for inter prediction of the associated picture or any picture following the associated picture in decoding order. For the decoding process of P slices, only reference picture list 0 is used for inter prediction. For the decoding process of B slices, both reference picture list 0 and reference picture list 1 are used for inter prediction. For the decoding of slice data of I slices, no reference picture list is used for inter prediction. Reference picture list 0 refers to the reference picture list used for inter prediction of P slices or the first reference picture list used for inter prediction of B slices. Reference picture list 1 refers to the second reference picture list used for inter prediction of B slices.

[0076] PH is a syntax structure containing syntax elements that apply to all slices of a coded picture, and SH is a part of a coded slice containing data elements related to all tiles or CTU rows within tiles represented in the slice. A slice is an integer number of complete tiles or consecutive complete CTU rows in an integer number of pictures of tiles that are completely contained in a single NAL unit.

[0077] In VVC Draft 8, there are the following syntax tables:

[0078] rpl_info_in_ph_flag equal to 1 specifies that the reference picture list information is present in the PH syntax structure and is not present in the slice header referring to a PPS that does not contain the PH syntax structure. rpl_info_in_ph_flag equal to 0 specifies that the reference picture list information is not present in the PH syntax structure and can be present in the slice header referring to a PPS that does not contain the PH syntax structure.

[0079] Picture header

[0080]

[0081]

[0082] ph_inter_slice_allowed_flag equal to 0 specifies that all coded slices of the picture have slice_type equal to 2. ph_inter_slice_allowed_flag equal to 1 specifies that there can be one or more coded slices of the picture with slice_type equal to 0 or 1, or there can be none.

[0083] ph_temporal_mvp_enabled_flag specifies whether temporal motion vector prediction can be used for inter prediction of slices associated with a PH. If ph_temporal_mvp_enabled_flag is equal to 0, the syntax elements of slices associated with a PH shall be constrained such that no temporal motion vector predictor is used in the decoding of the slices. Otherwise (ph_temporal_mvp_enabled_flag is equal to 1), a temporal motion vector predictor can be used in the decoding of slices associated with a PH. If not present, the value of ph_temporal_mvp_enabled_flag is inferred to be equal to 0. The value of ph_temporal_mvp_enabled_flag shall be equal to 0 when there is no reference picture in the DPB that has the same spatial resolution as the current picture.

[0084] ph_collocated_from_l0_flag equal to 1 specifies that the collocated picture for temporal motion vector prediction is derived from reference picture list 0. ph_collocated_from_l0_flag equal to 0 specifies that the collocated picture for temporal motion vector prediction is derived from reference picture list 1.

[0085] ph_collocated_ref_idx specifies the reference index of the collocated picture for temporal motion vector prediction.

[0086] When ph_collocated_from_l0_flag is equal to 1, ph_collocated_ref_idx refers to one entry in the reference picture list 0 and the value of ph_collocated_ref_idx shall be in the range of 0 to num_ref_entries[ 0 ][ PicRplsIdx[ 0 ] ] - 1, inclusive.

[0087] When ph_collocated_from_l0_flag is equal to 0, ph_collocated_ref_idx refers to one entry in the reference picture list 1 and the value of ph_collocated_ref_idx shall be in the range of 0 to num_ref_entries[ 1 ][ PicRplsIdx[ 1 ] ] - 1, inclusive.

[0088] If not present, the value of ph_collocated_ref_idx is inferred to be equal to 0.

[0089] slice header

[0090]

[0091]

[0092] sps_idr_rpl_present_flag equal to 1 specifies that the reference picture list syntax elements are present in the slice header of an IDR picture. sps_idr_rpl_present_flag equal to 0 specifies that the reference picture list syntax elements are not present in the slice header of an IDR picture.

[0093] num_ref_idx_active_override_flag equal to 1 specifies that the syntax elements num_ref_idx_active_minus1[0] are present for P and B slices and the syntax elements num_ref_idx_active_minus1[1] are present for B slices. num_ref_idx_active_override_flag equal to 0 specifies that the syntax elements num_ref_idx_active_minus1[0] and num_ref_idx_active_minus1[1] are not present. If not present, the value of num_ref_idx_active_override_flag is inferred to be equal to 1.

[0094] num_ref_idx_active_minus1[ i ] is used to derive the variable NumRefIdxActive[ i ], as specified in equation 143. The value of num_ref_idx_active_minus1[ i ] shall be in the range of 0 to 14, inclusive.

[0095] For i equal to 0 or 1, when the current slice is a B slice, num_ref_idx_active_override_flag is equal to 1, and num_ref_idx_active_minus1[ i ] is not present, num_ref_idx_active_minus1[ i ] is inferred to be equal to 0.

[0096] When the current slice is a P slice, num_ref_idx_active_override_flag is equal to 1 and num_ref_idx_active_minus1[ 0 ] is not present, num_ref_idx_active_minus1[ 0 ] is inferred to be equal to 0.

[0097] slice_collocated_from_l0_flag equal to 1 specifies that the collocated picture for temporal motion vector prediction is derived from reference picture list 0. slice_collocated_from_l0_flag equal to 0 specifies that the collocated picture for temporal motion vector prediction is derived from reference picture list 1.

[0098] When slice_type is equal to B or P, ph_temporal_mvp_enabled_flag is equal to 1 and slice_collocated_from_l0_flag is not present, the following conditions apply:

[0099] - If rpl_info_in_ph_flag is equal to 1, slice_collocated_from_l0_flag is inferred to be equal to ph_collocated_from_l0_flag.

[0100] - Otherwise rpl_info_in_ph_flag is equal to 0 and if slice_type is equal to P, the value of slice_collocated_from_l0_flag is inferred to be equal to 1.

[0101] slice_collocated_ref_idx specifies the reference index of the collocated picture for temporal motion vector prediction.

[0102] When slice_type is equal to P or slice_type is equal to B and collocated_from_l0_flag is equal to 1, slice_collocated_ref_idx refers to one entry in reference picture list 0 and the value of slice_collocated_ref_idx shall be in the range of 0 to NumRefIdxActive[ 0 ] - 1, inclusive.

[0103] When slice type is equal to B and slice collocated from lO flag is equal to 0, slice collocated ref idx refers to one entry in reference picture list 1, and the value of slice collocated ref idx shall be in the range of 0 to NumRefIdxActive[1] - 1, inclusive.

[0104] When slice collocated ref idx is not present, the following conditions apply:

[0105] - If rpl info in ph flag is equal to 1, the value of slice collocated ref idx is inferred to be equal to ph collocated ref idx.

[0106] - Otherwise (rpl info in ph flag is equal to 0), the value of slice collocated ref idx is inferred to be equal to 0.

[0107] It is a requirement of bitstream conformance that the picture referred to by slice collocated ref idx shall be the same for all slices of a coded picture.

[0108] It is a requirement of bitstream conformance that the values of pic width in luma samples and pic height in luma samples of the reference picture referred to by slice collocated ref idx shall be equal to the values of pic width in luma samples and pic height in luma samples of the current picture, respectively, and RprConstraintsActive[ slice collocated from lO flag? 0 : 1 ][ slice collocated ref idx ] shall be equal to 0.

[0109] The above implementation of PH and SH semantics can have some potential issues. As mentioned before, the TMVP enable flag is only signaled in PH, regardless of whether the RPL is signaled in PH or SH. In addition, VVC Draft 8 imposes a constraint where it is a requirement of bitstream conformance that the picture referred to by slice collocated ref idx shall be the same for all slices of a coded picture.

[0110] In the case where TMVP is enabled in PH and RPL is signaled in PH, the number of active reference pictures in a slice of a picture can be less than the number of reference pictures included in RPL in PH, and the collocated picture can not be included in the active reference pictures.

[0111] In another example of potential issues, for some coding scenarios, TMVP can be enabled in PH, but RPL can be signaled per slice. The number of active reference pictures can vary among slices of the same picture, and for some slices, the collocated picture can not be included in RPL. This disclosure describes techniques to disable TMVP for slices that violate the collocated picture constraint.

[0112] There can be inefficiencies in RPL signaling in SH. Current implementations (e.g., VVC Draft 8) allow RPL to be signaled for I-slices even when there is no RPL for an IDR picture, such as when sps_idr_rpl_present_flag is equal to 0. There can also be inefficiencies in the signaling of weighted prediction in PH, when there is no interrelated syntax, such as when ph_inter_slice_allowed_flag is equal to 0 and there is no RPL for an IDR picture, such as when sps_idr_rpl_present_flag is equal to 0. This disclosure describes techniques that can address the above issues. The various techniques described below can be used independently or in any combination.

[0113] In some examples, to address the redundant signaling of RPL in PH, this disclosure describes techniques to signal RPL when inter-picture prediction is allowed, such as when ph_inter_slice_allowed_flag is equal to 1, or when there is an RPL for an IDR picture, such as when sps_idr_rpl_present_flag is equal to 1.

[0114] In one example, this technique can be implemented with the following syntax changes related to VVC Draft 8. In the examples below, as well as the rest of this disclosure, <add>and< / add> Text between < and > represents text added to VVC Draft 8. <del>and< / del> Text between < and > represents text removed from VVC Draft 8.

[0115]

[0116] Video decoder 300 can be configured to infer that when ph_intra_slice_allowed_flag is not present, the value of ph_intra_slice_allowed_flag is inferred to be equal to 1, which means that when inter prediction is not allowed, then intra prediction should be allowed. Thus, the above condition can be simplified as follows:

[0117]

[0118] Furthermore, video encoder 200 can be configured to signal the weighting prediction table when inter prediction is allowed (e.g., when ph_inter_slice_allowed_flag is equal to 1) or when an IDR picture has RPL (e.g., when sps_idr_rpl_present_flag is equal to 1).

[0119] Example implementations of these techniques are shown below, which are modifications with respect to VVC Draft 8:

[0120]

[0121] Video decoder 300 can be configured to infer that when ph_intra_slice_allowed_flag is not present, the value of ph_intra_slice_allowed_flag is inferred to be equal to 1, which means that when inter prediction is not allowed, then intra prediction should be allowed. Thus, the above condition can be simplified as follows:

[0122]

[0123] The two techniques above use similar conditions, and can be combined to signal ref_pic_lists() and pred_weight_table(), which can be advantageous as it requires less condition checks to implement such conditions.

[0124] Example implementations of these techniques are shown below, which are modifications with respect to VVC Draft 8:

[0125]

[0126] In another example, given the above inference for ph_intra_slice_allowed_flag, these techniques can be implemented with respect to the following VVC Draft 8:

[0127]

[0128] pps_weighted_pred_flag equal to 0 specifies that weighted prediction is not applied to P slices referring to the PPS. pps_weighted_pred_flag equal to 1 specifies that weighted prediction is applied to P slices referring to the PPS. When sps_weighted_pred_flag is equal to 0, the value of pps_weighted_pred_flag shall be equal to 0.

[0129] pps_weighted_bipred_flag equal to 0 specifies that explicit weighted prediction is not applied to B slices referring to the PPS. pps_weighted_bipred_flag equal to 1 specifies that explicit weighted prediction is applied to B slices referring to the PPS. When sps_weighted_bipred_flag is equal to 0, the value of pps_weighted_bipred_flag shall be equal to 0.

[0130] According to the techniques of this disclosure, in order to avoid any inconsistency between the active entries of slices across the same picture, video encoder 200 can be configured to signal the number of active reference picture entries in the PH. In one example, video encoder 200 can signal the number of active reference entries in the PH on the condition that a RPL is present in the PH, e.g., rpl_info_in_ph_flag equal to 1, and, for example, when ph_inter_slice_allowed_flag is equal to 1, e.g., P or B slices can be present in the picture, or when a RPL is present in an IDR picture, e.g., when sps_idr_rpl_present_flag is equal to 1, inter prediction is allowed.

[0131] An example implementation of these techniques is shown below, which is a modification with respect to VVC Draft 8:

[0132]

[0133]

[0134] Video decoder 300 can be configured to infer that when ph_intra_slice_allowed_flag is not present, the value of ph_intra_slice_allowed_flag is inferred to be equal to 1, which means that when inter prediction is not allowed, then intra prediction should be allowed. Thus, the above conditions can be simplified as follows:

[0135]

[0136] <add>ph_num_ref_idx_active_override_flag equal to 1 specifies that the syntax element ph_num_ref_idx_active_minus1 is present in the PH When rpl1_present_flag is equal to 1, the syntax element ph_num_ref_idx_active_minus1[1] is present. ph_num_ref_idx_active_override_flag equal to 1 specifies that the syntax elements ph_num_ref_idx_active_minus1[0] and ph_num_ref_idx_active_minus1[1] are not present. If not present, the value of ph_num_ref_idx_active_override_flag is inferred to be equal to 0.< / add>

[0137] <add>ph_num_ref_idx_active_minus1[ i ] is used to derive the variable NumRefIdxActive[ i ] as specified in Equation 143. The value of ph_num_ref_idx_active_minus1[ i ] shall be in the range of 0, inclusive, to num_ref_entries[ i ][ RplsIdx[ i ] ], inclusive. If ph_num_ref_idx_active_minus1[ i ] is not present, it is inferred to be equal to 0.< / add>

[0138] Additionally, video encoder 200 can be configured to conditionally signal and video decoder 300 is configured to conditionally parse the override flag num_ref_idx_active_override_flag in SH based on the override flag ph_num_ref_idx_active_override_flag value in PH. In one example, video encoder 200 can be configured to signal the override flag num_ref_idx_active_override_flag in SH only if the PH override flag ph_num_ref_idx_active_override_flag is equal to 0.

[0139] Video decoder 300 can be configured to derive the number of active reference pictures in the slice, NumRefIdxActive[i], from the number of active reference pictures signaled in PH if the override flag ph_num_ref_idx_active_override_flag in PH is equal to 1, which is applied to RefPicListO and RefPicListl, respectively.

[0140] Example implementations of these techniques are shown below as modifications to VVC Draft 8:

[0141] The variable NumRefIdxActive[i] is derived as follows:

[0142]

[0143] In the weighted prediction table signaling, there are two syntax elements num_l0_weights and num_l1_weights that indicate the number of weights signaled in RefPicListO and RefPicListl, respectively. When the override flag is signaled in PH, the number of entries in the RPL is known. Therefore, the values of num_l0_weights and / or num_l1_weights can not be needed. Video encoder 200 can be configured to conditionally signal these syntax elements based on the presence of the PH override flag ph_num_ref_idx_active_override_flag, while video decoder 300 can be configured to infer the number of weights for RefPicListO and RefPicListl will be equal to ph_num_ref_idx_active_minus1, respectively.

[0144] Example implementations of these techniques are shown below, as modifications to VVC Draft 8:

[0145]

[0146]

[0147] num_l0_weights specifies the number of weights signaled for entries in the reference picture list. The value of num_l0_weights[ i ] shall be in the range of 0, inclusive, to num_ref_entries[ 0 ][ PicRplsIdx[ 0 ] ], inclusive.

[0148] <add>The variable NumWeightsL0 is derived as follows:

[0149]

[0150] num_l1_weights specifies the number of weights signaled for entries in reference picture list 1. The value of num_l1_weights shall be in the range of 0 to num_ref_entries[1][PicRplsIdx[1]] inclusive.

[0151] <add>The variable NumWeightsL0 is derived as follows:

[0152]

[0153]

[0154] According to the techniques of this disclosure, to address the issue that there can not be a collocated picture in the RPL when TMVP is enabled, video encoder 200 and video decoder 300 can be configured to operate according to a constraint that a collocated picture between active reference pictures should exist.

[0155] In one example, the ph collocated ref idx constraint is modified in the following way: the range of ph collocated ref idx is from 0 to the minimum number of active reference pictures of any slice belonging to the same picture. In this case, it is not possible to signal a collocated reference index in PH to indicate that a collocated picture can not exist in some slices of the same picture.

[0156] In one example, the semantic constraint of ph collocated ref idx can be implemented by modifying it as follows:

[0157] When ph collocated from l0 flag is equal to 1, ph collocated ref idx refers to one entry in the reference picture list 0, and the value of ph collocated ref idx shall be in the range of 0 <add>The minimum value of NumRefIdxActive[ 0 ] - 1, inclusive, over all slices in the picture< / add> ) for the picture.

[0158] When ph collocated from l0 flag is equal to 0, ph collocated ref idx refers to one entry in the reference picture list 1, and the value of ph collocated ref idx shall be in the range of 0 <add>The minimum value of NumRefIdxActive[ 1 ] - 1, inclusive, over all slices in the picture< / add> ) for the picture.

[0159] In some examples, the semantics of ph collocated ref idx can remain unchanged as they can be used in the process of parsing ph collocated ref idx, but bitstream conformance constraints can be added to reflect that ph collocated ref idx shall not exceed the value of num ref idx active minusl [0] [PicRplsIdx [0]] - 1 or the value of num ref idx active minusl [1] [PicRplsIdx [1]] - 1 for any slice in the picture.

[0160] In one example, such bitstream conformance constraints can be expressed as follows:

[0161] It is a requirement of bitstream conformance that the following conditions are met:

[0162] • When ph collocated from lO flag is equal to 1, ph collocated ref idx refers to one entry in reference picture list 0, and the value of ph collocated ref idx shall be in the range of 0 to <add>The minimum value of NumRefIdxActive[ 0 ] - 1, inclusive, over all slices in the picture< / add> .

[0163] • When ph collocated from lO flag is equal to 0, ph collocated ref idx refers to one entry in reference picture list 1, and the value of ph collocated ref idx shall be in the range of 0 to <add>The minimum value of NumRefIdxActive[ 1 ] - 1, inclusive, over all slices in the picture< / add> .

[0164] Thus, in some examples, video encoder 200 and video decoder 300 can operate according to a constraint on NumRefIdxActive of slices included in a picture, i.e., if TMVP is enabled, the number of active reference pictures for a reference picture list, NumRefIdxActive, shall be in the range of ph collocated ref idx to num ref idx active minusl.

[0165] In one example, the constraint can be expressed as follows:

[0166] <add>When ph_temporal_mvp_enabled_flag is equal to 1, NumRefIdxActive[ slice_collocated_from_l0_flag? 0 : 1 ] shall be greater than or equal to slice_collocated_ref_idx.< / add>

[0167] In yet another example, a constraint can be added to require that the collocated picture is an active reference picture, i.e., the collocated picture must be present in the RPL.

[0168] In one example, the constraint can be expressed as follows:

[0169] <add>The collocated picture of a slice shall be a valid entry.< / add>

[0170] Optionally, one example of a check whether TMVP is enabled can be added in the constraint, as follows:

[0171] <add>If TMVP is enabled, the collocated picture of a slice shall be a valid entry.< / add>

[0172] In VVC Draft 8, ph collocated from l0 flag can be used to signal in PH the reference picture list from which the collocated picture is derived. When this flag ph collocated from l0 flag is equal to 0, i.e., the collocated picture is derived from RefPicListl, and there are P and B slices in the picture, a problem can arise. In this case, because there is no RefPicListl for P slices, it is not possible to derive a collocated picture for P slices.

[0173] To solve this problem, video encoder 200 and video decoder 300 can be configured to operate according to the constraint that when there are P slices in the picture, ph collocated from l0 flag should not be equal to 1. In one example, the constraint can be expressed as follows:

[0174] ph collocated from l0 flag equal to 1 specifies that the collocated picture for temporal motion vector prediction is derived from reference picture list 0. ph collocated from l0 flag equal to 0 specifies that the collocated picture for temporal motion vector prediction is derived from reference picture list 1.

[0175] <add>When ph_temporal_mvp_enabled_flag is equal to 1, if there is at least one P slice in the picture, the value of ph_collocated_from_l0_flag shall be equal to 1.< / add>

[0176] Accordingly, video encoder 200 and video decoder 300 can be configured to operate according to the constraint that when slice type is P and TMVP is enabled, then ph collocated from l0 flag should be equal to 1. In one example, the constraint can be expressed as follows:

[0177] slice type specifies the coding type of the slice according to table 9.

[0178] Table 9 - Names associated with slice type

[0179]

[0180] <add>When slice_type is equal to 1, the value of slice_collocated_from_l0_flag shall be equal to 1.< / add>

[0181] In some examples, when slice collocated from l0 flag is not signaled due to ph collocated from l0 flag signaling in PH, video decoder 300 can be configured to infer slice collocated from l0 flag equal to 1 for P slices. In one example, the constraint can be expressed as follows:

[0182] slice_collocated_from_l0_flag equal to 1 specifies that the collocated picture for temporal motion vector prediction is derived from reference picture list 0. slice_collocated_from_l0_flag equal to 0 specifies that the collocated picture for temporal motion vector prediction is derived from reference picture list 1.

[0183] When slice_type is equal to B or P, ph_temporal_mvp_enabled_flag is equal to 1, and slice_collocated_from_l0_flag is not present, the following condition applies:

[0184] <add>

[0185] - If slice type is equal to P, the value of slice collocated from lO flag is inferred to be equal to 1.

[0186] - Otherwise, if rpl info in ph flag is equal to 1, slice collocated from lO flag is inferred to be equal to ph collocated from lO flag.

[0187] < / add>

[0188] That is, for P slices, video decoder 300 can be configured to set slice_collocated_from_l0_flag equal to 1, meaning that prediction is performed according to l0 regardless of the value of ph_collocated_from_l0_flag being 0 or 1.

[0189] To address the issue of no collocated picture in RPL, techniques of the present disclosure include disabling TMVP if there is no collocated picture in RPL but TMVP is enabled. In one example, video encoder 200 can be configured to signal a TMVP enabling flag in SH in addition to signaling the TMVP enabling flag in PH. In one example, if TMVP is enabled in PH, then TMVP can be allowed to be disabled on a slice-by-slice basis for slices that violate the collocated picture constraint condition. In another example, even if TMVP is enabled, video encoder 200 and video decoder 300 can be configured to not apply TMVP on a block basis if there is no collocated picture in RPL.

[0190] Video encoder 200 can be configured to conditionally signal a TMVP enabling flag in a slice based on the TMVP enabling flag in PH. In one example, if the TMVP enabling flag in PH is equal to 1, video encoder 200 can signal the TMVP enabling flag in the slice. When the TMVP enabling flag SH is not present, video decoder 300 can be configured to infer the value of TMVP to be equal to the TMVP flag signaled in PH.

[0191] In one example, the above techniques can be implemented as follows:

[0192]

[0193] <add>slice_temporal_mvp_enabled_flag equal to 1 specifies that temporal motion vector predictors can be used for inter prediction. slice_temporal_mvp_enabled_flag equal to 0 specifies that temporal motion vector predictors are disabled in the slice. When slice_temporal_mvp_enabled_flag is not present, it can be inferred to be equal to ph_temporal_mvp_enabled_flag.< / add>

[0194] In another example, the TMVP enable flag can be moved from PH to SH. If TMVP is disabled on slice level, video encoder 200 can be configured to not signal information related to TMVP, such as slice_collocated_from_l0_flag and slice_collocated_ref_idx.

[0195] Example implementations of these techniques are shown below, which are modifications related to VVC Draft 8:

[0196]

[0197]

[0198] In another example, video encoder 200 can be configured to always signal TMVP enable flag in SH, for example, by signaling it in SH, if RPL is signaled in SH, and signal TMVP enable flag in PH, if RPL is signaled in PH.

[0199] Additionally, video encoder 200 can be configured to signal TMVP enable flag in SH or PH in a mutually exclusive manner, such that TMVP enable flag cannot be signaled in both PH and SH.

[0200] When a collocated picture is indicated in PH, for example, by ph_collocated_ref_idx, the collocated picture can be indicated with a different picture size than the current picture. To address this issue, the present disclosure describes techniques for configuring video encoder 200 and video decoder 300 to operate according to the constraint that a collocated picture indicated in PH should have the same picture size as the current picture, not be scaled, or not use different values for the scaling window used to derive scaling ratios for the reference collocated picture and the current picture.

[0201] In an example, the above techniques can be implemented as follows:

[0202] <add>The requirement of bitstream conformance is that the values of pic width in luma samples and pic height in luma samples of the reference picture referred to by ph collocated ref idx shall be equal to the values of pic width in luma samples and pic height in luma samples of the current picture, respectively, and RprConstraintsActive[ph_collocated_from_l0_flag? 0 : 1][ph_collocated_ref_idx] shall be equal to 0.< / add>

[0203] In VVC Draft 8, RPL can be signaled in SPS, PH, or SH. When signaled in PH or SH, RPL can be derived from SPS. In the latter case, a signaling flag rpl_sps_flag[i] is signaled to indicate that RPL is derived from SPS, and rpl_idx[i] is signaled to indicate which RPL of the SPS will be used.

[0204] There can be two lists, RefPicListO and RefPicListl, which utilize this RPL derivation. In the case of RefPicListl, an additional syntax element flag rpl1_idx_present_flag is signaled to indicate whether the derivation of RPL from SPS has occurred. If RPL is not derived from SPS, RPL is explicitly signaled, which can introduce inconsistency in design by treating RefPicListO and RefPicListl signaling differently.

[0205] In one example, video encoder 200 and video decoder 300 can encode rplO_idx_present_flag, which indicates the presence of rpl_sps_flag[i] and rpl_idx[i] for RefPicListO, i.e., provides the same signaling function as RefPicListl.

[0206] In another example, rpl1_idx_present_flag can be replaced by another flag rpl1_present_flag, which indicates whether there is any RefPicListl related syntax element. In this case, the flag can additionally be used as a B slice in a picture.

[0207] In one example, such a flag can be signaled in PPS, and can be implemented as follows:

[0208] <add>rpl1_present_flag equal to 0 specifies that the RefPicListl related syntax elements are not present in the PH syntax structure or slice header of the picture referring to the PPS. rpl1_present_flag equal to 1 specifies that the RefPicListl related syntax elements can be present in the PH syntax structure or slice header of the picture referring to the PPS.< / add>

[0209] Then, video encoder 200 can be configured to conditionally signal, and video decoder 300 is configured to conditionally parse RefPicListl syntax elements based on this flag.

[0210] In one example, video encoder 200 can be configured to conditionally signal the co-located picture flag from RefPicListO in PH based on this flag, and if not present, video decoder 300 can be configured to infer that the co-located picture flag indicates the co-located picture is from RefPicListO. An example implementation of these techniques is shown below, which is a modification on VVC Draft 8:

[0211] <add>if ( rpl1_present_flag )< / add> ph_collocated_from_l0_flag u(1)

[0212] ph collocated from l0 flag equal to 1 specifies that the collocated picture for temporal motion vector prediction is derived from reference picture list 0. ph collocated from l0 flag equal to 0 specifies that the collocated picture for temporal motion vector prediction is derived from reference picture list 1. <add>If not present, ph_collocated_from_l0_flag is inferred to be equal to 1.< / add>

[0213] In another example, video encoder 200 can be configured to conditionally signal zero MVD flag for RefPicListl in PH based on ph collocated from l0 flag. Example implementations of these techniques are shown below, which are modifications related to VVC Draft 8:

[0214] <add>if ( rpl1_present_flag )< / add> mvd_l1_zero_flag u(1)

[0215] In another example, video encoder 200 and video decoder 300 can be configured to operate according to a constraint in which slice type is constrained such that when rpl1_present_flag is equal to 0, a slice can only have I-slice or P-slice type. For example, it can be expressed as follows:

[0216] <add>When rpl1_present_flag is equal to 0, the value of slice type shall be equal to 1 or 2.

[0217] < / add>

[0218] In another example, video encoder 200 can be configured to conditionally signal the weighted prediction table for RefPicListl based on rpl1_present_flag, while video decoder 300 conditionally parses the weighted prediction table for RefPicListl.

[0219] Example implementations of these techniques are shown below, which are modifications related to VVC Draft 8:

[0220] if(wp_info_in_ph_flag&& <add>rpl1_present_flag< / add> ) num_l1_weights ue(v)

[0221] num_l1_weights specifies the number of weights signaled for entries in reference picture list 1. The value of num_l1_weights shall be in the range of 0 to num_ref_entries[1][PicRplsIdx[1]] inclusive.

[0222] The variable NumWeightsL0 is derived as follows:

[0223]

[0224] In another example, video encoder 200 can be configured to conditionally signal RPL syntax elements for RefPicListl based on rpll_present_flag. Example implementations of these techniques are shown below, which are modifications to VVC Draft 8:

[0225]

[0226]

[0227] In this case, the conditional check in ref_pic_list signaling can be simplified.

[0228] In VVC Draft 8, RPL can be signaled for I-slices of non-IDR pictures. However, this RPL is not used in I-slice decoding and there is no way to avoid this redundant signaling. This disclosure describes techniques for configuring video encoder 200 to adjust RPL signaling on slices that are not I-slice type. This disclosure also describes techniques for configuring video encoder 200 to signal RPL to IDR pictures only when it is indicated that there is a RPL for such pictures, e.g., when sps_idr_rpl_present_flag is equal to 1.

[0229] Example implementations of these techniques are shown below, which are modifications to VVC Draft 8:

[0230]

[0231]

[0232] In VVC Draft 8, the reference picture override flag num_ref_idx_active_override_flag is signaled in SH only when RPL is signaled in PH. This can create a problem when RPL is signaled in SH and some pictures in the RPL should be signaled for future reference. Because the override flag is not signaled in SH, the number of active reference pictures cannot be changed. The only way to signal such pictures is to make the entire signaled RPL active, which requires extra overhead in reference index signaling for each block because some pictures included in the RPL are not used. Another problem is that num_ref_idx_active_override_flag is signaled even if I-slices are not needed.

[0233] According to the techniques of this disclosure, video encoder 200 can be configured to signal the override flag num_ref_idx_active_override_flag in the SH even when RPL is signaled in the SH. Additionally or alternatively, video encoder 200 can be configured to not signal (e.g., refrain from signaling) the override flag for I-slices of IDR pictures even when other signaling indicates that RPL is present for such pictures. For example, when sps_idr_rpl_present_flag is equal to 1, there is no need for RPL for I-slice decoding, and thus no reference index signaling in I-slices, and thus no additional overhead.

[0234] An example implementation of these techniques is shown below, which is a modification with respect to VVC Draft 8:

[0235]

[0236] In this example implementation, if signaling indicates that RPL information is present for I-slices, then the syntax element num_ref_idx_active_override_flag is not signaled when RPL information is signaled in the PH (e.g., when rpl_info_in_ph_flag is equal to 1) or RPL is present for IDR pictures (e.g., sps_idr_rpl_present_flag is equal to 1), and the number of RPL entries in any reference picture list is greater than 1. Or in other words, if signaling indicates that RPL is present in the PH (rpl_info_in_ph_flag is equal to 1) or RPL is present for IDR pictures (sps_idr_rpl_present_flag is equal to 1), and the number of RPL entries in any reference picture list is greater than 1, then num_ref_idx_active_override_flag is signaled only for P-slice or B-slice types.

[0237] In VVC Draft 8, there is a flag gdr_or_irap_pic_flag in the PH to indicate whether the picture is IRAP or GDR, which can be used to identify the starting point of decoding. However, the signaling of this flag is not constrained by the indication that mixed NAL types can be present in the picture. In this case, it is not possible to have IRAP or GDR pictures.

[0238] According to the techniques of this disclosure, video encoder 200 can be configured to conditionally signal gdr_or_irap_pic_flag based on an indication of mixed NAL unit types mixed_nalu_types_in_pic_flag. Video decoder 300 can be configured to infer gdr_or_irap_pic_flag equal to 0 when gdr_or_irap_pic_flag is not present in the bitstream.

[0239] Example implementations of these techniques are shown below, which are modifications to VVC Draft 8:

[0240] <add>if ( mixed_nalu_types_in_pic_flag )< / add> gdr_or_irap_pic_flag u(1)

[0241] gdr_or_irap_pic_flag equal to 1 specifies that the current picture is a GDR or IRAP picture. gdr_or_irap_pic_flag equal to 0 specifies that the current picture can or can not be a GDR or IRAP picture. <add>When gdr_irap_pic_flag is not present, it can be inferred to be equal to 0.< / add>

[0242] In some examples, video encoder 200 and video decoder 300 can be configured to operate according to a constraint in which the semantics of gdr_or_irap_pic_flag are constrained such that gdr_or_irap_pic_flag shall be 0 when mixed NAL unit types are indicated to be present. Example implementations of these techniques are shown below, which are modifications to VVC Draft 8:

[0243] gdr_or_irap_pic_flag equal to 1 specifies that the current picture is a GDR or IRAP picture. gdr_or_irap_pic_flag equal to 0 specifies that the current picture can or can not be a GDR or IRAP picture. <add>When mixed_nalu_types_in_pic_flag is equal to 1, gdr_irap_pic_flag shall be equal to 0.< / add>

[0244] According to the techniques described above, video encoder 200 can be configured to determine that reference picture list information is included in a PH syntax structure, and generate a first syntax element, such as the rpl_info_in_ph_flag described above, to indicate that the reference picture list information is included in the PH syntax structure. Video encoder 200 can generate a second syntax element, such as the ph_collocated_from_l0_flag described above, to include in the PH syntax structure, where a first value (e.g., 0) of the second syntax element indicates that a collocated picture for temporal motion vector prediction is to be derived from a first reference picture list (e.g., lo), and a second value (e.g., 1) of the second syntax element indicates that the collocated picture for temporal motion vector prediction is to be derived from a second reference picture list (e.g., li). Video encoder 200 can determine that a slice of video data referring to the PH syntax structure is a P slice, and in response to the slice being a P slice, determine a value of a third syntax element associated with the slice, such as slice_collocated_from_l0_flag, to be equal to a first value of the third syntax element, where the first value (e.g., 0) of the third syntax element indicates that a collocated picture for temporal motion vector prediction is to be derived from a first reference picture list (e.g., lo), and a second value (e.g., 1) of the third syntax element indicates that the collocated picture for temporal motion vector prediction is to be derived from a second reference picture list (e.g., li). Video encoder 200 can output a bitstream of encoded video data that includes the first syntax element and the PH syntax structure.

[0245] According to such techniques, video decoder 300 can be configured to receive a first syntax element, such as rpl_info_in_ph_flag above, indicating that reference picture list information is included in a PH syntax structure, and in response, receive a second syntax element, such as ph_collocated_from_l0_flag in the PH syntax structure above. For the second syntax element, a first value (e.g., 0) can indicate that a collocated picture for temporal motion vector prediction is to be derived from a first reference picture list (e.g., lo), and a second value (e.g., 1) can indicate that a collocated picture for temporal motion vector prediction is to be derived from a second reference picture list (e.g., li). Video decoder 300 can receive a slice of video data referring to the PH syntax structure, and in response to the slice being a P slice, set a value of a third syntax element associated with the slice, such as slice_collocated_from_l0_flag above, to the first value. The first value (e.g., 0) of the third syntax element can indicate that a collocated picture for temporal motion vector prediction is to be derived from the first reference picture list (e.g., lo), while the second value (e.g., 1) of the third syntax element would indicate that a collocated picture for temporal motion vector prediction is to be derived from the second reference picture list (e.g., li).

[0246] Additionally, video decoder 300 can be configured to receive, in response to receiving a first syntax element indicating that reference picture list information is included in a picture header syntax structure and in response to the slice being a P slice, an instance of a fourth syntax element, where a first value of the fourth syntax element indicates that a fifth syntax element is included in a slice header and a second value of the fourth syntax element indicates that the fifth syntax element is not included in the slice header. In response to the instance of the fourth syntax element being equal to the first value of the fourth syntax element, video decoder 300 can receive an instance of the fifth syntax element and determine a number of active reference pictures for the slice based on a value for the instance of the fifth syntax element.

[0247] FIG. 3 A block diagram of an example video encoder 200 that can perform the techniques of this disclosure is shown. The video encoder 200 is provided by way of illustration and should not be considered limiting. For purposes of illustration, the video encoder 200 is described in accordance with the techniques of VVC (ITU-T H.266 under development) and HEVC (ITU-T H.265). However, the techniques of this disclosure can be performed by video encoding devices configured to other video coding standards. FIG. 3 is for purposes of illustration and should not be considered limiting of the techniques of this disclosure. For purposes of illustration, this disclosure describes video encoder 200 in accordance with the techniques of VVC (ITU-T H.266 under development) and HEVC (ITU-T H.265). However, the techniques of this disclosure can be performed by video encoding devices configured to other video coding standards.

[0248] In FIG. 3 In the example of FIG. 2, video encoder 200 includes video data memory 230, mode select unit 202, residual generation unit 204, transform processing unit 206, quantization unit 208, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, filter unit 216, DPB 218, and entropy encoding unit 220. Any or all of video data memory 230, mode select unit 202, residual generation unit 204, transform processing unit 206, quantization unit 208, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, filter unit 216, decoded picture buffer (DPB) 218, and entropy encoding unit 220 can be implemented in one or more processors or in processing circuitry. For instance, the units of video encoder 200 can be implemented as one or more circuits or logic elements as part of hardware circuitry, or as part of a processor, ASIC, FPGA. Also, video encoder 200 can include additional or alternative processors or processing circuitry to perform these and other functions.

[0249] Video data memory 230 can store video data to be encoded by the components of video encoder 200. Video encoder 200 can receive the video data stored in video data memory 230 from, for example, video source 104 FIG. 1 ) that is to be encoded by video encoder 200. DPB 218 can act as a reference picture memory that stores reference video data for use in prediction of subsequent video data by video encoder 200. Video data memory 230 and DPB 218 can be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM) including synchronous DRAM (SDRAM), magneto resistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. Video data memory 230 and DPB 218 can be provided by the same memory device or separate memory devices. In various examples, video data memory 230 can be on-chip with other components of video encoder 200, as illustrated, or off-chip relative to those components.

[0250] In this disclosure, reference to video data memory 230 should not be interpreted as being limited to memory internal to video encoder 200 (unless so specifically described), or memory external to video encoder 200 (unless so specifically described). Rather, reference to video data memory 230 should be understood as reference memory that stores video data that video encoder 200 receives for encoding (e.g., video data of a current block that is to be encoded). FIG. 1 Memory 106 of source device 102 can also provide temporary storage of the outputs from the various units of video encoder 200.

[0251] FIG. 3 The various units of FIG. 2 are illustrated to assist with understanding the operations performed by video encoder 200. The units can be implemented as fixed- function circuits, programmable circuits, or a combination thereof. Fixed-function circuits refer to circuits that provide particular functionality, and are preset on the operations that can be performed. Programmable circuits refer to circuits that can be programmed to perform various tasks and provide flexible functionality in the operations that can be performed. For instance, a programmable circuit can run software or firmware that cause the programmable circuit to operate in the manner defined by instructions of the software or firmware. Fixed-function circuits can run software instructions (e.g., receive parameters or output parameters), but the types of operations that the fixed-function circuits perform are generally immutable. In some examples, one or more of these units can be distinct circuit blocks (fixed-function or programmable), and in some examples one or more of these units can be integrated circuits.

[0252] Video encoder 200 can include arithmetic logic units (ALUs), elementary function units (EFUs), digital circuits, analog circuits, and / or programmable cores formed from programmable circuits. In examples where the operations of video encoder 200 are performed using software run by the programmable circuits, memory 106 FIG. 1 ) can store the instructions (e.g., object code) of the software that video encoder 200 receives and runs, or another memory within video encoder 200 (not shown) can store such instructions.

[0253] Video data memory 230 is configured to store video data to be encoded. Video encoder 200 can retrieve pictures of the video data from video data memory 230 and provide the video data to residual generation unit 204 and mode selection unit 202. Video data in video data memory 230 can be raw video data that is to be encoded.

[0254] Mode selection unit 202 includes motion estimation unit 222, motion compensation unit 224, and intra-prediction unit 226. Mode selection unit 202 can include additional functional units to perform video prediction according to other prediction modes. As examples, mode selection unit 202 can include a palette unit, an intra-block copy unit (which can be part of motion estimation unit 222 and / or motion compensation unit 224), an affine unit, a linear model (LM) unit, etc.

[0255] The mode selection unit 202 generally coordinates multiple encoding passes to test combinations of encoding parameters and the resulting rate-distortion values for such combinations. The encoding parameters can include partitioning of CTUs into CUs, prediction modes for CUs, transform types for residual data of CUs, quantization parameters for residual data of CUs, and so on. The mode selection unit 202 can ultimately select the combination of encoding parameters that has a better rate-distortion value than other tested combinations.

[0256] The video encoder 200 can partition a picture retrieved from the video data memory 230 into a series of CTUs, and encapsulate one or more CTUs within a slice. The mode selection unit 202 can partition the CTUs of the picture according to a tree structure, such as the QTBT structure or the quad-tree structure of HEVC described above. As described above, the video encoder 200 can form one or more CUs by partitioning a CTU according to the tree structure. Such CUs can also be commonly referred to as "video blocks" or "blocks."

[0257] In general, the mode selection unit 202 also controls its components (e.g., the motion estimation unit 222, the motion compensation unit 224, and the intra prediction unit 226) to generate a prediction block for a current block (e.g., a current CU, or an overlapping portion of a PU and a TU in HEVC). For inter prediction of a current block, the motion estimation unit 222 can perform a motion search to identify one or more closely matching reference blocks in one or more reference pictures (e.g., one or more previously coded pictures stored in the DPB 218). Specifically, the motion estimation unit 222 can calculate values representative of how similar a candidate reference block is to the current block, e.g., according to a sum of absolute difference (SAD), a sum of squared difference (SSD), a mean absolute difference (MAD), a mean squared difference (MSD), and so on. The motion estimation unit 222 can generally perform these calculations using sample-by-sample differences between the current block and a considered reference block. The motion estimation unit 222 can identify the reference block with the lowest value resulting from these calculations, which indicates the reference block that most closely matches the current block.

[0258] Motion estimation unit 222 can form one or more motion vectors (MVs) that define a position of a reference block in a reference picture relative to a position of the current block in the current picture. Motion estimation unit 222 can then provide the motion vector(s) to motion compensation unit 224. For example, for single predictive motion compensation, motion estimation unit 222 can provide a single motion vector, while for bi-predictive motion compensation, motion estimation unit 222 can provide two motion vectors. Motion compensation unit 224 can then use the motion vector(s) to generate the predictive block. For example, motion compensation unit 224 can use the motion vector(s) to retrieve data for the reference block. As another example, if the motion vector(s) have fractional sample precision, motion compensation unit 224 can interpolate values for the predictive block according to one or more interpolation filters. Further, for bi-predictive motion compensation, motion compensation unit 224 can retrieve data for two reference blocks identified by the respective motion vectors and combine the retrieved data by, for example, a sample-wise average or weighted average.

[0259] As another example, for intra-prediction, or intra-prediction coding, intra-prediction unit 226 can generate the predictive block from samples neighboring the current block. For example, for directional modes, intra-prediction unit 226 can mathematically combine values of the neighboring samples and fill these computed values on the current block in a defined direction to produce the predictive block. As another example, for a DC mode, intra-prediction unit 226 can compute an average of the neighboring samples of the current block and generate the predictive block to include the produced average for each sample of the predictive block.

[0260] Mode select unit 202 provides the predictive block to residual generation unit 204. Residual generation unit 204 receives the original, unencoded version of the current block from video data memory 230 and the predictive block from mode select unit 202. Residual generation unit 204 computes a sample-wise difference between the current block and the predictive block. The resulting sample-wise difference defines a residual block for the current block. In some examples, residual generation unit 204 can also determine a difference between sample values in the residual block to generate the residual block using residual differential pulse code modulation (RDPCM). In some examples, residual generation unit 204 can be formed using one or more subtractor circuits that perform binary subtraction.

[0261] In examples where the mode selection unit 202 partitions a CU into PUs, each PU can be associated with a luma prediction unit and corresponding chroma prediction units. Video encoder 200 and video decoder 300 can support PUs having various sizes. As described above, the size of a CU can refer to the size of the CU's luma coding block, and the size of a PU can refer to the size of the PU's luma prediction unit. Assuming that a particular CU has a size of 2Nx2N, video encoder 200 can support PU sizes of 2Nx2N or NxN for intra prediction, and 2Nx2N, 2NxN, Nx2N, NxN, or similar symmetric PU sizes for inter prediction. Video encoder 200 and video decoder 300 can also support asymmetric partitioning for inter prediction for PU sizes of 2NxnU, 2NxnD, nLx2N, and nRx2N.

[0262] In examples where the mode selection unit 202 does not further partition a CU into PUs, each CU can be associated with a luma coding block and corresponding chroma coding blocks. As described above, the size of a CU can refer to the size of the CU's luma coding block. Video encoder 200 and video decoder 300 can support CU sizes of 2Nx2N, 2NxN, or Nx2N.

[0263] For other video coding techniques, such as intra block copy mode coding, affine mode coding, and linear model (LM) mode coding, as a few examples, the mode selection unit 202 generates, via a respective unit associated with the coding technique, a prediction block for the current block being encoded. In some examples, such as palette mode coding, the mode selection unit 202 can not generate a prediction block, but rather generate syntax elements that indicate a way to reconstruct the block based on a selected palette. In such modes, the mode selection unit 202 can provide the syntax elements to the entropy encoding unit 220 for encoding.

[0264] As described above, the residual generation unit 204 receives video data for the current block and the corresponding prediction block. The residual generation unit 204 then generates a residual block for the current block. To generate the residual block, the residual generation unit 204 calculates, sample by sample, differences between the prediction block and the current block.

[0265] The transform processing unit 206 applies one or more transforms to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). The transform processing unit 206 can apply various transforms to the residual block to form the transform coefficient block. For example, the transform processing unit 206 can apply a discrete cosine transform (DCT), a directional transform, a Karhunen-Loeve Transform (KLT), or a conceptually similar transform. In some examples, the transform processing unit 206 can perform multiple transforms, e.g., a primary transform and a secondary transform such as a rotational transform, on the residual block. In some examples, the transform processing unit 206 does not apply a transform to the residual block.

[0266] Quantization unit 208 can quantize the transform coefficients in a transform coefficient block, to produce a quantized transform coefficient block. Quantization unit 208 can quantize transform coefficients of a transform coefficient block according to a quantization parameter (QP) value associated with the current block. Video encoder 200 (e.g., via mode select unit 202) can adjust the degree of quantization applied to transform coefficient blocks associated with the current block by adjusting the QP value associated with the CU. Quantization introduces loss of information, therefore quantized transform coefficients can have lower precision than the original transform coefficients produced by transform processing unit 206.

[0267] Inverse quantization unit 210 and inverse transform processing unit 212 can apply inverse quantization and inverse transform, respectively, to a quantized transform coefficient block to reconstruct a residual block from the transform coefficient block. Reconstruction unit 214 can generate a reconstructed block corresponding to the current block (albeit with some degree of distortion) based on the reconstructed residual block and the prediction block generated by mode select unit 202. For example, reconstruction unit 214 can add samples of the reconstructed residual block to corresponding samples from the prediction block generated by mode select unit 202 to produce the reconstructed block.

[0268] Filter unit 216 can perform one or more filtering operations on the reconstructed block. For example, filter unit 216 can perform a deblocking operation to reduce blocking artifacts along CU boundaries. In some examples, the operations of filter unit 216 can be skipped.

[0269] Video encoder 200 stores the reconstructed block in DPB 218. For example, in examples in which the operations of filter unit 216 are not performed, reconstruction unit 214 can store the reconstructed block to DPB 218. In examples in which the operations of filter unit 216 are performed, filter unit 216 can store the filtered reconstructed block to DPB 218. Motion estimation unit 222 and motion compensation unit 224 can retrieve reference pictures from DPB 218 that are formed of reconstructed (and possibly filtered) blocks to inter-predict blocks of subsequently encoded pictures. Moreover, intra-prediction unit 226 can use reconstructed blocks in DPB 218 of the current picture to intra-predict other blocks in the current picture.

[0270] In general, entropy encoding unit 220 can entropy encode syntax elements received from other functional components of video encoder 200. For example, entropy encoding unit 220 can entropy encode quantized transform coefficient blocks from quantization unit 208. As another example, entropy encoding unit 220 can entropy encode prediction syntax elements (e.g., motion information for inter prediction or intra-mode information for intra prediction) from mode select unit 202. Entropy encoding unit 220 can perform one or more entropy encoding operations on syntax elements, as another example of video data, to generate entropy encoded data. For example, entropy encoding unit 220 can perform a context- adaptive variable length coding (CAVLC) operation, a CABAC operation, a variable- to-variable (V2V) coding operation, a syntax-based context-adaptive binary arithmetic coding (SBAC) operation, a Probability Interval Partitioning Entropy (PIPE) coding operation, an Exponential-Golomb coding operation, or another type of entropy coding operation on the data. In some examples, entropy encoding unit 220 can operate in a bypass mode in which syntax elements are not entropy encoded.

[0271] Video encoder 200 can output a bitstream that includes the entropy encoded syntax elements needed to reconstruct blocks of a slice or picture. Specifically, entropy encoding unit 220 can output the bitstream.

[0272] The operations described above are described with respect to blocks. Such description should be understood to be operations for luma coding blocks and / or chroma coding blocks. As described above, in some examples, the luma coding blocks and the chroma coding blocks are luma and chroma components of a CU. In some examples, the luma coding blocks and the chroma coding blocks are luma and chroma components of a PU.

[0273] In some examples, operations performed with respect to luma coding blocks need not be repeated for chroma coding blocks. As one example, operations to identify a motion vector (MV) and a reference picture for a luma coding block need not be repeated to identify an MV and a reference picture for a chroma coding block. Rather, the MV of the luma coding block can be scaled to determine the MV of the chroma block, and the reference picture can be the same. As another example, intra prediction processing can be the same for luma coding blocks and chroma coding blocks.

[0274] Video encoder 200 represents an example of a device configured to encode video data, including a memory configured to store video data, and one or more processing units implemented in circuitry and configured to perform the techniques of the present disclosure, including the techniques in the claims section below.

[0275] FIG. 4 A block diagram of an example video decoder 300 that can perform the techniques of the present disclosure is shown. FIG. 4 The techniques are provided for purposes of illustration and description, and are not limiting of the technology in the present disclosure. For illustration, the present disclosure describes a video decoder 300 according to the techniques of VVC (ITU-T H.266 under development) and HEVC (ITU-T H.265). However, the technology of the present disclosure can be performed by video coding devices configured to other video coding standards.

[0276] In FIG. 4 In the example of FIG. 3, video decoder 300 includes coded picture buffer (CPB) memory 320, entropy decoding unit 302, prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, filter unit 312, and decoded picture buffer (DPB) 314. Any or all of CPB memory 320, entropy decoding unit 302, prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, filter unit 312, and DPB 314 can be implemented in one or more processors or in processing circuitry. For instance, the units of video decoder 300 can be implemented as one or more circuits or logic elements as part of hardware circuitry, or as part of a processor, ASIC, FPGA. Also, video decoder 300 can include additional or alternative processors or processing circuitry to perform these and other functions.

[0277] Prediction processing unit 304 includes motion compensation unit 316 and intra-prediction unit 318. Prediction processing unit 304 can include additional units to perform prediction from other prediction modes. As examples, prediction processing unit 304 can include a palette unit, an intra-block copy unit (which can form a part of motion compensation unit 316), an affine unit, a linear model (LM) unit, etc. In other examples, video decoder 300 can include more, less, or different functional components.

[0278] CPB memory 320 can store video data, such as an encoded video bitstream, to be decoded by the components of video decoder 300. The video data stored in CPB memory 320 can be obtained, for example, from computer- readable medium 110 (FIG. 1). In some examples, CPB memory 320 can store data for multiple video bitstreams to be decoded by video decoder 300. FIG. 1 ) is obtained. CPB memory 320 can include a CPB that stores encoded video data, e.g., syntax elements, from an encoded video bitstream. Moreover, CPB memory 320 can store video data other than syntax elements of a coded picture, such as temporary data representing outputs from the various units of video decoder 300. DPB 314 generally stores decoded pictures, which video decoder 300 can output and / or use as reference video data when decoding subsequent data or pictures of the encoded video bitstream. CPB memory 320 and DPB 314 can be formed by any of a variety of memory devices, such as DRAM, including SDRAM, MRAM, RRAM, or other types of memory devices. CPB memory 320 and DPB 314 can be provided by the same memory device or separate memory devices. In various examples, CPB memory 320 can be on-chip with other components of video decoder 300, or off-chip relative to those components.

[0279] Additionally or alternatively, in some examples, video decoder 300 can retrieve coded video data from memory 120 FIG. 1 ) as described above with CPB memory 320. Also, in examples where some or all of the functionality of video decoder 300 is implemented in software executed by the processing circuitry of video decoder 300, memory 120 can store the instructions to be executed by video decoder 300.

[0280] FIG. 4 The various units illustrated in FIG. 3 are illustrated to assist with understanding the operations performed by video decoder 300. The units can be implemented as fixed- function circuits, programmable circuits, or a combination thereof. Similar to FIG. 3 , fixed-function circuits refer to circuits that provide particular functionality and are preset to operate on only the data of particular types. Programmable circuits refer to circuits that can be programmed to perform various tasks and provide flexible functionality in the operations that can be performed. For instance, a programmable circuit can execute software or firmware that cause the programmable circuit to operate in ways defined by instructions of the software or firmware. Fixed-function circuits can execute software instructions (e.g., receive parameters or output parameters), but the types of operations that the fixed-function circuits perform are generally immutable. In some examples, one or more of the units can be distinct circuit blocks (fixed-function or programmable) and, in some examples, one or more of the units can be integrated circuits.

[0281] Video decoder 300 can include ALUs, EFUs, digital circuits, analog circuits, and / or programmable cores formed from programmable circuitry. In examples where the operations of video decoder 300 are performed by software running on the programmable circuitry, on-chip or off-chip memory can store instructions (e.g., object code) of the software that video decoder 300 receives and executes.

[0282] Entropy decoding unit 302 can receive encoded video data from the CPB and entropy decode the video data to reproduce syntax elements. Prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, and filter unit 312 can generate decoded video data based on the syntax elements extracted from the bitstream.

[0283] In general, video decoder 300 reconstructs a picture on a block-by-block basis. Video decoder 300 can perform reconstruction operations separately on each block (where the block that is currently being reconstructed (i.e., decoded) can be referred to as the "current block").

[0284] Entropy decoding unit 302 can entropy decode syntax elements defining quantized transform coefficients of a quantized transform coefficient block, as well as transform information such as a quantization parameter (QP) and / or transform mode indication(s). Inverse quantization unit 306 can use the QP associated with the quantized transform coefficient block to determine a degree of quantization and, likewise, a degree of inverse quantization for inverse quantization unit 306 to apply. Inverse quantization unit 306 may, for example, perform a bit- shift operation to the right by the QP to inverse quantize the quantized transform coefficients. Inverse quantization unit 306 may

[0285] After inverse quantization unit 306 forms the transform coefficient block, inverse transform processing unit 308 can apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, inverse transform processing unit 308 can apply an inverse DCT, an inverse integer transform, an inverse Karhunen-Loeve Transform (KLT), an inverse rotational transform, an inverse directional transform, or another inverse transform to the transform coefficient block.

[0286] Further, prediction processing unit 304 generates the prediction block from the prediction information syntax elements entropy decoded by entropy decoding unit 302. For example, if the prediction information syntax elements indicate that the current block is inter-predicted, motion compensation unit 316 can generate the prediction block. In this case, the prediction information syntax elements can indicate a reference picture in the DPB 314 from which to retrieve a reference block, and a motion vector identifying a location of the reference block in the reference picture relative to a location of the current block in the current picture. Motion compensation unit 316 generally can perform inter-prediction processing in a manner generally similar to that described above with respect to motion compensation unit 224. FIG. 3 ) described above with respect to motion compensation unit 224.

[0287] As another example, if the prediction information syntax element indicates that the current block is intra-predicted, the intra-prediction unit 318 can generate the prediction block according to an intra-prediction mode indicated by the prediction information syntax element. Again, the intra-prediction unit 318 can generally perform intra-prediction processing in a manner generally similar to that described with respect to the intra-prediction unit 226( FIG. 3 ) of the video encoder 200.

[0288] The reconstruction unit 310 can reconstruct the current block using the prediction block and the residual block. For example, the reconstruction unit 310 can add the samples of the residual block to corresponding samples of the prediction block to reconstruct the current block.

[0289] The filter unit 312 can perform one or more filtering operations on the reconstructed block. For example, the filter unit 312 can perform a deblocking operation to reduce blocking artifacts along the edges of the reconstructed block. The operations of the filter unit 312 are not necessarily performed in all examples.

[0290] The video decoder 300 can store the reconstructed block in the DPB 314. For example, in examples in which the operations of the filter unit 312 are not performed, the reconstruction unit 310 can store the reconstructed block to the DPB 314. In examples in which the operations of the filter unit 312 are performed, the filter unit 312 can store the filtered reconstructed block to the DPB 314. As described above, the DPB 314 can provide reference information to the prediction processing unit 304, such as samples of the current picture for intra-prediction and previously decoded pictures for subsequent motion compensation. In addition, the video decoder 300 can output decoded pictures (e.g., decoded video) from the DPB 314 for subsequent presentation on a display device, such as the display device 118 of FIG. 1. FIG. 1

[0291] In this way, the video decoder 300 represents an example of a video decoding apparatus that includes a memory configured to store video data, and one or more processing units implemented in circuitry and configured to perform the techniques of this disclosure, including the techniques of the following claims section.

[0292] FIG. 5 A flowchart of an example method for encoding a current block is shown. The current block can include a current CU. Although described with respect to the video encoder 200( FIG. 1 and 3 ), it should be understood that other devices can be configured to perform similar methods. FIG. 5

[0293] ​​In this example, video encoder 200 initially predicts the current block (350). For example, video encoder 200 can form a prediction block for the current block. Video encoder 200 can then calculate a residual block for the current block (352). To calculate the residual block, video encoder 200 can calculate the difference between the original, unencoded block and the prediction block for the current block. Video encoder 200 can then transform the residual block and quantize the transform coefficients of the residual block (354). Next, video encoder 200 can scan the quantized transform coefficients of the residual block (356). During the scan, or after the scan, video encoder 200 can entropy encode the transform coefficients (358). For example, video encoder 200 can encode the transform coefficients using CAVLC or CABAC. Video encoder 200 can then output the entropy encoded data for the block (360).

[0294] FIG. 6 is a flowchart illustrating an example method for decoding a current block of video data. The current block can include a current CU. Although video decoder 300 is described, it should be understood that other devices can be configured to perform similar methods as FIG. 1 and 4 are described, it should be understood that other devices can be configured to perform similar methods as FIG. 6

[0295] Video decoder 300 can receive entropy encoded data for the current block (370), such as entropy encoded prediction information and entropy encoded data for transform coefficients of a residual block corresponding to the current block. Video decoder 300 can entropy decode the entropy encoded data to determine prediction information for the current block and reproduce transform coefficients of the residual block (372). Video decoder 300 can predict the current block (374), e.g., using an intra or inter prediction mode indicated by the prediction information for the current block, to calculate a prediction block for the current block. Video decoder 300 can then inverse scan the reproduced transform coefficients (376) to create a block of quantized transform coefficients. Video decoder 300 can then inverse quantize and inverse transform the transform coefficients to produce a residual block (378). Video decoder 300 can finally decode the current block by combining the prediction block and the residual block (380).

[0296] FIG. 7 is a flowchart illustrating an example method for encoding a current block. The current block can include a current CU. Although video encoder 200 is described, it should be understood that other devices can be configured to perform similar methods as FIG. 1 and 3 are described, it should be understood that other devices can be configured to perform similar methods as FIG. 7

[0297] In this example, video encoder 200 initially predicts the current block (350). For example, video encoder 200 can form a prediction block for the current block. Video encoder 200 can then calculate a residual block for the current block (352). To calculate the residual block, video encoder 200 can calculate the difference between the original, unencoded block and the prediction block for the current block. Video encoder 200 can then transform the residual block and quantize the transform coefficients of the residual block (354). Next, video encoder 200 can scan the quantized transform coefficients of the residual block (356). During the scan, or after the scan, video encoder 200 can entropy encode the transform coefficients (358). For example, video encoder 200 can encode the transform coefficients using CAVLC or CABAC. Video encoder 200 can then output the entropy encoded data for the block (360). FIG. 7 ​​In the example, in response to determining that the reference image list information is included in the PH syntax structure, the video encoder 200 generates a first syntax element, such as rpl_info_in_ph_flag above, indicating that the reference image list information is included in the PH syntax structure (400).

[0298] Video encoder 200 generates a second syntax element, such as ph_collocated_from_l0_flag, included in the PH syntax structure. This second syntax element is set to a value indicating that the co-located image used for temporal motion vector prediction will be derived from a second list of reference images (e.g., l1) of the two reference image lists (402). Video encoder 200 determines that the video data stripe referring to the PH syntax structure is a P-strip (404). In response to the stripe being a P-stripe, video encoder 200 determines the value of a third syntax element associated with the stripe, such as slice_collocated_from_l0_flag, equal to the value indicating that the co-located image used for temporal motion vector prediction will be derived from a first list of reference images (e.g., l0) of the two reference image lists (406). Video encoder 200 outputs a bitstream of encoded video data including the first syntax element and the PH syntax structure (408).

[0299] FIG. 8 This is a flowchart illustrating an example method for decoding the current block of video data. The current block may include the current CU. Although for video decoder 300 ( FIG. 1 and 4 This has been described, but it should be understood that other devices can be configured to perform the same actions. FIG. 8 A similar approach.

[0300] exist FIG. 8 In the example, in response to receiving a first syntax element, such as rpl_info_in_ph_flag, indicating that reference picture list information is included in the PH syntax structure, the video decoder 300 receives a second syntax element, such as ph_collocated_from_l0_flag in the PH syntax structure (410). The video decoder 300 determines, based on the value of the second syntax element, that the co-located picture for temporal motion vector prediction will be derived from the second reference picture list (e.g., l1) (412). The video decoder 300 receives a slice of video data referring to the PH syntax structure (414). In response to the slice being a P slice, the video decoder 300 sets the value of a third syntax element associated with the slice, such as slice_collocated_from_l0_flag, equal to the value indicating that the co-located picture for temporal motion vector prediction will be derived from the second reference picture list (416).

[0301] Video decoder 300 can decode the slice of video data based on the value of the third syntax element. For example, for a block of the slice, video decoder 300 can determine a temporal motion vector candidate to be included in a motion vector candidate list. To determine the temporal motion vector candidate to be included in the motion vector candidate list, video decoder 300 can identify a collocated picture for temporal motion vector prediction from the first reference picture list, identify a collocated block in the collocated picture, and derive the temporal motion vector candidate based on a motion vector used to decode the collocated block.

[0302] Video decoder 300 can output decoded video data that includes a decoded version of the slice. For example, video decoder 300 can output a decoded picture that includes the decoded slice for display or storage.

[0303] The following clauses represent example implementations of the described techniques and apparatuses.

[0304] Clause 1 : A method of decoding video data, the method comprising: receiving a first syntax element; in response to the first syntax element indicating that reference picture list information is included in a picture header syntax structure, receiving a second syntax element in the picture header syntax structure, wherein a first value for the second syntax element indicates that a collocated picture used for temporal motion vector prediction is to be derived from a first reference picture list and a second value for the second syntax element indicates that the collocated picture used for temporal motion vector prediction is to be derived from a second reference picture list; receiving a slice of video data referring to the picture header syntax structure; and in response to the slice being a P slice, setting a value for a third syntax element associated with the slice to the first value for the third syntax element, wherein the first value for the third syntax element indicates that the collocated picture used for temporal motion vector prediction is to be derived from the first reference picture list and a second value for the third syntax element indicates that the collocated picture used for temporal motion vector prediction is to be derived from the second reference picture list.

[0305] Clause 2: The method of clause 1, wherein the value for the second syntax element is equal to the second value for the second syntax element.

[0306] Clause 3: The method of clause 1 or clause 2, wherein setting the value for the third syntax element associated with the slice to the second value comprises inferring a value for a second flag to be the second value.

[0307] Clause 4: The method of any of clauses 1-3, wherein setting the value for the third syntax element associated with the slice to the second value comprises setting the value for the third syntax element to the second value without receiving an instance of the third syntax in a slice header of the slice.

[0308] Clause 5. The method of any of clauses 1-4, further comprising, for a block of the slice, determining a temporal motion vector candidate for inclusion in a motion vector candidate list, wherein determining the temporal motion vector candidate comprises: identifying, from a first reference picture list, a collocated picture used as temporal motion vector prediction; identifying, in the collocated picture, a collocated block; and deriving the temporal motion vector candidate based on a motion vector used to decode the collocated block.

[0309] Clause 6. The method of any of clauses 1-5, further comprising, in response to receiving a first syntax element indicating that reference picture list information is included in a picture header syntax structure, and in response to the slice being a P slice, receiving an instance of a fourth syntax element, wherein a first value for the fourth syntax element indicates that a fifth syntax element is included in the slice header, and a second value for the fourth syntax element indicates that the fifth syntax element is not included in the slice header; in response to the instance of the fourth syntax element being equal to the first value for the fourth syntax element, receiving an instance of the fifth syntax element; and determining a number of active reference pictures for the slice based on a value for the instance of the fifth syntax element.

[0310] Clause 7. The method of clause 6, wherein the slice is a first slice, the instance of the fourth syntax element is a first instance of the fourth syntax element, and the instance of the fifth syntax element is a first instance of the fifth syntax element, the method further comprising: receiving a second slice of video data referring to the picture header syntax structure; and in response to receiving the first syntax element indicating that reference picture list information is included in the picture header syntax structure, and in response to the second slice being an I slice, receiving a slice header for the second slice without a second instance of the fourth syntax element.

[0311] Clause 8: A device for decoding video data, comprising: a memory configured to store video data; and one or more processors implemented in circuitry and configured to: receive a first syntax element; responsive to the first syntax element indicating that reference picture list information is included in a picture header syntax structure, receive a second syntax element in the picture header syntax structure, wherein a first value for the second syntax element indicates that a collocated picture used for temporal motion vector prediction is to be derived from a first reference picture list and a second value for the second syntax element indicates that the collocated picture used for temporal motion vector prediction is to be derived from a second reference picture list; receive a slice of video data referring to the picture header syntax structure; and responsive to the slice being a P slice, set a value for a third syntax element associated with the slice to the first value for the third syntax element, wherein the first value for the third syntax element indicates that the collocated picture used for temporal motion vector prediction is to be derived from the first reference picture list and a second value for the third syntax element indicates that the collocated picture used for temporal motion vector prediction is to be derived from the second reference picture list.

[0312] Clause 9: The device of clause 8, wherein the value for the second syntax element is equal to the second value for the second syntax element.

[0313] Clause 10: The device of clause 8 or 9, wherein to set the value for the third syntax element associated with the slice to the second value, the one or more processors are further configured to infer the value for the third syntax element to be the second value.

[0314] Clause 11: The device of any of clauses 8-10, wherein to set the value for the third syntax element associated with the slice to the second value comprises the one or more processors being further configured to set the value for the third syntax element to the second value without receiving an instance of the third syntax in a slice header of the slice.

[0315] Clause 12: The device of any of clauses 8-11, wherein the one or more processors are further configured to, for a block of the slice, determine a temporal motion vector candidate for inclusion in a motion vector candidate list, wherein to determine the temporal motion vector candidate, the one or more processors are further configured to: identify a collocated picture used for temporal motion vector prediction from the first reference picture list; identify a collocated block in the collocated picture; and derive the temporal motion vector candidate based on a motion vector used to decode the collocated block.

[0316] Clause 13: The device of any of clauses 8-12, wherein the one or more processors are further configured to: in response to receiving the first syntax element indicating that reference picture list information is included in the picture header syntax structure, and in response to the slice being a P slice, receive an instance of a fourth syntax element, wherein a first value for the fourth syntax element indicates that a fifth syntax element is included in the slice header, and a second value for the fourth syntax element indicates that the fifth syntax element is not included in the slice header; in response to the instance of the fourth syntax element being equal to the first value for the fourth syntax element, receive an instance of the fifth syntax element; and determine a number of active reference pictures for the slice based on a value for the instance of the fifth syntax element.

[0317] Clause 14: The device of clause 13, wherein the slice is a first slice, the instance of the fourth syntax element is a first instance of the fourth syntax element, and the instance of the fifth syntax element is a first instance of the fifth syntax element, wherein the one or more processors are further configured to: receive a second slice of video data referring to the picture header syntax structure; in response to receiving the first syntax element indicating that reference picture list information is included in the picture header syntax structure, and in response to the second slice being an I slice, receive a slice header for the second slice without a second instance of the fourth syntax element.

[0318] Clause 15: The device of any of clauses 8-14, wherein the apparatus comprises a wireless communication device, further comprising a receiver configured to receive the encoded video data.

[0319] Clause 16: The device of clause 15, wherein the wireless communication device comprises a telephone handset, and wherein the receiver is configured to demodulate a signal comprising the encoded video data according to a wireless communication standard.

[0320] Clause 17: The device of any of clauses 8-16, further comprising a display configured to display the decoded video data.

[0321] Clause 18: The device of any of clauses 8-17, wherein the apparatus comprises one or more of a video camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

[0322] Clause 19: A method of encoding video data, comprising: in response to determining that reference picture list information is included in a picture header syntax structure, generating a first syntax element indicating that the reference picture list information is included in the picture header syntax structure; generating a second syntax element for inclusion in the picture header syntax structure, wherein a first value for the second syntax element indicates that a collocated picture used for temporal motion vector prediction is to be derived from a first reference picture list and a second value for the second syntax element indicates that the collocated picture used for temporal motion vector prediction is to be derived from a second reference picture list; determining that a slice of video data referring to the picture header syntax structure is a P slice; in response to the slice being a P slice, determining that a value for a third syntax element associated with the slice is equal to a first value for the third syntax element, wherein the first value for the third syntax element indicates that the collocated picture used for temporal motion vector prediction is to be derived from the first reference picture list and a second value for the third syntax element indicates that the collocated picture used for temporal motion vector prediction is to be derived from the second reference picture list; and outputting a bitstream of encoded video data, wherein the encoded video data includes the first syntax element and the picture header syntax structure.

[0323] Clause 20: The method of clause 19, wherein the value for the second syntax element is equal to the second value for the second syntax element.

[0324] Clause 21 : The method of clause 19 or 20, further comprising generating a slice header for the slice for inclusion in the bitstream of encoded video data, without including an instance of the third syntax element in the slice header.

[0325] Clause 22: The method of any of clauses 19-21, further comprising, for a block of the slice, determining a temporal motion vector candidate for inclusion in a motion vector candidate list, wherein determining the temporal motion vector candidate comprises: identifying a collocated picture used for temporal motion vector prediction from the first reference picture list; identifying a collocated block in the collocated picture; and deriving the temporal motion vector candidate based on a motion vector used to decode the collocated block.

[0326] Clause 23: The method of any of clauses 19-22, further comprising, in response to determining that reference picture list information is included in the picture header syntax structure and in response to determining that the slice is a P slice, generating an instance of a fourth syntax element, wherein a first value for the fourth syntax element indicates that a fifth syntax element is included in the slice header and a second value for the fourth syntax element indicates that the fifth syntax element is not included in the slice header; determining a number of active reference pictures for the slice; and in response to the instance of the fourth syntax element being equal to the first value for the fourth syntax element, generating an instance of the fifth syntax element for inclusion in the bitstream of encoded video data, wherein a value for the instance of the fifth syntax element indicates the number of active reference pictures for the slice.

[0327] Clause 24: The method of clause 23, wherein the slice is a first slice, the instance of the fourth syntax element is a first instance of the fourth syntax element, and the instance of the fifth syntax element is a first instance of the fifth syntax element, the method further comprising, for a second slice of the video data referring to the picture header syntax structure, in response to determining that reference picture list information is included in the picture header syntax structure and in response to the second slice being an I slice, generating the slice header for inclusion in the bitstream of encoded video data without a second instance of the fourth syntax element.

[0328] Clause 25: A device for encoding video data, comprising: a memory configured to store video data; and one or more processors implemented in circuitry and configured to: in response to determining that reference picture list information is included in a picture header syntax structure, generate a first syntax element indicating that the reference picture list information is included in the picture header syntax structure; generate a second syntax element for inclusion in the picture header syntax structure, wherein a first value for the second syntax element indicates that a collocated picture used for temporal motion vector prediction is to be derived from a first reference picture list and a second value for the second syntax element indicates that the collocated picture used for temporal motion vector prediction is to be derived from a second reference picture list; determine that a slice of video data referring to the picture header syntax structure is a P slice; in response to the slice being a P slice, determine that a value for a third syntax element associated with the slice is equal to a first value for the third syntax element, wherein the first value for the third syntax element indicates that the collocated picture used for temporal motion vector prediction is to be derived from the first reference picture list and a second value for the third syntax element indicates that the collocated picture used for temporal motion vector prediction is to be derived from the second reference picture list; and output a bitstream of encoded video data, wherein the encoded video data includes the first syntax element and the picture header syntax structure.

[0329] Clause 26: The device of clause 25, wherein the value for the second syntax element is equal to a second value for the second syntax element.

[0330] Clause 27: The device of clause 25 or 26, further comprising generating, for inclusion in a bitstream of encoded video data, a slice header for the slice, the slice header not including an instance of the third syntax element.

[0331] Clause 28: The device of any of clauses 25-27, further comprising determining, for a block of the slice, a temporal motion vector candidate for inclusion in a motion vector candidate list, wherein determining the temporal motion vector candidate comprises identifying, from the first reference picture list, a collocated picture used as temporal motion vector prediction, identifying, in the collocated picture, a collocated block, and deriving the temporal motion vector candidate based on a motion vector used to decode the collocated block.

[0332] Clause 29: The device of any of clauses 25-28, further comprising, in response to determining that reference picture list information is included in the picture header syntax structure and in response to determining that the slice is a P slice, generating an instance of a fourth syntax element, wherein a first value for the fourth syntax element indicates that a fifth syntax element is included in the slice header and a second value for the fourth syntax element indicates that the fifth syntax element is not included in the slice header, determining a number of active reference pictures for the slice, and in response to the instance of the fourth syntax element being equal to the first value for the fourth syntax element, generating, for inclusion in a bitstream of encoded video data, an instance of the fifth syntax element, wherein a value for the instance of the fifth syntax element indicates the number of active reference pictures for the slice.

[0333] Clause 30: The device of clause 29, wherein the slice is a first slice, the instance of the fourth syntax element is a first instance of the fourth syntax element, and the instance of the fifth syntax element is a first instance of the fifth syntax element, the method further comprising, for a second slice of video data referring to the picture header syntax structure, in response to determining that reference picture list information is included in the picture header syntax structure and in response to the second slice being an I slice, generating, without a second instance of the fourth syntax element, a slice header for inclusion in a bitstream of encoded video data.

[0334] Clause 31 : The device of any of clauses 25-30, wherein the device comprises a wireless communication device, further comprising a transmitter configured to transmit the encoded video data.

[0335] Clause 32: The device of clause 31, wherein the wireless communication device comprises a telephone handset, and wherein the transmitter is configured to modulate a signal comprising the encoded video data according to a wireless communication standard.

[0336] Clause 33: The device of any of clauses 25-32, further comprising a video camera configured to capture the video data.

[0337] Clause 34: The device of any of clauses 25-33, wherein the device comprises one or more of a video camera, a computer, or a mobile device.

[0338] Clause 35: A method of decoding video data, the method comprising: receiving a first flag; receiving a second flag; determining, based on the first and second flags, whether reference picture lists are signaled in a syntax structure of the video data.

[0339] Clause 36: The method of clause 35, wherein the first flag indicates whether inter- frame slices are allowed for the picture.

[0340] Clause 37: The method of clause 35 or 36, wherein the second flag indicates whether reference picture list syntax elements are present in a slice header of an instantaneous decoder refresh picture.

[0341] Clause 38: The method of any combination of clauses 35-37, wherein determining, based on the first and second flags, whether reference picture lists are signaled in a data structure of the video data comprises: determining that reference picture lists are signaled in the data structure of the video data in response to at least one of the first flag or the second flag being true.

[0342] Clause 39: The method of any combination of clauses 35-37, further comprising: receiving a third flag that indicates whether reference picture list information is present in a picture header syntax structure or a slice header.

[0343] Clause 40: The method of clause 39, wherein determining, based on the first and second flags, whether reference picture lists are signaled in a data structure of the video data comprises: determining that reference picture lists are signaled in the data structure of the video data in response to (1) at least one of the first flag or the second flag being true and (2) the third flag being true.

[0344] Clause 41: The method of any combination of clauses 35-40, wherein the syntax structure comprises a picture header syntax structure.

[0345] Clause 42: A method of coding video data, the method comprising: receiving a first flag; receiving a second flag; determining, based on the first and second flags, whether a weighted prediction table is signaled in a syntax structure of the video data.

[0346] Clause 43: The method of clause 42, wherein the first flag indicates whether inter- frame slices are allowed for the picture.

[0347] Clause 44: The method of clauses 42 or 43, wherein the second flag indicates whether a reference picture list syntax element is present in a slice header of an instant decoder refresh picture.

[0348] Clause 45: The method of any combination of clauses 42-44, wherein determining, based on the first and second flags, whether to signal a weighting prediction table in a data structure of the video data comprises determining to signal the weighting prediction table in the data structure of the video data in response to at least one of the first flag or the second flag being true.

[0349] Clause 46: The method of any combination of clauses 42-45, further comprising receiving a third flag that indicates whether to apply weighted prediction.

[0350] Clause 47: The method of clause 46, wherein determining, based on the first and second flags, whether to signal a weighting prediction table in a data structure of the video data comprises determining to signal a reference picture list in the data structure of the video data in response to (1) at least one of the first flag or the second flag being true and (2) the third flag being true.

[0351] Clause 48: The method of any combination of clauses 42-47, wherein the syntax structure comprises a picture header syntax structure.

[0352] Clause 49: A device for coding video data, the device comprising one or more means for performing the method of any combination of clauses 35-48.

[0353] Clause 50: The device of clause 49, wherein the one or more means are included in one or more processors implemented in circuitry.

[0354] Clause 51: The device of any of clauses 49-50, further comprising a memory for storing video data.

[0355] Clause 52: The device of any of clauses 49-51, further comprising a display configured to display decoded video data.

[0356] Clause 53: The device of any of clauses 49-52, wherein the device comprises one or more of a video camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

[0357] Clause 54: The device of any of clauses 49-53, wherein the device comprises a video decoder.

[0358] Clause 55: The device of any of clauses 49-54, wherein the device comprises a video encoder.

[0359] Clause 56: A computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to perform the method of any of clauses 35-48.

[0360] It should be appreciated that, in accordance with examples, certain actions or events of any of the techniques described herein can be performed in a different order, can be added, merged, or omitted altogether (e.g., not all described actions or events are required to practice a technique). Further, in some examples, actions or events can be performed concurrently, e.g., through multi-threaded processing, interrupt processing, or multiple processors, rather than sequentially.

[0361] In one or more examples, the functions described can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored on or transmitted over as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media can include computer-readable storage media, which corresponds to a tangible medium such as data storage media, or communication media including any medium that facilitates transfer of a computer program from one place to another, e.g., according to a communication protocol. In this manner, computer- readable media generally can correspond to (1) tangible computer-readable storage media which is non-transitory or (2) a communication medium such as a signal or carrier wave. Data storage media can be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementation of the techniques described in this disclosure. A computer program product can include a computer-readable medium.

[0362] By way of example, and not limitation, such computer-readable storage media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other storage medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any

[0363] Instructions can be executed by one or more processors, such as one or more (DSPs), general purpose microprocessors, ASICs, FPGAs, or other equivalent integrated or discrete logic circuitry. Accordingly, the terms "processor" and "processing circuitry," as used herein can refer to any of the foregoing structures or any other structure suitable for implementation of the techniques described herein. In addition, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec. Also, the techniques could be fully implemented in one or more circuits or logic elements.

[0364] The techniques of this disclosure can be implemented in a variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC), or a set of one or more integrated circuits (e.g., a chip set). Various components, modules, or units described herein can be implemented as hardware, software, or a combination thereof. In some aspects, the various components, modules, or units described herein can be implemented as software modules, hardware modules, or any combination thereof. In some aspects, the various components, modules, or units described herein can be implemented as hardware modules.

[0365] Various examples have been described. These and other examples are within the scope of the following claims.< / add> < / add>

Claims

1. A method of decoding video data, the method comprising: receiving a first syntax element that indicates that reference picture list information is included in a first picture header syntax structure for a first slice and in a second picture header syntax structure for a second slice; in response to the first syntax element, receiving a second syntax element in the first picture header syntax structure, wherein a first value for the second syntax element indicates that a collocated picture used for temporal motion vector prediction is to be derived from a first reference picture list and a second value for the second syntax element indicates that the collocated picture used for temporal motion vector prediction is to be derived from a second reference picture list; in response to the first slice being a P slice, setting a value for a first instance of a third syntax element associated with the first slice to a first value for the third syntax element, wherein the first value for the third syntax element indicates that the collocated picture used for temporal motion vector prediction is to be derived from the first reference picture list and a second value for the third syntax element indicates that the collocated picture used for the temporal motion vector prediction is to be derived from the second reference picture list; after receiving the first syntax element that indicates that reference picture list information is included in the first picture header syntax structure and in response to the first slice being a P slice, receiving a fourth syntax element, wherein a first value for the fourth syntax element indicates that a fifth syntax element is included in a first slice header for the first slice and a second value for the fourth syntax element indicates that the fifth syntax element is not included in the first slice header; in response to the fourth syntax element being equal to the first value for the fourth syntax element, receiving the fifth syntax element; determining a number of active reference pictures for the first slice based on a value for an instance of the fifth syntax element; and in response to receiving the first syntax element that indicates that reference picture list information is included in the second picture header syntax structure and in response to the second slice being an I slice, receiving a second slice header for the second slice without a second instance of the third syntax element.

2. The method of claim 1, wherein, the value for the second syntax element is equal to the second value for the second syntax element.

3. The method of claim 1, wherein, setting the value for the first instance of the third syntax element associated with the first slice to the second value comprises inferring the value for the first instance of the third syntax element to be the second value.

4. The method of claim 1, wherein setting a value for the first instance of a third syntax element associated with the first slice to the second value comprises: setting the value for the first instance of the third syntax element to the second value without receiving an instance of the third syntax element in the first slice header for the first slice.

5. The method of claim 1, further comprising: for a block of the first slice, determining a temporal motion vector candidate for inclusion in a motion vector candidate list, wherein determining the temporal motion vector candidate comprises: identifying the collocated picture used for the temporal motion vector prediction from the first reference picture list; identifying a collocated block in the collocated picture; and deriving the temporal motion vector candidate based on a motion vector used to decode the co-located block.

6. A device for decoding video data, the device comprising: a memory configured to store video data; and one or more processors implemented in circuitry and configured to: receive a first syntax element indicating that reference picture list information is included in a first picture header syntax structure for a first slice and in a second picture header syntax structure for a second slice; in response to the first syntax element, receive a second syntax element in the first picture header syntax structure, wherein a first value for the second syntax element indicates that a co-located picture used for temporal motion vector prediction is to be derived from a first reference picture list and a second value for the second syntax element indicates that a co-located picture used for temporal motion vector prediction is to be derived from a second reference picture list; in response to the first slice being a P slice, set a value for a first instance of a third syntax element associated with the first slice to a first value for the third syntax element, wherein the first value for the third syntax element indicates that a co-located picture used for temporal motion vector prediction is to be derived from the first reference picture list and a second value for the third syntax element indicates that a co-located picture used for temporal motion vector prediction is to be derived from a second reference picture list; after receiving the first syntax element indicating that reference picture list information is included in the first picture header syntax structure and in response to the first slice being a P slice, receive a fourth syntax element, wherein a first value for the fourth syntax element indicates that a fifth syntax element is included in a first slice header for the first slice and a second value for the fourth syntax element indicates that the fifth syntax element is not included in the first slice header; in response to the fourth syntax element being equal to the first value for the fourth syntax element, receive the fifth syntax element; determine a number of active reference pictures for the first slice based on a value for the fifth syntax element; and in response to receiving the first syntax element indicating that reference picture list information is included in the second picture header syntax structure and in response to the second slice being an I slice, receive a second slice header for the second slice without a second instance of the third syntax element.

7. The apparatus of claim 6, wherein, the value for the second syntax element is equal to the second value for the second syntax element.

8. The device of claim 6, wherein to set the value for the first instance of a third syntax element associated with the first slice to the second value, the one or more processors are further configured to infer the value for the first instance of the third syntax element to be the second value.

9. The apparatus of claim 6, wherein setting a value for the first instance of a third syntax element associated with the first slice to the second value comprises: the one or more processors are further configured to set the value for the first instance of the third syntax element to the second value without receiving an instance of a third syntax element in the first slice header for the first slice.

10. The apparatus of claim 6, wherein, the one or more processors are further configured to: For a block of the first slice, determining a temporal motion vector candidate for inclusion in a motion vector candidate list, wherein, to determine the temporal motion vector candidate, the one or more processors are further configured to: identify, from the first reference picture list, the collocated picture used as the temporal motion vector prediction; identify a collocated block in the collocated picture; and derive the temporal motion vector candidate based on a motion vector used to decode the collocated block.

11. The apparatus of claim 6, wherein, The device comprises a wireless communication device, further comprising a receiver configured to receive encoded video data.

12. The apparatus of claim 11, wherein, The wireless communication device comprises a telephone handset, and wherein the receiver is configured to demodulate a signal comprising the encoded video data according to a wireless communication standard.

13. The device of claim 6, further comprising: a display configured to display decoded video data.

14. The apparatus of claim 6, wherein, The device comprises one or more of a video camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

15. A method of encoding video data, the method comprising: in response to determining that reference picture list information is included in a first picture header syntax structure for a first slice and in a second picture header syntax structure for a second slice, generating a first syntax element that indicates that reference picture list information is included in the first picture header syntax structure and in the second picture header syntax structure; generating a second syntax element for inclusion in the first picture header syntax structure, wherein a first value for the second syntax element indicates that a collocated picture used as temporal motion vector prediction is to be derived from a first reference picture list and a second value for the second syntax element indicates that a collocated picture used as temporal motion vector prediction is to be derived from a second reference picture list; determining that the first slice of the video data referred to by the first picture header syntax structure is a P slice; in response to the first slice being a P slice, determining that a value of a first instance of a third syntax element associated with the first slice is equal to a first value for the third syntax element, wherein the first value for the third syntax element indicates that the collocated picture used as temporal motion vector prediction is to be derived from a first reference picture list and a second value for the third syntax element indicates that a collocated picture used as temporal motion vector prediction is to be derived from a second reference picture list; outputting a bitstream of encoded video data that includes the first syntax element and the first picture header syntax structure; after determining that the reference picture list information is included in the first picture header syntax structure, and in response to determining that the first slice is a P slice, generating a fourth syntax element, wherein a first value for the fourth syntax element indicates that a fifth syntax element is included in a first slice header for the first slice and a second value for the fourth syntax element indicates that the fifth syntax element is not included in the first slice header; determining a number of active reference pictures for the first slice; and in response to a first instance of the fourth syntax element being equal to the first value for the fourth syntax element, generating the fifth syntax element for inclusion in a bitstream of encoded video data, wherein a value for the fifth syntax element indicates a number of active reference pictures for the first slice; and in response to determining that reference picture list information is included in the second picture header syntax structure and in response to the second slice being an I slice, generating a second slice header for inclusion in a bitstream of encoded video data without a second instance of the third syntax element.

16. The method of claim 15, wherein, a value for the second syntax element is equal to a second value for the second syntax element.

17. The method of claim 15, further comprising: in the absence of an instance of the third syntax element in the first slice header, generating the first slice header for the first slice for inclusion in a bitstream of encoded video data.

18. The method of claim 15, further comprising: for a block of the first slice, determining a temporal motion vector candidate for inclusion in a motion vector candidate list, wherein determining the temporal motion vector candidate comprises: identifying, from the first reference picture list, the collocated picture used as the temporal motion vector predictor; identifying a collocated block in the collocated picture; and deriving the temporal motion vector candidate based on a motion vector used to decode the collocated block.

19. A device for encoding video data, the device comprising: a memory configured to store video data; and one or more processors implemented in circuitry and configured to: in response to determining that reference picture list information is included in a first picture header syntax structure for a first slice and in a second picture header syntax structure for a second slice, generate a first syntax element indicating that reference picture list information is included in the first picture header syntax structure and in the second picture header syntax structure; generate a second syntax element for inclusion in the first picture header syntax structure, wherein a first value for the second syntax element indicates that a collocated picture used as a temporal motion vector predictor is to be derived from a first reference picture list and a second value for the second syntax element indicates that a collocated picture used as a temporal motion vector predictor is to be derived from a second reference picture list; determine that the first slice of the video data referred to by the first picture header syntax structure is a P slice; in response to the first slice being a P slice, determine that a value for a first instance of a third syntax element associated with the first slice is equal to a first value for the third syntax element, wherein the first value for the third syntax element indicates that a collocated picture used as a temporal motion vector predictor is to be derived from the first reference picture list and a second value for the third syntax element indicates that a collocated picture used as a temporal motion vector predictor is to be derived from the second reference picture list; output a bitstream of encoded video data including the first syntax element and the first picture header syntax structure; generate a fourth syntax element after determining that the reference picture list information is included in the first picture header syntax structure and in response to determining that the first slice is a P slice, wherein a first value for the fourth syntax element indicates that a fifth syntax element is included in a first slice header for the first slice, and a second value for the fourth syntax element indicates that the fifth syntax element is not included in the first slice header; determine a number of active reference pictures for the first slice; and in response to the first instance of the fourth syntax element being equal to the first value for the fourth syntax element, generate the fifth syntax element for inclusion in a bitstream of encoded video data, wherein a value for the fifth syntax element indicates the number of active reference pictures for the first slice; and for a second slice of the video data referring to the second picture header syntax structure, in response to determining that reference picture list information is included in the second picture header syntax structure and in response to the second slice being an I slice, generate a second slice header for inclusion in a bitstream of encoded video data without a second instance of the third syntax element.

20. The apparatus of claim 19, wherein, the value for the second syntax element is equal to a second value for the second syntax element.

21. The device of claim 19, wherein the one or more processors are further configured to: generate the first slice header for a first slice for inclusion in a bitstream of encoded video data without including an instance of the third syntax element in the first slice header.

22. The device of claim 19, wherein the one or more processors are further configured to: For a block of the first slice, determining a temporal motion vector candidate for inclusion in a motion vector candidate list, wherein, determining that the temporal motion vector candidate comprises: identifying, from the first reference picture list, the collocated picture used as the temporal motion vector predictor; identifying a collocated block in the collocated picture; and deriving the temporal motion vector candidate based on a motion vector used to decode the collocated block.

23. The apparatus of claim 19, wherein, the device comprises a wireless communication device, further comprising a transmitter configured to transmit the encoded video data.

24. The apparatus of claim 23, wherein, the wireless communication device comprises a telephone handset, and wherein the transmitter is configured to modulate a signal comprising the encoded video data according to a wireless communication standard.

25. The device of claim 19, further comprising: a video camera configured to capture the video data.

26. The apparatus of claim 19, wherein, the device comprises one or more of a video camera, a computer, or a mobile device.

27. A device for decoding video data, the device comprising means for performing the method of any of claims 1-5.

28. A device for encoding video data, the device comprising means for performing the method of any of claims 15-18.

29. A computer program product comprising computer readable instructions, which when executed by a processor, cause the processor to perform the method of any of claims 1-5.

30. A computer program product comprising computer readable instructions which, when executed by a processor, cause the processor to perform the method of any one of claims 15-18.