Encoder, decoder, and corresponding method
By employing GDR technology and using GDR pictures within video coding, the challenges of achieving efficient random access in video coding are addressed, resulting in reduced latency and improved bitrate consistency.
Patent Information
- Application Number
- JP2025051297
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-07-05
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-26
- Estimated Expiration
- 2040-03-11
AI Technical Summary
Existing video coding technologies face challenges in achieving efficient random access without the need for intra random access point (IRAP) pictures, which often result in higher bitrate and increased latency.
The implementation of gradual decoding refresh (GDR) technology, which allows for progressive intra refresh by using GDR pictures instead of IRAP pictures, enabling random access without the need for IRAP pictures. This is achieved by defining a Video Coding Layer (VCL) Network Abstraction Layer (NAL) unit with a GDR Network Abstraction Layer (NAL) unit type (GDR_NUT) to indicate the presence of a GDR picture.
The use of GDR pictures in video coding reduces end-to-end delay and achieves a smoother, more consistent bitrate, thereby improving the efficiency of the video coding process and enhancing user experience.
Smart Images

Figure 2025096295000001_ABST
Abstract
Description
Technical Field
[0001] Technical Field In general, this disclosure describes techniques for supporting gradual decoding refresh in video coding. More specifically, this disclosure permits progressive intra refresh that enables random access without the need to use intra random access point (IRAP) pictures.
Background Art
[0002] Even the amount of video data required to depict relatively short videos can be quite large, which can cause difficulties when the data is streamed or otherwise communicated over a communication network with limited bandwidth capacity. Thus, video data is generally compressed before being communicated over modern telecommunications networks. The size of the video can also be a problem when the video is stored in a storage device because memory resources may be limited. Video compression devices often use software and / or hardware at the source to code the video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompression device that decodes the video data. Improved compression and decompression techniques that improve the compression ratio with little or no sacrifice in image quality are desirable due to limited network resources and the ever-increasing demand for higher video quality.
Summary of the Invention
[0003] The first aspect relates to a method of decoding a coded video bitstream implemented by a video decoder. The method includes the steps of: determining by the video decoder that a coded video sequence (CVS) of the coded video bitstream includes a video coding layer (VCL) network abstraction layer (NAL) unit having a gradual decoding refresh (GDR) network abstraction layer (NAL) unit type (GDR_NUT), wherein the VCL NAL unit having the GDR_NUT includes a GDR picture; starting, by the video decoder, decoding of the CVS in the GDR picture; and generating, by the video decoder, an image according to the decoded CVS.
[0004] The method provides a technique that allows for progressive intra refresh, enabling random access without the need to use an intra random access point (IRAP) picture. A video coding layer (VCL) network abstraction layer (NAL) unit has a gradual decoding refresh (GDR) network abstraction layer (NAL) unit type (GDR_NUT) to indicate that the VCL NAL unit includes a GDR picture. By using a GDR picture instead of an IRAP picture, a smoother and more consistent bitrate may be achieved, for example, due to the size of the GDR picture compared to the size of the IRAP picture, which allows for a reduced end-to-end delay (i.e., latency). Thus, a coder / decoder (also known as a "codec") in video coding is improved compared to current codecs. As a practical matter, an improved video coding process provides a better user experience for the user when the video is sent, received, and / or viewed.
[0005] In a first implementation form of the method by the first side itself, the GDR picture is the first picture in the CVS.
[0006] In a second implementation form of the method by the first side itself or any preceding implementation form of the first side, the GDR picture is the first picture in the GDR period.
[0007] In a third implementation form of the method by the first side itself or any preceding implementation form of the first side, the GDR picture has a temporal identifier (ID) equal to zero.
[0008] In a fourth implementation form of the method by the first side itself or any preceding implementation form of the first side, an access unit containing a VCL NAL unit having a GDR_NUT is referred to as a GDR access unit.
[0009] In a fifth implementation form of the method by the first side itself or any preceding implementation form of the first side, the GDR_NUT indicates to the video decoder that a VCL NAL unit having the GDR_NUT contains a GDR picture.
[0010] The second aspect relates to a method of encoding a video bitstream implemented by a video encoder. The method includes the steps of: determining, by the video encoder, a random access point for a video sequence; encoding, by the video encoder, a gradual decoding refresh (GDR) picture into a video coding layer (VCL) network abstraction layer (NAL) unit having a gradual decoding refresh (GDR) network abstraction layer (NAL) unit type (GDR_NUT) at the random access point for the video sequence; generating, by the video encoder, a bitstream including the video sequence having the GDR picture in the VCL NAL unit having the GDR_NUT at the random access point; and storing, by the video encoder, the bitstream for transmission to a video decoder.
[0011] This method provides a technique that allows progressive intra-refresh to enable random access without the need to use an Intra-Random Access Point (IRAP) picture. A Video Coding Layer (VCL) Network Abstraction Layer (NAL) unit has a Gradual Decoding Refresh (GDR) Network Abstraction Layer (NAL) unit type (GDR_NUT) to indicate that the VCL NAL unit contains a GDR picture. By using a GDR picture instead of an IRAP picture, for example, a smoother and more consistent bitrate may be achieved due to the size of the GDR picture compared to the size of the IRAP picture, which allows for a reduced end-to-end delay (i.e., latency). Thus, the coder / decoder (also known as the "codec") in video coding is improved compared to current codecs. As a practical matter, the improved video coding process provides a better user experience for the user when the video is sent, received, and / or viewed.
[0012] In a first implementation of the method according to the second aspect itself, the GDR picture is the first picture in the CVS.
[0013] In a second implementation of the method according to the second aspect itself or any preceding implementation of the second aspect, the GDR picture is the first picture in the GDR period.
[0014] In a third implementation of the method according to the second aspect itself or any preceding implementation of the second aspect, the GDR picture has a temporal identifier (ID) equal to zero.
[0015] In a fourth implementation of the method according to the second aspect itself or any preceding implementation of the second aspect, a VCL NAL unit having a GDR_NUT is referred to as a GDR access unit.
[0016] In a fifth implementation form of the method according to the second aspect itself or any previous implementation form of the second aspect, the GDR_NUT indicates to the video decoder that the VCL NAL unit having the GDR_NUT includes a GDR picture.
[0017] A third aspect relates to a decoding device. The decoding device includes a receiver configured to receive a coded video bitstream, a memory coupled to the receiver that stores instructions, and a processor coupled to the memory. The processor is configured to execute the instructions to cause the decoding device to: determine that the coded video sequence (CVS) of the coded video bitstream includes a video coding layer (VCL) network abstraction layer (NAL) unit having a gradual decoding refresh (GDR) network abstraction layer (NAL) unit type (GDR_NUT), wherein the VCL NAL unit having the GDR_NUT includes a GDR picture; start decoding the CVS in the GDR picture; and generate an image according to the decoded CVS.
[0018] The decoding device provides a technique that allows for progressive intra refresh to enable random access without the need to use an Intra Random Access Point (IRAP) picture. The Video Coding Layer (VCL) Network Abstraction Layer (NAL) unit has a Gradual Decoding Refresh (GDR) Network Abstraction Layer (NAL) unit type (GDR_NUT) to indicate that the VCL NAL unit contains a GDR picture. By using GDR pictures instead of IRAP pictures, for example, due to the size of the GDR picture compared to the size of the IRAP picture, a smoother and more consistent bitrate may be achieved, which allows for a reduced end-to-end delay (i.e., latency). Thus, the coder / decoder (also known as the "codec") in video coding is improved compared to the current codec. As a practical matter, the improved video coding process provides a better user experience for the user when the video is sent, received, and / or viewed.
[0019] In a first implementation form of the decoding device according to the third aspect itself, the decoding device further has a display configured to display the generated image.
[0020] The fourth aspect relates to an encoding device. The encoding device includes a memory storing instructions; and a processor coupled to the memory, the processor configured to execute the instructions to cause the encoding device to: determine a random access point for a video sequence; encode a gradual decoding refresh (GDR) picture into a video coding layer (VCL) network abstraction layer (NAL) unit having a GDR network abstraction layer (NAL) unit type (GDR_NUT) at the random access point for the video sequence; generate a bitstream including the video sequence having the GDR picture in the VCL NAL unit having the GDR_NUT at the random access point; and a transmitter coupled to the processor and configured to transmit the bitstream towards a video decoder.
[0021] The encoding device provides a technique that allows progressive intra-refresh to enable random access without the need to use an Intra Random Access Point (IRAP) picture. The Video Coding Layer (VCL) Network Abstraction Layer (NAL) unit has a Gradual Decoding Refresh (GDR) Network Abstraction Layer (NAL) unit type (GDR_NUT) to indicate that the VCL NAL unit contains a GDR picture. By using a GDR picture instead of an IRAP picture, for example, due to the size of the GDR picture compared to the size of the IRAP picture, a smoother and more consistent bitrate may be achieved, which allows for a reduced end-to-end delay (i.e., latency). Thus, the coder / decoder (also known as a "codec") in video coding is improved compared to current codecs. As a practical matter, the improved video coding process provides a better user experience for the user when the video is sent, received, and / or viewed.
[0022] In a first implementation form of the encoding device according to the fourth aspect itself, the memory stores the bitstream before the transmitter transmits the bitstream towards the video decoder.
[0023] The fifth aspect relates to a coding device. The coding device includes a receiver configured to receive a picture to be encoded or a bitstream to be decoded; a transmitter coupled to the receiver, the transmitter being configured to transmit the bitstream to a decoder or transmit a decoded image to a display; a memory coupled to at least one of the receiver or the transmitter, the memory being configured to store instructions; and a processor coupled to the memory, the processor being configured to execute the instructions stored in the memory to execute any of the methods disclosed herein.
[0024] The coding device provides a technique that allows progressive intra-refresh to enable random access without the need to use an Intra Random Access Point (IRAP) picture. The Video Coding Layer (VCL) Network Abstraction Layer (NAL) unit has a Gradual Decoding Refresh (GDR) Network Abstraction Layer (NAL) unit type (GDR_NUT) to indicate that the VCL NAL unit contains a GDR picture. By using a GDR picture instead of an IRAP picture, for example, due to the size of the GDR picture compared to the size of the IRAP picture, a smoother and more consistent bitrate may be achieved, which allows for a reduced end-to-end delay (i.e., latency). Thus, the coder / decoder (also known as a "codec") in video coding is improved compared to current codecs. As a practical matter, the improved video coding process provides a better user experience for the user when the video is sent, received, and / or viewed.
[0025] A sixth aspect relates to a system. The system includes an encoder and a decoder that communicates with the encoder, and the encoder or the decoder includes a decoding device, an encoding device, or a coding device disclosed in this application.
[0026] The system provides a technique that allows for progressive intra-refresh to enable random access without the need to use an Intra Random Access Point (IRAP) picture. A Video Coding Layer (VCL) Network Abstraction Layer (NAL) unit has a Gradual Decoding Refresh (GDR) Network Abstraction Layer (NAL) unit type (GDR_NUT) to indicate that the VCL NAL unit contains a GDR picture. By using GDR pictures instead of IRAP pictures, a smoother and more consistent bitrate may be achieved, for example, due to the size of the GDR picture compared to the size of the IRAP picture, which allows for reduced end-to-end latency (i.e., latency). Thus, a coder / decoder (also known as a "codec") in video coding is improved compared to current codecs. As a practical matter, the improved video coding process provides a better user experience for the user when the video is sent, received, and / or viewed.
[0027] A seventh aspect relates to coding means. The coding means includes receiving means configured to receive a picture to be encoded or a bitstream to be decoded; transmitting means coupled to the receiving means, the transmitting means being configured to transmit the bitstream to decoding means or to transmit the decoded image to display means; storage means coupled to at least one of the receiving means or the transmitting means, the storage means being configured to store instructions; and processing means coupled to the storage means, the processing means being configured to execute the instructions stored in the storage means to perform any of the methods disclosed herein.
[0028] The coding means provides a technique that allows progressive intra-refresh, which enables random access without the need to use an Intra Random Access Point (IRAP) picture. The Video Coding Layer (VCL) Network Abstraction Layer (NAL) unit has a Gradual Decoding Refresh (GDR) Network Abstraction Layer (NAL) unit type (GDR_NUT) to indicate that the VCL NAL unit contains a GDR picture. By using a GDR picture instead of an IRAP picture, a smoother and more consistent bitrate may be achieved, for example, due to the size of the GDR picture compared to the size of the IRAP picture, which allows for a reduced end-to-end delay (i.e., latency). Thus, the coder / decoder (also known as the "codec") in video coding is improved compared to current codecs. As a practical matter, the improved video coding process provides a better user experience for the user when the video is sent, received, and / or viewed.
Brief Description of the Drawings
[0029] For a more complete understanding of the present disclosure, reference is made to the following brief description, taken in conjunction with the accompanying drawings and detailed description. Here, like reference numerals represent like parts.
[0030]
Figure 1
[0031]
Figure 2
[0032]
Figure 3
[0033]
Figure 4
[0034]
Figure 5
[0035]
Figure 6
[0036]
Figure 7
[0037]
Figure 8
[0038]
Figure 9
[0039]
Figure 10
[0040]
Figure 11
Mode for Carrying Out the Invention
[0041] FIG. 1 is a block diagram showing an exemplary coding system 10 that can utilize the video coding techniques described herein. As shown in FIG. 1, the coding system 10 includes a source device 12 that provides encoded video data to be decoded by a destination device 14 at a later time. Specifically, the source device 12 may provide the video data to the destination device 14 via a computer-readable medium 16. The source device 12 and the destination device 14 can include any of a wide range of devices, including desktop computers, notebook (e.g., laptop) computers, tablet computers, set-top boxes, telephone terminals such as so-called "smart" phones, so-called "smart" pads, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, and the like. In some cases, the source device 12 and the destination device 14 may be equipped for wireless communication.
[0042] The destination device 14 may receive the encoded video data to be decoded via the computer-readable medium 16. The computer-readable medium 16 can include any type of medium or device that can move the encoded video data from the source device 12 to the destination device 14. In one example, the computer-readable medium 16 may include a communication medium that enables the source device 12 to directly transmit the encoded video data to the destination device 14 in real time. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and transmitted to the destination device 14. The communication medium can include any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium can include routers, switches, base stations, or other facilities that may be useful for facilitating communication from the source device 12 to the destination device 14.
[0043] In some examples, the encoded data may be output from the output interface 22 to a storage device. Similarly, the encoded data may be accessed from the storage device by an input interface. The storage device may include any of a variety of distributed or locally accessible data storage media, such as a hard drive, a Blu-ray disk, a digital video disk (DVD), a compact disk read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In a further example, the storage device may correspond to a file server or another intermediate storage device that can store the encoded video generated by the source device 12. The destination device 14 can access the stored video data from the storage device via streaming or downloading. The file server can be any type of server that stores the encoded video data and can transmit the encoded video data to the destination device 14. Exemplary file servers include web servers (e.g., for websites), file transfer protocol (FTP) servers, network attached storage (NAS) devices, or local disk drives. The destination device 14 can access the encoded video data through any standard data connection including an Internet connection. This can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage device may be a streaming transmission, a download transmission, or a combination thereof.
[0044] The technology of the present disclosure is not necessarily limited to wireless applications or settings. This technology can be applied to video coding that supports any of a variety of multimedia applications, such as wireless television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission, e.g., Dynamic Adaptive Streaming over HTTP (DASH), digital video encoded on a data storage medium, decoding of digital video stored on a data storage medium, or other applications. In some examples, the coding system 10 may be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.
[0045] In the example of FIG. 1, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The destination device 14 includes an input interface 28, a video decoder 30, and a display device 32. According to the present disclosure, the video encoder 20 of the source device 12 and / or the video decoder 30 of the destination device 14 may be configured to apply techniques for video coding. In other examples, the source device and the destination device may include other components or arrangements. For example, the source device 12 may receive video data from an external video source such as an external camera. Similarly, the destination device 14 may interface with an external display device rather than including an integrated display device.
[0046] The coding system 10 shown in FIG. 1 is merely an example. The techniques for video coding can be performed by any digital video encoding and / or decoding device. The techniques of the present disclosure are generally performed by a video coding device, but the techniques may typically be performed by a video encoder / decoder, commonly referred to as a “codec”. Further, the techniques of the present disclosure may be performed by a video preprocessor. The video encoder and / or decoder may be a graphics processing unit (GPU) or a similar device.
[0047] The source device 12 and the destination device 14 are merely examples of such coding devices, where the source device 12 generates coded video data for transmission to the destination device 14. In some examples, the source device 12 and the destination device 14 can operate in a substantially symmetric manner, such that each of the source device 12 and the destination device 14 includes video encoding and decoding components. Thus, the coding system 10 can support one-way or two-way video transmission between the video devices 12, 14, for example, for video streaming, video playback, video broadcast, or video telephony.
[0048] The video source 18 of the source device 12 can include a video capture device such as a video camera, a video archive including previously captured video, and / or a video feed interface for receiving video from a video content provider. As a further alternative, the video source 18 may generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video.
[0049] In some cases, when the video source 18 is a video camera, the source device 12 and the destination device 14 can form a so-called camera phone or video phone. However, as described above, the techniques described in this disclosure may be generally applicable to video coding and may also be applied to wireless and / or wired applications. In each case, the captured, pre-captured, or computer-generated video may be encoded by the video encoder 20. The encoded video information may then be output to the computer-readable medium 16 by the output interface 22.
[0050] The computer-readable medium 16 can include a temporary medium such as a wireless broadcast or a wired network transmission, or a storage medium such as a hard disk, a flash drive, a compact disk, a digital video disk, a Blu-ray disk, or other computer-readable media. In some examples, a network server (not shown) can receive the encoded video data from the source device 12 and provide the encoded video data to the destination device 14, for example, via a network transmission. Similarly, a computing device of a media manufacturing facility such as a disk stamping facility can receive the encoded video data from the source device 12 and produce a disk containing the encoded video data. Thus, the computer-readable medium 16 can be understood to include one or more computer-readable media in various forms in various examples.
[0051] The input interface 28 of the destination device 14 receives information from the computer-readable medium 16. The information on the computer-readable medium 16 may include syntax information defined by the video encoder 20, and this syntax information is also used by the video decoder 30 and includes syntax elements that describe the characteristics and / or processing of blocks and / or other coding units, such as groups of pictures (GOPs). The display device 32 displays the decoded video data to the user and may include any of a variety of display devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.
[0052] Video encoder 20 and video decoder 30 may operate according to a video coding standard such as the currently under - development High Efficiency Video Coding (HEVC) standard, or may comply with the HEVC Test Model (HM). Alternatively, video encoder 20 and video decoder 30 may also operate according to other proprietary standards or industry standards, such as the International Telecommunication Union Telecommunication Standardization Sector (ITU - T) H.264 standard, also known as Moving Picture Experts Group (MPEG) - 4, Part 10, Advanced Video Coding (AVC), or extensions of such standards. However, the technology of the present disclosure is not limited to any particular coding standard. Other examples of video coding standards include MPEG - 2 and ITU - T H.263. Although not shown in FIG. 1, in some aspects, video encoder 20 and video decoder 30 may each be integrated with an audio encoder and decoder, and may include a suitable multiplexer - demultiplexer (MUX - DEMUX) unit or other hardware and software to handle the encoding of both audio and video in a common data stream or separate data streams. If applicable, the MUX - DEMUX unit may comply with the ITU H.223 multiplexer protocol, or other protocols such as the User Datagram Protocol (UDP).
[0053] Video encoder 20 and video decoder 30 may each be implemented as any of a variety of suitable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. Where the technology is implemented partially in software, the apparatus may store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of the present disclosure. Each of video encoder 20 and video decoder 30 may be included in one or more encoders or decoders, or any of these may be integrated as part of an integrated encoder / decoder (codec) within each respective apparatus. Apparatuses including video encoder 20 and / or video decoder 30 may include integrated circuits, microprocessors, and / or wireless communication devices such as cellular telephones.
[0054] FIG. 2 is a block diagram illustrating an example of a video encoder 20 that can implement video coding techniques. Video encoder 20 can perform intra coding and inter coding of video blocks within a video slice. Intra coding relies on spatial prediction to reduce or remove spatial redundancy in the video within a given video frame or picture. Inter coding relies on temporal prediction to reduce or remove temporal redundancy in the video within adjacent frames or pictures of a video sequence. Intra mode (I mode) can refer to any of several spatial-based coding modes. Inter modes, such as uni-directional (also known as single prediction) prediction (P mode) or bi-prediction (also known as bi-prediction) (B mode), can refer to any of several time-based coding modes.
[0055] As shown in FIG. 2, video encoder 20 receives a current video block within a video frame to be encoded. In the example of FIG. 2, video encoder 20 includes a mode selection unit 40, a reference frame memory 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. Mode selection unit 40 further includes a motion compensation unit 44, a motion estimation unit 42, an intra prediction (also known as internal prediction) unit 46, and a splitting unit 48. For video block reconstruction, video encoder 20 also includes an inverse quantization unit 58, an inverse transform unit 60, and an adder 62. A deblocking filter (not shown in FIG. 2) may also be included to filter block boundaries so as to remove blocky artifacts from the reconstructed video. If desired, the deblocking filter typically filters the output of adder 62. In addition to the deblocking filter, additional filters (either within the loop or after the loop) may also be used. Such filters are not shown for simplicity, but if desired, may filter the output of adder 50 (as a loop filter).
[0056] During the encoding process, video encoder 20 receives a video frame or slice to be coded. The frame or slice may be divided into a plurality of video blocks. Motion estimation unit 42 and motion compensation unit 44 perform inter-prediction coding of the received video block with respect to one or more blocks in one or more reference frames to provide temporal prediction. Intra prediction unit 46 may alternatively perform intra-prediction coding of the received video block with respect to one or more neighboring blocks in the same frame or slice as the block to be coded to provide spatial prediction. Video encoder 20 may perform multiple coding passes, for example, to select an appropriate coding mode for each block of video data.
[0057] Furthermore, the splitting unit 48 may split the blocks of video data into sub-blocks based on the evaluation of the previous splitting method in the previous coding path. For example, the splitting unit 48 may first split a frame or a slice into the largest coding units (LCUs), and then split each LCU into sub-coding units (sub-CUs) based on rate-distortion analysis (e.g., rate-distortion optimization). The mode selection unit 40 may further generate a quadtree data structure indicating the splitting of the LCU into sub-CUs. The leaf node CUs of the quadtree may include one or more prediction units (PUs) and one or more transform units (TUs).
[0058] This disclosure uses the term "block" to refer to any of the CUs, PUs, or TUs in the context of HEVC, or similar data structures in the context of other standards (e.g., macroblocks and their sub-blocks in H.264 / AVC). A CU includes a coding node, a PU, and a TU associated with the coding node. The size of a CU corresponds to the size of the coding node and is square. The size of a CU ranges from 8×8 pixels to a size of a tree block of up to 64×64 pixels or more. Each CU can include one or more PUs and one or more TUs. The syntax data associated with a CU may describe, for example, the splitting of the CU into one or more PUs. The splitting mode may vary depending on whether the CU is encoded in skip or direct mode, intra prediction mode, or inter prediction (also known as mutual prediction) mode. A PU may be split into a non-square shape. The syntax data associated with a CU may describe, for example, the splitting of the CU into one or more TUs according to a quadtree. A TU may be square or non-square (e.g., rectangular) in shape.
[0059] The mode selection unit 40 may select one of the coding modes, for example, intra or inter, based on the error result, and provide the resulting intra or inter-coded block to the adder 50 that generates the residual block data, and provide the encoded block to the adder 62 that reconstructs it for use as a reference frame. The mode selection unit 40 also provides syntax elements such as motion vectors, intra mode indicators, partition information, and other such syntax information to the entropy coding unit 56.
[0060] The motion estimation unit 42 and the motion compensation unit 44 may be highly integrated but are shown separately for conceptual purposes. The motion estimation performed by the motion estimation unit 42 is a process that generates a motion vector and estimates the motion for a video block. The motion vector may indicate, for example, the displacement of the PU of the video block in the current video frame or picture with respect to the predicted block (or other coding unit) in the reference frame for the currently coded current block (or other coding unit) within the current frame. The predicted block is a block that is found to closely match the block to be coded with respect to the pixel difference, and the match can be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference metrics. In some examples, the video encoder 20 can calculate the values of pixel positions finer than the integers of the reference picture stored in the reference frame memory 64. For example, the video encoder 20 can interpolate the values of the 1 / 4 pixel position, 1 / 8 pixel position, or other fractional pixel positions of the reference picture. Thus, the motion estimation unit 42 can perform a motion search for all pixel positions and fractional pixel positions and output a motion vector with fractional pixel accuracy.
[0061] The motion estimation unit 42 calculates the motion vector for the PU of the video block in the inter-coded slice by comparing the position of the PU with the position of the prediction block in the reference picture. The reference picture may be selected from the first reference picture list (list 0) or the second reference picture list (list 1), each of which identifies one or more reference pictures stored in the reference frame memory 64. The motion estimation unit 42 sends the calculated motion vector to the entropy coding unit 56 and the motion compensation unit 44.
[0062] The motion compensation executed by the motion compensation unit 44 may include fetching or generating a prediction block based on the motion vector determined by the motion estimation unit 42. Here too, the motion estimation unit 42 and the motion compensation unit 44 may be functionally integrated in some examples. Upon receiving the motion vector for the current video block's PU, the motion compensation unit 44 can locate the prediction block indicated by the motion vector in one of the reference picture lists. The adder 50 forms a residual video block by subtracting the pixel values of the prediction block from the pixel values of the current video block to be coded to form a pixel difference value, as will be described later. Generally, the motion estimation unit 42 performs motion estimation for the luma component, and the motion compensation unit 44 uses the motion vector calculated based on the luma component for both the chroma component and the luma component. The mode selection unit 40 can also generate syntax elements related to the video block and the video slice for use by the video decoder 30 when decoding the video blocks of the video slice.
[0063] As an alternative to the inter prediction performed by the motion estimation unit 42 and the motion compensation unit 44 described above, the intra prediction unit 46 may perform intra prediction on the current block. Specifically, the intra prediction unit 46 may determine an intra prediction mode to be used for encoding the current block. In some examples, the intra prediction unit 46 may encode the current block using different intra prediction modes, for example, between different encoding paths, and the intra prediction unit 46 (or, in some examples, the mode selection unit 40) may select an appropriate intra prediction mode for use from among the tested modes.
[0064] For example, the intra prediction unit 46 may calculate rate-distortion values using rate-distortion analysis for different tested intra prediction modes and select the intra prediction mode having the best rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block that was encoded to generate the encoded block, as well as the bit rate (i.e., the number of bits) used to generate the encoded block. The intra prediction unit 46 can calculate a ratio from the distortion and rate for different encoded blocks and determine the intra prediction mode that exhibits the best rate-distortion value for that block.
[0065] In addition, the intra prediction unit 46 may be configured to code depth blocks of a depth map using a depth modeling mode (DMM). The mode selection unit 40 can determine, for example, using rate-distortion optimization (RDO), whether an available DMM mode results in better coding results than the intra prediction mode and other DMM modes. Data for a texture image corresponding to the depth map may be stored in the reference frame memory 64. The motion estimation unit 42 and the motion compensation unit 44 may be configured to inter-predict depth blocks of the depth map.
[0066] After selecting an intra prediction mode (e.g., one of a conventional intra prediction mode or a DMM mode) for a block, the intra prediction unit 46 can provide information indicating the selected intra prediction mode for that block to the entropy coding unit 56. The entropy coding unit 56 can encode the information indicating the selected intra prediction mode. The video encoder 20 may include configuration data in the transmitted bitstream. The configuration data may include a plurality of intra prediction mode index tables and a plurality of modified intra prediction mode index tables (also referred to as codeword mapping tables), definitions of encoding contexts for various blocks, and indications of the most likely intra prediction mode, intra prediction mode index table, and modified intra prediction mode index table to be used for each context.
[0067] The video encoder 20 forms a residual video block by subtracting the prediction data from the mode selection unit 40 from the original video block to be coded. The adder 50 represents the component(s) that perform this subtraction operation.
[0068] The transformation processing unit 52 applies a transformation such as a discrete cosine transform (DCT) or a conceptually similar transform to the residual block to generate a video block including residual transform coefficient values. The transformation processing unit 52 may perform other transforms conceptually similar to the DCT. A wavelet transform, an integer transform, a subband transform, or other types of transforms may also be used.
[0069] The transformation processing unit 52 applies a transformation to the residual block to generate a block of residual transform coefficients. The transformation may transform the residual information from the pixel value domain to a transform domain, such as a frequency domain. The transformation processing unit 52 may send the resulting transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting the quantization parameter. In some examples, the quantization unit 54 may then perform a scan of the matrix including the quantized transform coefficients. Alternatively, the entropy coding unit 56 may perform the scan.
[0070] After quantization, the entropy coding unit 56 entropy-codes the quantized transform coefficients. For example, the entropy coding unit 56 may perform context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding techniques. In the case of context-based entropy coding, the context may be based on neighboring blocks. Following the entropy coding by the entropy coding unit 56, the encoded bit stream may be sent to another device (e.g., the video decoder 30) or archived for later transmission or retrieval.
[0071] The inverse quantization unit 58 and the inverse transform unit 60 respectively apply inverse quantization and inverse transform to reconstruct the residual block in the pixel region. For example, this is for later use as a reference block. The motion compensation unit 44 can calculate the reference block by adding the residual block to a predicted block of one of the frames in the reference frame memory 64. The motion compensation unit 44 may also apply one or more interpolation filters to the reconstructed residual block to calculate pixel values finer than integers for use in motion estimation. The adder 62 adds the reconstructed residual block to the motion-compensated predicted block generated by the motion compensation unit 44 to generate a reconstructed video block for storage in the reference frame memory 64. The reconstructed video block may be used by the motion estimation unit 42 and the motion compensation unit 44 as a reference block for inter-coding blocks in subsequent video frames.
[0072] FIG. 3 is a block diagram showing an example of a video decoder 30 that can implement video coding techniques. In the example of FIG. 3, the video decoder 30 includes an entropy decoding unit 70, a motion compensation unit 72, an intra prediction unit 74, an inverse quantization unit 76, an inverse transform unit 78, a reference frame memory 82, and an adder 80. The video decoder 30 can execute a decoding path that is generally inverse to the encoding path described with respect to the video encoder 20 (FIG. 2) in some examples. The motion compensation unit 72 may generate prediction data based on the motion vector received from the entropy decoding unit 70, while the intra prediction unit 74 may generate prediction data based on the intra prediction mode indicator received from the entropy decoding unit 70.
[0073] During the decoding process, video decoder 30 receives, from video encoder 20, an encoded video bitstream representing video blocks of an encoded video slice and associated syntax elements. Entropy decoding unit 70 of video decoder 30 decodes the bitstream and generates quantized coefficients, motion vectors or intra prediction mode indicators, and other syntax elements. Entropy decoding unit 70 transfers the motion vectors and other syntax elements to motion compensation unit 72. Video decoder 30 may receive syntax elements at the video slice level and / or at the video block level.
[0074] When the video slice is coded as an intra-coded (I) slice, intra prediction unit 74 may generate prediction data for the video blocks of the current video slice based on the signaled intra prediction mode and data from blocks decoded prior to the current frame or picture. When the video frame is coded as an inter-coded (e.g., B, P, or GPB) slice, motion compensation unit 72 generates a prediction block for the video blocks of the current video slice based on the motion vectors and other syntax elements received from entropy decoding unit 70. The prediction block may be generated from one of the reference pictures in one of the reference picture lists. Video decoder 30 can construct reference frame lists, list 0 and list 1, using a default construction technique based on the reference pictures stored in reference frame memory 82.
[0075] The motion compensation unit 72 determines prediction information for video blocks of the current video slice by parsing motion vectors and other syntax elements, and generates a prediction block for the currently decoded video block using the prediction information. For example, the motion compensation unit 72 uses some of the received syntax elements to determine the prediction mode (e.g., intra prediction or inter prediction) used to code the video blocks of the video slice, the inter prediction slice type (e.g., B slice, P slice, or GPB slice), the construction information for one or more of the reference picture lists for the slice, the motion vectors for each inter-encoded video block of the slice, the inter prediction status for each inter-coded video block of the slice, and other information for decoding the video blocks in the current video slice.
[0076] The motion compensation unit 72 may also perform interpolation based on an interpolation filter. The motion compensation unit 72 can use the interpolation filter used by the video encoder 20 during encoding of the video block to calculate interpolated values for pixels finer than the integers of the reference block. In this case, the motion compensation unit 72 may determine the interpolation filter used by the video encoder 20 from the received syntax elements and use the interpolation filter to generate the prediction block.
[0077] Data for the texture image corresponding to the depth map may be stored in the reference frame memory 82. The motion compensation unit 72 may be configured to inter-predict depth blocks of the depth map.
[0078] Taking the above into consideration, video compression techniques perform spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or remove redundancy inherent in a video sequence. For block-based video coding, a video slice (i.e., a video picture or a part of a video picture) may be divided into video blocks, and such video blocks may also be referred to as tree blocks, coding tree blocks (CTBs), coding tree units (CTUs), coding units (CUs) and / or coding nodes. Video blocks within an intra-coded (I) slice of a picture are encoded using spatial prediction with respect to reference samples in neighboring blocks within the same picture. Video blocks within an inter-coded (P or B) slice of a picture can use spatial prediction with respect to reference samples in neighboring blocks within the same picture, or temporal prediction with respect to reference samples in other reference pictures. A picture may be referred to as a frame, and a reference picture may be referred to as a reference frame.
[0079] Spatial or temporal prediction gives a prediction block for the block to be coded. Residual data represents the pixel difference between the original block to be coded and the prediction block. An inter-coded block is coded according to a motion vector indicating a block of reference samples forming the prediction block and residual data indicating the difference between the block to be coded and the prediction block. An intra-coded block is coded according to an intra-coding mode and residual data. For further compression, the residual data may be transformed from a pixel domain to a transform domain to give residual transform coefficients, which may then be quantized. The quantized transform coefficients are initially arranged in a two-dimensional array and may be scanned to generate a one-dimensional vector of transform coefficients, and entropy coding may be applied to achieve further compression.
[0080] Image and video compression has experienced rapid growth, leading to various coding standards. Such video coding standards include ITU-T H.261, ISO / IEC MPEG-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, Advanced Video Coding (AVC) also known as ITU-T H.264 or ISO / IEC MPEG-4 Part 10, and High Efficiency Video Coding (HEVC) also known as ITU-T H.265 or MPEG-H Part 2. AVC includes extensions such as Scalable Video Coding (SVC), Multi-View Video Coding (MVC) and Multi-View Video Coding Plus Depth (MVC+D), and 3D AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multi-View HEVC (MV-HEVC), and 3D HEVC (3D-HEVC).
[0081] There is also a new video coding standard named Versatile Video Coding (VVC) being developed by the Joint Video Experts Team (JVET) of ITU-T and ISO / IEC. There are several working drafts for the VVC standard. Here, in particular, a certain working draft of VVC (Working Draft, WD), namely, B. Bross, J. Chen, and S. Liu, "Versatile Video Coding (Draft 4)," JVET-M1001-v5, 13th JVET Meeting, January 2019 (VVC Draft 4) is referred to.
[0082] The description of the technology disclosed herein is based on the Versatile Video Coding (VVC), a video coding standard under development by the Joint Video Experts Team (JVET) of ITU-T and ISO / IEC. However, this technology is also applicable to other video codec specifications.
[0083] FIG. 4 is a representation 400 of the relationship between an IRAP picture 402 and a leading picture 404 and a subsequent picture 406 in decode order 408 and presentation order 410. In some embodiments, the IRAP picture 402 is called an instantaneous decoder refresh (IDR) picture having a clean random access (CRA) picture, or a random access decodable leading (RADL) picture. In HEVC, all IDR pictures, CRA pictures, and broken link access (BLA) pictures are considered IRAP pictures 402. For VVC, it was agreed at the 12th JVET meeting in October 2018 that both IDR pictures and CRA pictures are IRAP pictures.
[0084] As shown in FIG. 4, the leading pictures 404 (e.g., pictures 2 and 3) come after the IRAP picture 402 in decode order 408, but precede the IRAP picture 402 in presentation order 410. The subsequent picture 406 comes after the IRAP picture 402 in both decode order 408 and presentation order 410. Two leading pictures 404 and one subsequent picture 406 are depicted in FIG. 4, but those skilled in the art will understand that in actual applications, more or fewer leading pictures 404 and / or subsequent pictures 406 may exist in decode order 408 and presentation order 410.
[0085] The leading picture 404 in FIG. 4 is divided into two types, namely, random access skipped lading (RASL) and RADL. When decoding starts with an IRAP picture 402 (e.g., picture 1), the RADL picture (e.g., picture 3) can be decoded properly, but the RASL picture (e.g., picture 2) cannot be decoded properly. Thus, the RASL picture is discarded. In light of the distinction between the RADL picture and the RASL picture, the type of the leading picture associated with the IRAP picture should be identified as either RADL or RASL for efficient and proper coding. In HEVC, when RASL and RADL pictures exist, for the RASL and RADL pictures associated with the same IRAP picture, it is constrained that the RASL picture precedes the RADL picture in the presentation order 410.
[0086] The IRAP picture 402 provides the following two important functions / advantages. First, the presence of the IRAP picture 402 indicates that the decoding process can start from that picture. This function allows the decoding process to have a random access function that starts not necessarily at the beginning of the bitstream but at that position in the bitstream as long as the IRAP picture 402 exists at that position. Second, the presence of the IRAP picture 402 refreshes the decoding process, so that the coded pictures starting with the IRAP picture 402, excluding the RASL pictures, are coded without referring to the previous picture. As a result, when the IRAP picture 402 exists in the bitstream, even if there are errors that may occur during the decoding of the pictures coded before the IRAP picture 402, it will prevent those errors from propagating to the IRAP picture 402 and the pictures that come after the IRAP picture 402 in the decoding order 408.
[0087] The IRAP picture 402 provides important functions but comes with a penalty to compression efficiency. The presence of the IRAP picture 402 causes a surge in the bitrate. This penalty to compression efficiency is due to two reasons. First, since the IRAP picture 402 is an intra-predicted picture, when the picture itself is compared with other pictures (e.g., the leading picture 404, the subsequent picture 406) that are inter-predicted pictures, relatively more bits are required to represent it. Second, the presence of the IRAP picture 402 disrupts temporal prediction (this is because the decoder refreshes the decoding process. One of the actions of the decoding process for this purpose is to remove the previous reference pictures in the decoded picture buffer (DPB)). Therefore, due to the IRAP picture 402, the coding of the pictures following the IRAP picture 402 in decoding order 408 is less efficient (i.e., requires more bits to represent) because there are fewer reference pictures for inter-prediction coding.
[0088] Among the picture types considered as the IRAP picture 402, the IDR picture in HEVC has different signaling and derivation compared to other picture types. Some of the differences are as follows.
[0089] For the signaling and derivation of the picture order count (POC) value of the IDR picture, the most significant bit (MSB) part of the POC is not derived from the previous key picture and is simply set equal to 0.
[0090] Regarding the signaling information necessary for reference picture management, the slice header of an IDR picture does not contain the information that needs to be signaled to assist in reference picture management. For other picture types (i.e., CRA, Trailing, temporal sub-layer access (TSA), etc.), for the reference picture marking process (i.e., the process of determining the state of reference pictures in the decoded picture buffer (DPB) regardless of whether they are used for reference or not), information such as the reference picture set (RPS) described below or other forms of similar information (e.g., reference picture lists) is required. However, for IDR pictures, such information does not need to be signaled. This is because the presence of an IDR indicates that the decoding process should simply mark all reference pictures in the DPB as not being used for reference.
[0091] In HEVC and VVC, the IRAP picture 402 and the leading picture 404 may each be contained within a single Network Abstraction Layer (NAL) unit. A set of NAL units may be referred to as an access unit. The IRAP picture 402 and the leading picture 404 are given different NAL unit types and can thus be easily identified by system-level applications. For example, a video splicer needs to understand the coded picture type without the need to understand the excessive details of the syntax elements within the coded bitstream, and in particular, to identify the IRAP picture 402 from non-IRAP pictures and to identify the leading picture 404 from subsequent pictures 406, including discriminating RASL and RADL pictures. The subsequent pictures 406 are pictures that are associated with the IRAP picture 402 and come after the IRAP picture 402 in presentation order 410. A picture may come after a particular IRAP picture 402 in decode order 408 and may precede any other IRAP picture 402 in decode order 408. For this purpose, giving the IRAP picture 402 and the leading picture 404 their own NAL unit types helps such applications.
[0092] For HEVC, the NAL unit types of IRAP pictures include the following: BLA with leading picture (BLA_W_LP): The NAL unit of a Broken Link Access (BLA) picture where one or more leading pictures may follow in decode order BLA with RADL (BLA_W_RADL): The NAL unit of a BLA picture where one or more RADL pictures may follow in decode order but no RASL pictures follow BLA without leading picture (BLA_N_LP): The NAL unit of a BLA picture where no leading picture follows in decode order IDR with RADL (IDR_W_RADL): A NAL unit of an IDR picture where one or more RADL pictures may follow in decoding order, but no RASL picture follows. IDR without leading picture (IDR_N_LP): A NAL unit of an IDR picture where no leading picture follows in decoding order. CRA: A NAL unit of a clean random access (CRA) picture where a leading picture (i.e., a RASL picture or a RADL picture, or both) may follow. RADL: A NAL unit of a RADL picture. RASL: A NAL unit of a RASL picture.
[0093] For VVC, the NAL unit types of IRAP picture 402 and leading picture 404 are as follows: IDR with RADL (IDR_W_RADL): A NAL unit of an IDR picture where one or more RADL pictures may follow in decoding order, but no RASL picture follows. IDR without leading picture (IDR_N_LP): A NAL unit of an IDR picture where no leading picture follows in decoding order. CRA: A NAL unit of a clean random access (CRA) picture where a leading picture (i.e., a RASL picture or a RADL picture, or both) may follow. RADL: A NAL unit of a RADL picture. RASL: A NAL unit of a RASL picture.
[0094] Progressive intra refresh / Progressive decode refresh is discussed below.
[0095] For low-latency applications, it is desirable to avoid coding a picture as an IRAP picture (e.g., IRAP picture 402). This is because its bitrate requirement is relatively large compared to non-IRAP (i.e., P / B) pictures, and as a result, it causes a larger latency / delay. However, it may not be possible to completely avoid the use of IRAP in all low-latency applications. For example, for conversational applications such as multiparty remote conferencing, it is necessary to provide regular points at which new users can join the remote conference.
[0096] To provide access to a bitstream that allows new users to join a multiparty remote conferencing application, one possible strategy is to use Progressive Intra Refresh (PIR) technology instead of using IRAP pictures in order to avoid having a peak in the bitrate. PIR is also sometimes referred to as gradual decoding refresh (GDR). The terms PIR and GDR may be used interchangeably in this disclosure.
[0097] FIG. 5 shows the Gradual Decode Refresh (GDR) technique 500. As shown, the GDR technique 500 is depicted using GDR pictures 502, one or more subsequent pictures 504, and recovery point pictures 506 within a coded video sequence 508 of a bitstream. In certain embodiments, the GDR pictures 502, subsequent pictures 504, and recovery point pictures 506 can define a GDR period within the CVS 508. The CVS 508 is a series of pictures (or portions thereof) starting with the GDR picture 502 and including all pictures (or portions thereof) up to, but not including, the next GDR picture, or up to the end of the bitstream. The GDR period is a series of pictures starting with the GDR picture 502 and including all pictures up to and including the recovery point picture 506.
[0098] As shown in FIG. 5, the GDR technique 500 or principle functions over a series of pictures starting with the GDR picture 502 and ending with the recovery point picture 506. The GDR picture 502 includes a refresh / clean area 510 containing blocks coded using only intra prediction (i.e., intra prediction blocks), and an unrefreshed / dirty area 512 containing blocks coded using only inter prediction (i.e., inter prediction blocks).
[0099] The subsequent picture 504 that is immediately adjacent to the GDR picture 502 includes a refresh / clean area 510 that has a first portion 510A coded using intra prediction and a second portion 510B coded using inter prediction. The second portion 510B is coded, for example, by referring to the refresh / clean area 510 of a preceding picture within the GDR period of the CVS 508. As shown in the figure, the refresh / clean area 510 of the subsequent picture 504 expands as the coding process moves or progresses in a consistent direction (e.g., from left to right), and correspondingly, the unrefreshed / dirty area 512 is shrunk. Eventually, a recovery point picture 506 that only includes the refresh / clean area 510 is obtained from the coding process. It should be noted that, as further discussed below, the second portion 510B of the refresh / clean area 510 coded as an inter prediction block may simply refer to the refresh area / clean area 510 in the reference picture.
[0100] In HEVC, the GDR technique 500 of FIG. 5 is non-normatively supported using a recovery point supplementary enhancement information (SEI) message and an area refresh information SEI message. These two SEI messages do not define how GDR is executed. Rather, these two SEI messages simply provide a mechanism to indicate the first and last pictures (i.e., provided by the recovery point SEI message) and the refreshed areas (i.e., provided by the area refresh information SEI message) during the GDR period.
[0101] In practice, GDR technique 500 is implemented by using two techniques together. Those two techniques are constraint intra prediction (CIP) and encoder constraint conditions for motion vectors. CIP can be used for the purpose of GDR, especially to code areas that are coded only as intra prediction blocks (e.g., the first part 510A of the refresh / clean area 510). Because CIP allows areas that do not use samples from non-refreshed areas (e.g., the non-refreshed / dirty area 512) to be used for reference. However, the use of CIP causes severe coding performance degradation because the constraint conditions for intra blocks must be applied not only to intra blocks in the refreshed area but also to all intra blocks in the picture. The encoder constraint conditions for motion vectors restrict the encoder from using any samples in the reference picture located outside the refresh area. Such constraint conditions cause non-optimal motion search.
[0102] Figure 6 is a schematic diagram showing non-optimal motion search 600 when encoder constraints are used to support GDR. As shown in the figure, motion search 600 depicts the current picture 602 and the reference picture 604. The current picture 602 and the reference picture 604 each include a refresh area 606 coded by intra prediction, a refresh area 608 coded by inter prediction, and a non-refresh area 608. The refresh area 604, the refresh area 606, and the non-refresh area 608 are similar to the first part 510A of the refresh / clean area 510, the second part 510B of the refresh / clean area 510, and the non-refresh / dirty area 512 in FIG. 5.
[0103] During the motion search process, the encoder is constrained or prevented from selecting any motion vector 610 that leads to a sample of the reference block 612 located outside the refresh region 606. This occurs even if the reference block 612 provides the best rate-distortion cost criterion when predicting the current block 614 within the current picture 602. Thus, FIG. 6 shows the reason for the non-optimality in motion search 600 when encoder constraints are used to support GDR.
[0104] Contributions to JVET JVET-K0212 and JVET-L0160 describe the implementation of GDR based on the use of CIP and the encoder constraint approach. The implementation can be summarized as follows: The intra prediction mode is forced for coding units on a column-by-column basis, constrained intra prediction is enabled to ensure the reconstruction of intra CUs, and motion vectors are constrained to point within the refresh region while taking into account an additional margin (e.g., 6 pixels) to avoid error diffusion for filters, and the previous reference pictures are removed when re-looping the intra columns.
[0105] Contribution to JVET JVET-M0529 proposed a way to indicate canonically that a picture is the first and last during the GDR period. The proposed idea works as follows.
[0106] As a non-video coding layer (VCL) NAL unit, a new NAL unit with a NAL unit type recovery point indication is defined. The payload of the NAL unit contains syntax elements for specifying information that can be used to derive the POC value of the last picture in the GDR period. An access unit containing a non-VCL NAL unit with a type recovery point indication is called a recovery point begin (RBP) access unit (AU), and the pictures within the RBP access unit are called RBP pictures. The decoding process can start from the RBP AU. When decoding starts from the RBP AU, all pictures within the GDR period except the last picture are not output.
[0107] Some of the problems of the existing GDR design are discussed.
[0108] The existing designs / approaches for supporting GDR have at least the following problems.
[0109] The method of defining GDR normatively in JVET-M0529 has the following problems. The proposed method does not describe how GDR is executed. Instead, the proposed method only provides some signaling to indicate the first and last pictures in the GDR period. A new non-VCL NAL unit is required to indicate the first and last pictures in the GDR period. This is redundant because the information contained in the recovery point indication (RPI) NAL unit can simply be contained in the tile group header of the first picture in the GDR period. Also, the proposed method cannot describe which regions in the pictures during the GDR period are the refreshed regions and the non-refreshed regions.
[0110] The GDR approach described in JVET-K0212 and JVET-L0160 has the following problems. First, there is the use of CIP. In order to prevent any sample from the non-refreshed area from being used for spatial reference, it is necessary to code the refresh area by intra prediction with some constraints. When CIP is used, the coding is picture-based, which means that all intra blocks within a picture must be coded as CIP intra blocks. This results in performance degradation. Furthermore, when the samples of the reference blocks related to the motion vectors are not completely within the refresh area in the reference picture, the encoder is prevented from selecting the best motion vector by using encoder constraint conditions that limit motion search. Also, the refresh area coded only by intra prediction is not the CTU size. Instead, the refresh area can be smaller than the CTU size and can even be the minimum CU size. This may require instructions at the block level, making the implementation unnecessarily complex.
[0111] Disclosed herein are techniques for supporting Gradual Decoding Refresh (GDR) in video coding. The disclosed techniques allow for Progressive Intra Refresh to enable random access without the need to use an intra random access point picture. As will be explained in more detail below, a Video Coding Layer (VCL) Network Abstraction Layer (NAL) unit has a Gradual Decoding Refresh (GDR) Network Abstraction Layer (NAL) unit type (GDR_NUT) to indicate that the VCL NAL unit contains a GDR picture. That is, the GDR_NUT directly indicates to the video decoder that the VCL NAL unit having the GDR_NUT contains a GDR picture. By using GDR pictures instead of IRAP pictures, a smoother and more consistent bitrate can be achieved, for example, due to the size of the GDR picture compared to the size of the IRAP picture. This allows for reducing the end-to-end delay (i.e., latency). Thus, a coder / decoder (also known as a "codec") in video coding is improved compared to current codecs. As a practical matter, the improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed.
[0112] To solve one or more of the problems discussed above, the present disclosure discloses the following aspects. Each aspect can be applied individually and some of them can be applied in combination.
[0113] 1) A VCL NAL unit having the type GDR_NUT is defined.
[0114] a. A picture having the NAL unit type GDR_NUT is referred to as a GDR picture, i.e., the first picture in a GDR period.
[0115] b. The GDR picture has a temporalID (temporal identifier) equal to 0.
[0116] c. The access unit containing the GDR picture is called a GDR access unit. As described above, an access unit is a set of NAL units. Each NAL unit can contain a single picture.
[0117] 2) The coded video sequence (CVS) may start with a GDR access unit.
[0118] 3) The GDR access unit is the first access unit in the CVS if any of the following is true.
[0119] a. The GDR access unit is the first access unit in the bitstream.
[0120] b. The GDR access unit comes immediately after the end-of-sequence (EOS) access unit.
[0121] c. The GDR access unit comes immediately after the end-of-bitstream (EOB) access unit.
[0122] d. A decoder flag called NoIncorrectPicOutputFlag is associated with the GDR picture, and the value of this flag is set to 1 (i.e., true) by an entity external to the decoder.
[0123] 4) If the GDR picture is the first access unit in the CVS, the following apply.
[0124] a. All reference pictures in the DPB are marked "not used for reference".
[0125] b. The MSB of the POC of the picture is set equal to 0.
[0126] c. The GDR picture and all pictures following the GDR picture in output order (excluding the last picture in the GDR period) are not output (i.e., marked as "not needed for output") until the last picture in the GDR period.
[0127] 5) A flag that specifies whether GDR is enabled is signaled in a sequence level parameter set (e.g., in the SPS).
[0128] a. The flag may be specified as gdr_enabled_flag.
[0129] b. If the flag is equal to 1, the GDR picture may be present in the CVS. Otherwise, if the flag is equal to 0, GDR is not enabled and the GDR picture is not present in the CVS.
[0130] 6) Information that can be used to derive the POC value of the last picture in the GDR period is signaled in the tile group header of the GDR picture.
[0131] a. The information is signaled as the delta POC between the last picture in the GDR period and the GDR picture. The information can be signaled using a syntax element designated as recovery_point_cnt.
[0132] b. The presence of the syntax element recovery_point_cnt in the tile group header may be conditional based on the value of the gdr_enabled flag and the NAL unit type of the picture. That is, the flag is present only if the gdr_enabled_flag is equal to 1 and the nal_unit_type of the NAL unit containing the tile group is GDR_NUT.
[0133] 7) A flag that specifies whether a tile group is part of the refresh region is signaled in the tile group header.
[0134] a. The flag may be designated as refreshed_region_flag.
[0135] b. The presence of the flag may be conditional based on the value of gdr_enabled_flag and whether the picture containing the tile group is within the GDR period. Thus, the flag is present only if all of the following are true.
[0136] i. The value of gdr_enabled_flag is equal to 1.
[0137] ii. The POC of the current picture is greater than or equal to the POC value of the last GDR picture (if the current picture is a GDR picture, the last GDR picture is the current picture), and less than the POC of the last picture in the GDR period.
[0138] c. If the flag does not exist in the tile group header, the value of the flag is presumed to be equal to 1.
[0139] 8) All tile groups for which refreshed_region_flag is equal to 1 cover a contiguous region. Similarly, all tile groups for which refreshed_region_flag is equal to 0 also cover a contiguous region.
[0140] 9) Tile groups with refreshed_region_flag can be of type I (i.e., intra-tile group) or B or P (i.e., inter-tile group).
[0141] 10) Each picture starting from the GDR picture up to the last picture in the GDR period includes at least one tile group where the refreshed_region_flag is equal to 1.
[0142] 11) The GDR picture includes at least one tile group where the refreshed_region_flag is equal to 1 and the tile_group_type is equal to I (i.e., an intra-tile group).
[0143] 12) When the gdr_enabled_flag is equal to 1, it is allowed that the information of the rectangular tile group, i.e., the number of tile groups and their addresses, is signaled either in the Picture Parameter Set (PPS) or in the tile group header. To do this, a flag for specifying whether the rectangular tile group information exists in the PPS is signaled in the PPS. This flag may be called rect_tile_group_in_pps_flag. This flag may be constrained to be equal to 1 when the gdr_enabled_flag is equal to 1.
[0144] a. In an alternative, instead of signaling whether the rectangular tile group information exists in the PPS, a more general flag for specifying whether tile group information (i.e., any type of tile group such as a rectangular tile group, a raster scan tile group, etc.) exists in the PPS can be signaled in the PPS.
[0145] 13) If the tile group information does not exist in the PPS, it may be further constrained that there is no signaling of explicit tile group identifier (ID) information. The explicit tile group ID information includes signaled_tile_group_id_flag, signaled_tile_group_id_length_minus1, and tile_group_id[i].
[0146] 14) A flag is signaled to specify whether a loop filtering operation across the boundary between the refreshed region and the non-refreshed region within a picture is allowed.
[0147] a. This flag is signaled in the PPS and is called loop_filter_across_refreshd_region_enabled_flag.
[0148] b. The presence of loop_filter_across_refreshed_region_enabled_flag may be conditioned on the value of loop_filter_across_tile_enabled_flag. If loop_filter_across_tile_enabled_flag is equal to 0, loop_filter_across_refreshed_region_enabled_flag may not be present and its value is assumed to be equal to 0.
[0149] c. In an alternative, the flag may be signaled in the tile group header and its presence may be conditioned based on the value of refreshed_region_flag. That is, the flag is present only when the value of refreshed_region_flag is equal to 1.
[0150] 15) If a tile group is indicated to be a refreshed region and it is indicated that a loop filter across the refreshed region is not allowed, the following applies.
[0151] a. Deblocking of edges at the tile group boundary is not performed if the neighboring tile group sharing the edge is a non-refreshed tile group.
[0152] b. The sample adaptive offset (SAO) process for blocks at the tile group boundary does not use any samples from outside the refreshed region boundary.
[0153] c. For the block at the boundary of the tile group, the adaptive loop filtering (ALF) process does not use any samples from outside the refresh region boundary.
[0154] 16) If gdr_enabled_flag is equal to 1, each picture is associated with variables for determining the boundary of the refresh region within the picture. These variables may be called as follows.
[0155] a. For the left boundary position of the refresh region within the picture, PicRefreshedLeftBoundaryPos.
[0156] b. For the right boundary position of the refresh region within the picture, PicRefreshedRightBoundaryPos.
[0157] c. For the top boundary position of the refresh region within the picture, PicRefreshedTopBoundaryPos.
[0158] d. For the bottom boundary position of the refresh region within the picture, PicRefreshedBotBoundaryPos.
[0159] 17) The boundaries of the refresh region within the picture may be derived. The boundaries of the refresh region of the picture are updated by the decoder after the tile group header is parsed and the value of the refreshd_region_flag of the tile group becomes equal to 1.
[0160] 18) In an alternative to solution 17, the boundaries of the refreshed region within the picture are explicitly signaled in each tile group of the picture.
[0161] a. A flag indicating whether the picture to which the tile group belongs contains an unrefreshed area may be signaled. If it is specified that the picture does not contain an unrefreshed area, the refreshed boundary information is not signaled and can simply be assumed to be equal to the picture boundary.
[0162] 19) For the current picture, the boundaries of the refresh area are used as follows in the in-loop filter process.
[0163] a. For the deblocking process, to determine the edges of the refresh area in order to determine whether an edge needs to be deblocked.
[0164] b. For the SAO process, to determine the boundaries of the refresh area so that the clipping process can be applied to avoid using samples from the unrefreshed area when a loop filter across the refresh area is not allowed.
[0165] c. For the ALF process, to determine the boundaries of the refresh area so that the clipping process can be applied to avoid using samples from the unrefreshed area when a loop filter across the refresh area is not allowed.
[0166] 20) For the motion compensation process, information regarding the boundaries of the refresh area, particularly the boundaries of the refresh area in the reference picture, is used as follows. If the current block in the current picture is in a tile group with a refreshed_region_flag equal to 1 and the reference block is in a reference picture that contains an unrefreshed area, the following applies.
[0167] a. The motion vector from the current block to its reference picture is clipped by the boundaries of the refresh area within that reference picture.
[0168] b. For the fractional interpolation filter for samples within the reference picture, it is clipped by the boundaries of the refresh region within the reference picture.
[0169] A detailed description of embodiments of the present disclosure is provided. This description is with respect to the base text which is JVET's contribution JVET-M1001-v5. That is, only the differences are described, and the text in the base text not mentioned below is applied as it is. The modified text with respect to the base text is italicized [underlined in the text].
[0170] Definitions are given.
[0171] 3.1 Clean Random Access (CRA) Picture: An IRAP picture in which each VCL NAL unit has a nal_unit_type equal to CRA_NUT.
[0172] Note - A CRA picture does not reference any picture other than itself for inter prediction in its decoding process, and may be the first picture in the bitstream in decoding order, or may appear later in the bitstream. A CRA picture may have an associated RADL or RASL picture. If the CRA picture has a value equal to 1 NoIncorrectPicOutputFlag for, the associated RASL picture is not output by the decoder. This is because an RASL picture may contain references to pictures that do not exist within the bitstream and thus may not be decodable.
[0173] 3.2 Coding Video Sequence (CVS): In decoding order, an IRAP access unit having a value equal to 1 NoIncorrectPicOutputFlag and the subsequent IRAP access units having a value equal to 1 or a GDR access unit with NoIncorrectPicOutputFlag equal to 1 and the subsequent IRAP access units having a value equal to 1 NoIncorrectPicOutputFlag and the subsequent IRAP access units having a value equal to 1 or a GDR access unit with NoIncorrectPicOutputFlag equal to 1Composed of zero or more access units that are not, and an IRAP access unit with a NoIncorrectPicOutputFlag equal to 1 or a GDR access unit with NoIncorrectPicOutputFlag equal to 1 A coding video sequence (CVS) that includes all subsequent access units up to, but not including, any subsequent access unit that is such.
[0174] Note 1 - The IRAP access unit can be an IDR access unit or a CRA access unit. NoIncorrectPicOutputFlag The value of is equal to 1 for each IDR access unit and each CRA access unit that is the first access unit in the bitstream in decoding order, the first access unit following the end of the sequence NAL unit in decoding order, or has a HandleCraAsCvsStartFlag equal to 1.
[0175] Note 2 - NoIncorrectPicOutputFlag is equal to 1 for the first access unit in the bitstream in decode order, the first access unit following the end of the sequence NAL unit in decode order, or for each GDR access unit with HandleGdrAsCvsStartFlag equal to 1.
[0176] 3.3 Progressive Decoding Refresh (GDR) Access Unit: An access unit in which the coding picture is a GDR picture.
[0177] 3.4 Progressive Decoding Refresh (GDR) Picture: A picture in which each VCL NAL unit has a nal_unit_type equal to GDR_NUTS.
[0178] 3.5 Random Access Skip Leading (RASL) Picture: A coded picture in which each VCL NAL unit has a nal_unit_type equal to RASL_NUT.
[0179] Note - All RASL pictures are the leading pictures of the associated CRA picture. If the associated CRA picture has a value equal to 1 NoIncorrectPicOutputFlag the RASL picture is not output and may not be correctly decodable. This is because the RASL picture may contain references to pictures that do not exist in the bitstream. The RASL picture is not used as a reference picture for the decoding process of non-RASL pictures. If present, all RASL pictures come before, in decoding order, all subsequent pictures of the same associated CRA picture.
[0180] Raw byte sequence payload (RBSP) syntax and semantic content of the sequence parameter set [Table 1]
[0181] That the gdr_enabled_flag is equal to 1 specifies that a GDR picture may exist in the coded video sequence. That the gdr_enabled_flag is equal to 0 specifies that a GDR picture does not exist in the coded video sequence.
[0182] RBSP syntax and semantic content of the picture parameter set [Table 2]
[0183] That the rect_tile_group_info_in_pps_flag is equal to 1 specifies that rectangular tile group information is signaled in the PPS. That the rect_tile_group_info_in_pps_flag is equal to 0 specifies that rectangular tile group information is not signaled in the PPS.
[0184] That the value of rect_tile_group_info_in_pps_flag is equal to 0 when the value of gdr_enabled_flag in the active SPS is equal to 0 is a requirement for bitstream conformance.
[0185] The loop_filter_across_refreshed_region_enabled_flag being equal to 1 specifies that the in-loop filtering operation may be performed across the boundaries of tile groups with a refreshed_region_flag equal to 1 in the picture that references the PPS. The loop_filter_across_refreshed_region_enabled_flag being equal to 0 specifies that the in-loop filtering operation is not performed across the boundaries of tile groups with a refreshed_region_flag equal to 1 in the picture that references the PPS. The in-loop filtering operation includes deblocking filter, sample adaptive offset filter, and adaptive loop filter operations. If not present, the value of the loop_filter_across_refreshed_region_enabled_flag is assumed to be equal to 0.
[0186] When signalled_tile_group_id_flag is equal to 1, it specifies that the tile group ID for each tile group is signalled. When signalled_tile_group_index_flag is equal to 0, it specifies that the tile group ID is not signalled. Not present In the case, the value of signalled_tile_group_index_flag is presumed to be equal to 0.
[0187] One plus signalled_tile_group_length_minus1 specifies the number of bits used to represent the syntax element tile_group_id[i] and the syntax element tile_group_address in the tile group header, if present. The value of signalled_tile_group_index_length_minus1 is in the range 0 to 15 (inclusive). If not present, the value of signalled_tile_group_index_length_minus1 is assumed as follows.
[0188] If rect_tile_group_info_in_pps_flag is equal to 1, Ceil(Log2(num_tile_groups_in_pic_minus1 + 1)) - 1.
[0189] Otherwise, Ceil(Log2(NumTilesInPic)) - 1.
[0190] General tile group header syntax and semantic content [Table 3]
[0191] tile_group_address specifies the tile address of the first tile in the tile group. If not present, the value of tile_group_address is assumed to be equal to 0.
[0192] When rect_tile_group_flag is equal to 0, the following applies: tile_group_address is the tile ID specified by Equation 6-7. The length of tile_group_address is Ceil(Log2(NumTilesInPic)) bits. The value of tile_group_address is in the range 0 to NumTilesInPic - 1 (inclusive).
[0193] Otherwise, when rect_tile_group_flag is equal to 1 and rect_tile_group_info_in_pps is equal to 0, the following applies: The tile_group_address is the tile index of the tile located at the upper left corner of the i-th tile group. The length of the tile_group_address is signalled_tile_group_index_length_minus1 + 1 bits. If signelled_tile_group_id_flag is equal to 0, the value of tile_group_address ranges from 0 to NumTilesInPic - 1 (both ends inclusive). Otherwise, the value of tile_group_address ranges from 0 to 2 (signalled_tile_group_index_length_minus1+1) - 1 (both ends inclusive).
[0194] Otherwise (when rect_tile_group_flag is equal to 1 and rect_tile_group_in_pps is equal to 1), The following applies. The tile_group_address is the tile group ID of the tile group. The length of the tile_group_address is signalled_tile_group_index_length_minus1 + 1 bits. When the signalled_tile_group_id_flag is equal to 0, the value of the tile_group_address shall be in the range from 0 to num_tile_groups_in_pic_minus1 (inclusive at both ends). Otherwise, the value of the tile_group_address shall be in the range from 0 to 2 (signalled_tile_group_index_length_minus1+1) - 1 (inclusive at both ends).
[0195] bottom_right_tile_id specifies the tile index of the tile located at the bottom - right corner of the tile group. When single_tile_per_tile_group_flag is equal to 1, bottom_right_tile_id is assumed to be equal to tile_group_address. The length of the bottom_right_tile_id syntax element is Ceil(Log2(NumTilesInPic)) bits.
[0196] The variable NumTilesInCurrTileGroup that specifies the number of tiles in the current tile group, the TopLeftTileIdx that specifies the tile index of the top - left tile of the tile group, the BottomRightTileIdx that specifies the tile index of the bottom - right tile of the tile group, and the TgTileIdx[i] that specifies the tile index of the i - th tile of the current tile group are derived as follows. [Table 4]
[0197] recovery_poc_cnt specifies the recovery point of the decoded picture in output order. In CVS, if there is a picture picA that follows the current picture (i.e., the GDR picture) in decode order and has a PicOrderCntVal equal to the value obtained by adding the value of recovery_poc_cnt to the PicOrderCntVal of the current picture, then picture picA is called a recovery - point picture. Otherwise, the first picture in output order that has a PicOrderCntVal greater than the value obtained by adding the value of recovery_poc_cnt to the PicOrderCntVal of the current picture is called a recovery - point picture. The recovery - point picture does not precede the current picture in decode order. All decoded pictures in output order are shown to be correct or nearly correct in the content starting from the output - order position of the recovery - point picture. The value of recovery_poc_cnt shall be in the range from - MaxPicOrderCntLsb / 2 to MaxPicOrderCntLsb / 2 - 1 (both ends inclusive).
[0198] The value of RecoveryPointPocVal is derived as follows.
[0199] RecoveryPointPocVal = PicOrderCntVal+recovery_poc_cnt.
[0200] The fact that the refreshed_region_flag is equal to 1 specifies that the decoding of the tile group generates correct reconstructed sample values regardless of the value of NoIncorrectPicOutputFlag of the associated GDR. The fact that the refreshed_region_flag is equal to 0 specifies that the decoding of the tile group may generate incorrect reconstructed sample values when starting from an associated GDR with NoIncorrectPicOutputFlag equal to 1. If not present, the value of the refreshed_region_flag is assumed to be equal to 1.
[0201] Note x - The current picture itself can be a GDR picture with NoIncorrectPicOutputFlag equal to 1.
[0202] The tile group refreshed boundary is derived as follows:
Table 5
[0203] Meaning content of NAL unit header
Table 6
[0204] …
[0205] If the nal_unit_type is equal to GDR_NUT, the coded tile group belongs to a GDR picture and the TemporalId is equal to 0.
[0206] The order of access units and their association with CVS are discussed.
[0207] A bitstream that conforms to this specification (i.e., the JVET contribution JVET-M1001-v5) contains one or more CVSs.
[0208] A CVS contains one or more access units. The order of NAL units and coded pictures and their association with access units are described in Section 7.4.2.4.4.
[0209] The first access unit of the CVS is one of the following.
[0210] · An IRAP access unit with NoBrokenPictureOutputFlag equal to 1.
[0211] · A GDR access unit with NoIncorrectPicOutputFlag equal to 1.
[0212] The bitstream conformance requirement, if present, is that the next access unit after an access unit including the end of the sequence NAL unit or the end of the bitstream NAL unit is one of the following.
[0213] · An IRAP access unit. This can be an IDR access unit or a CRA access unit.
[0214] · A GDR access unit.
[0215] 8.1.1 The decoding process of the coded picture is discussed.
[0216] …
[0217] If the current picture is an IRAP picture, the following applies.
[0218] · If the current picture is an IDR picture, the first picture in the bitstream in decoding order, or the first picture following the end of the sequence NAL unit in decoding order, the variable NoIncorrectPicOutputFlag is set equal to 1.
[0219] · Otherwise, if some external means (e.g., user input) not specified in this specification is available to set the variable HandleCraAsCvsStartFlag to the value of the current picture, the variable HandleCraAsCvsStartFlag is set equal to the value provided by the external means, and the variable NoIncorrectPicOutputFlag is set equal to HandleCraAsCvsStartFlag.
[0220] · Otherwise, the variable HandleCraAsCvsStartFlag is set equal to 0, and the variable NoIncorrectPicOutputFlag is set equal to 0.
[0221] If the current picture is a GDR picture, the following applies.
[0222] · If the current picture is a GDR picture, the first picture in the bitstream in decoding order, or the first picture following the end of the sequence NAL unit in decoding order, the variable NoIncorrectPicOutputFlag is set equal to 1.
[0223] · Otherwise, if some external means not specified in this specification is available to set the variable HandleGdrAsCvsStartFlag to the value for the current picture, the variable HandleGdrAsCvsStartFlag is set equal to the value provided by the external means, and the variable NoIncorrectPicOutputFlag is set equal to HandleGdrAsCvsStartFlag.
[0224] · Otherwise, the variable HandleGdrAsCvsStartFlag is set equal to 0, and the variable NoIncorrectPicOutputFlag is set equal to 0.
[0225] …
[0226] The decoding process operates as follows for the current picture CurrPic.
[0227] 1. The decoding of NAL units is specified in Section 8.2.
[0228] 2. The process in Section 8.3 uses the tile group header layer and syntax elements above it to define the following decoding process.
[0229] · Variables and functions related to the picture order count are derived as specified in Section 8.3.1. This needs to be called only for the first tile group of a picture.
[0230] · At the start of the decoding process for each tile group of a non - IDR picture, the decoding process for constructing the reference picture lists defined in Section 8.3.2 is called to derive reference picture list 0 (RefPicList[0]) and reference picture list 1 (RefPicList[1]).
[0231] · The decoding process for reference picture marking in Section 8.3.3 is called. Here, a reference picture can be marked as "not used for reference" or "used for long - term reference". This needs to be called only for the first tile group of a picture.
[0232] · PicOutputFlag is set as follows.
[0233] · If any of the following conditions is true, PictureOutputFlag is set to 0:
[0234] · The current picture is a RASL picture and the NoIncorrectPicOutputFlag of the associated IRAP picture is equal to 1.
[0235] · gdr_enabled_flag is equal to 1 and the current picture is a GDR picture with a NoIncorrectPicOutputFlag equal to 1.
[0236] · gdr_enabled_flag is equal to 1 and the current picture includes one or more tile groups with a refreshed_region_flag equal to 0, and the NoBrokenPictureOutputFlag of the associated GDR picture is equal to 1.
[0237] · Otherwise, PicOutputFlag is set equal to 1.
[0238] 3. The various decoding processes are called to code tree units, scaling, conversion, in-loop filtering, etc.
[0239] 4. After all tile groups of the current picture have been decoded, the current decoded picture is marked as "used for short-term reference".
[0240] The decoding process for picture order count is discussed.
[0241] The output of this process is PicOrderCntVal, the picture order count of the current picture.
[0242] Each coded picture is associated with a picture order count variable indicated as PicOrderCntVal.
[0243] If the current picture is 1 an IRAP picture with NoIncorrectPicOutputFlag equal to Or a GDR picture with a NoIncorrectPicOutputFlag equal to 1 otherwise, the variables prevPicOrderCntLsb and prevPicOrderCntMsb are derived as follows.
[0244] · Assume prevTid0Pic is the picture immediately preceding in decode order that has a TemporalId equal to 0 and is not a RASL or RADL picture.
[0245] · The variable prevPicOrderCntLsb is set equal to the tile_group_pic_order_cnt_lsb of prevTid0Pic.
[0246] · The variable prevPicOrderCntMsb is set equal to the PicOrderCntMsb of prevTid0Pic.
[0247] The variable PicOrderCntMsb of the current picture is derived as follows.
[0248] · If the current picture is an IRAP picture NoIncorrectPicOutputFlag with a value equal to 1 Or a GDR picture with a NoIncorrectPicOutputFlag equal to 1 then PicOrderCntMsb is set equal to 0.
[0249] · Otherwise, PicOrderCntMsb is derived as follows.
Table 7
[0250] Note 1 - All IRAP pictures NoIncorrectPicOutputFlag with a value equal to 1 have a PicOrderCntVal equal to tile_group_pic_order_cnt_lsb. This is because for IRAP pictures with a NoIncorrectPicOutputFlag equal to 1, PicOrderCntMsb is set equal to 0.
[0251] Note 1 - All GDR pictures with a NoIncorrectPicOutputFlag equal to 1 have a PicOrderCntVal equal to tile_group_pic_order_cnt_lsb. This is because for GDR pictures with a NoIncorrectPicOutputFlag equal to 1, PicOrderCntMsb is set equal to 0.
[0252] The value of PicOrderCntVal ranges from -2 31 to 2 31 - 1 (inclusive at both ends).
[0253] If the current picture is a GDR picture, the value of LastGDRPocVal is set equal to PicOrderCntVal.
[0254] The decoding process for the refresh boundary position of the picture is discussed.
[0255] This process is called only when gdr_enabled_flag is equal to 1.
[0256] This process is called after the parsing of the tile group header is completed.
[0257] The output of this process is the boundary positions of the refresh region of the current picture, namely PicRefreshedLeftBoundaryPos, PicRefreshedRightBoundaryPos, PicRefreshedTopBoundaryPos, and PicRefreshedBotBoundaryPos.
[0258] Each coded picture is associated with a set of refresh region boundary position variables indicated as PicOrderCntVal.
[0259] PicRefreshedLeftBoundaryPos, PicRefreshedRightBoundaryPos, PicRefreshedTopBoundaryPos, and PicRefreshedBotBoundaryPos are derived as follows.
[0260] If the tile group is the first received tile group of the current picture with a refreshed_region_flag equal to 1, the following applies: PicRefreshedLeftBoundaryPos = TGRefreshedLeftBoundary PicRefreshedRightBoundaryPos = TGRefreshedRightBoundary PicRefreshedTopBoundaryPos = TGRefreshedTopBoundary PicRefreshedBotBoundaryPos = TileGroupBotBoundary Otherwise, if refreshed_region_flag is equal to 1, the following applies. PicRefreshedLeftBoundaryPos = TGRefreshedLeftBoundary < PicRefreshedLeftBoundaryPos? TGRefreshedLeftBoundary : PicRefreshedLeftBoundaryPos PicRefreshedRightBoundaryPos = TGRefreshedRightBoundary > PicRefreshedRightBoundaryPos? TGRefreshedRightBoundary : PicRefreshedRightBoundaryPos PicRefreshedTopBoundaryPos = TGRefreshedTopBoundary < PicRefreshedTopBoundaryPos? TGRefreshedTopBoundary : RefreshedRegionTopBoundaryPos PicRefreshedBotBoundaryPos = TileGroupBotBoundary > PicRefreshedBotBoundaryPos? TileGroupBotBoundary : PicRefreshedBotBoundaryPos
[0261] The decoding process for constructing a reference picture list is discussed. …
[0262] equal to 1 NoIncorrectPicOutputFlag IRAP picture having or a GDR picture with NoIncorrectPicOutputFlag equal to 1 For each current picture that is not an IRAP picture having a NoIncorrectPicOutputFlag equal to 1, it is a bitstream compliance requirement that the value of maxPicOrderCnt - minPicOrderCnt be less than MaxPicOrderCntLsb / 2. …
[0263] The decoding process for reference picture marking … The current picture is an IRAP picture having a NoIncorrectPicOutputFlag equal to 1or a GDR picture with NoIncorrectPicOutputFlag equal to 1 If so, all reference pictures (if any) currently in the DPB are marked as "not used for reference". …
[0264] The derivation process for the temporal luma motion vector prediction is discussed. … The variable currCb specifies the current luma coding block at the luma position (xCb, yCb).
[0265] The variables mvLXCol and availableFlagLXCol are derived as follows: · If tile_group_temporal_mvp_enabled_flag is equal to 0, both components of mvLXCol are set equal to 0 and availableFlagLXCol is set equal to 0. · Otherwise (tile_group_temporal_mvp_enabled_flag is equal to 1), the following ordered steps are applied: 1. The co-located motion vector at the bottom right is derived as follows:
Table 8
[0266] · If yCb >> CtbLog2SizeY is equal to yColBr >> CtbLog2Size, yColBr is in the range from topBoundaryPos to botBoundaryPos (both ends inclusive), and xColBr is in the range from leftBoundaryPos to rightBoundaryPos (both ends inclusive), the following applies.
[0267] · The variable colCb specifies the luma coding block covering the modified position given by ((xColBr >> 3) << 3, (yColBr >> 3) << 3) within the co-located picture specified by ColPic.
[0268] · The luma position (xColCb, yColCb) is set equal to the top-left sample of the luma coding block at the co-location specified by colCb with respect to the top-left luma sample of the picture at the co-location specified by ColPic.
[0269] · The derivation process for the co-located motion vectors specified in clause 8.5.2.12 is called with currCb, colCb, (xColCb, yColCb), refIdxLX, and sbFlag (set to 0) as inputs, and the outputs are assigned to mvLXCol and availableFlagLXCol.
[0270] · Otherwise, both components of mvLXCol are set equal to 0, and availableFlagLXCol is set equal to 0.
[0271] 2. … The luma sample bilinear interpolation process is discussed.
[0272] The inputs to this process are as follows: · The luma position in full sample units (xInt L , yInt L ) · The luma position in fractional sample units (xFrac L , yFrac L ) · The luma reference sample array refPicLX L · The refresh region boundaries PicRefreshedLeftBoundaryPos, PicRefreshedTopBoundaryPos, PicRefreshedRightBoundaryPos, and PicRefreshedBotBoundaryPos of the reference picture …
[0273] The luma position in full sample units (xInt i , yInt i ) is derived as follows for i = 0..1:
Table 9
[0274] The luma sample 8-tap interpolation filtering process is discussed.
[0275] The inputs to this process are as follows: · The luma position in full sample units (xInt L , yInt L ) · The luma position in fractional sample units (xFrac L , yFrac L ) · The luma reference sample array refPicLX L , · A list padVal[dir] with dir = 0, 1 specifying the reference sample padding direction and amount · The refresh region boundaries PicRefreshedLeftBoundaryPos, PicRefreshedTopBoundaryPos, PicRefreshedRightBoundaryPos, and PicRefreshedBotBoundaryPos of the reference picture …
[0276] The luma position in full sample units (xInt i , yInt i ) is derived as follows for i = 0..7.
Table 10
[0277] The chroma sample interpolation process is discussed.
[0278] The inputs to this process are as follows: · The chroma position in full sample units (xInt C , yInt C ) · The chroma position in fractional sample units of 1 / 32 (xFrac C , yFrac C ) · The chroma reference sample array refPicLX C · The refresh region boundaries PicRefreshedLeftBoundaryPos, PicRefreshedTopBoundaryPos, PicRefreshedRightBoundaryPos, and PicRefreshedBotBoundaryPos of the reference picture …
[0279] The variable xOffset is set equal to ((sps_ref_wraparound_offset_minus1 + 1)*MinCbSizeY) / SubWidthC.
[0280] The chroma position (xInt i , yInt i ) in full sample units is derived as follows for i = 0..3:
Table 11
[0281] The deblocking filter process is discussed.
[0282] General process …
[0283] The deblocking filter process is applied to all coding sub-block edges and transform block edges of the picture, except for the following types of edges: · Edges at the picture boundary · Edges that coincide with the upper boundary of tile group tgA, if all of the following are satisfied: · gdr_enabled_flag is equal to 1 · loop_filter_across_refreshed_region_enabled_flag is equal to 0 · The edge coincides with the lower boundary of tile group tgB and the value of the refreshed_region_flag of tgB is different from the value of the refreshed_region_flag of tgA · Edges that coincide with the left boundary of tile group tgA, if all of the following are satisfied: · gdr_enabled_flag is equal to 1 · loop_filter_across_refreshed_region_enabled_flag is equal to 0 ·When the edge coincides with the right boundary of tile group tgB and the value of the refreshed_region_flag of tgB is different from the value of the refreshed_region_flag of tgA ·Edges that coincide with tile boundaries when loop_filter_cross_tiles_enabled_flag is equal to 0 ·Edges that coincide with the upper or left boundary of a tile group with a tile_group_loop_filter_across_tile_groups_enabled_flag equal to 0 or a tile_group_deblocking_filter_disabled_flag equal to 1 ·Edges within tile groups with a tile_group_deblocking_filter_disabled_flag equal to 1 ·Edges that do not correspond to the 8×8 sample grid boundaries of the component under consideration ·Edges within the chroma component where both sides of the edge use inter prediction ·Edges of chroma transform blocks that are not edges of the associated transform unit ·Edges that cross the luma transform blocks of coding units with an IntraSubPartitionsSplit value not equal to ISP_NO_SPLIT
[0284] The deblocking filter process for a certain direction is discussed … For each coding unit with coding block width log2CbW, coding block height log2CbH, and the position (xCb, yCb) of the top-left sample of the coding block, when edgeType is equal to EDGE_VER and xCb % 8 is equal to 0, or when edgeType is equal to EDGE_HOR and yCb % 8 is equal to 0, the edge is filtered by the following ordered steps
[0285] 1. The coding block width nCbW is set to 1 << log2CbW, and the coding block height nCbH is set to 1 << log2CbH.
[0286] 2. The variable filterEdgeFlag is derived as follows. · When edgeType is equal to EDGE_VER and one or more of the following conditions are true, filterEdgeFlag is set to 0: · The left boundary of the current coding block is the left boundary of the picture. · The left boundary of the current coding block is the left boundary of the tile and loop_filter_across_tiles_enabled_flag is equal to 0. · The left boundary of the current coding block is the left boundary of the tile group and tile_group_loop_filter_across_tile_groups_enabled_flag is equal to 0. · The left boundary of the current coding block is the left boundary of the current tile group, and all of the following conditions are met: · gdr_enabled_flag is equal to 1 · The loop_filter_across_refreshed_region_enabled_flag is equal to 0 · There exists a tile group that shares a boundary with the left boundary of the current tile group, and the value of its refreshed_region_flag is different from the value of the refreshed_region_flag of the current tile group.
[0287] · Otherwise, when edgeType is equal to EDGE_HOR and one or more of the following conditions are true, the variable filterEdgeFlag is set to 0: · The upper boundary of the current coding block is the upper boundary of the picture. · The upper boundary of the current coding block is the upper boundary of the tile and loop_filter_across_tiles_enabled_flag is equal to 0. · The upper boundary of the current coding block is the upper boundary of the tile group and tile_group_loop_filter_cross_tile_groups_enabled_flag is equal to 0. · The upper boundary of the current coding block is the upper boundary of the current tile group, and all of the following conditions are satisfied: · The gdr_enabled_flag is equal to 1. · The loop_filter_across_refreshed_region_enabled_flag is equal to 0. · There exists a tile group that shares a boundary with the upper boundary of the current tile group, and the value of its refreshed_region_flag is different from the value of the refreshed_region_flag of the current tile group.
[0288] Otherwise, the filterEdgeFlag is set to 1.
[0289] Once the tile is integrated, adapt the syntax.
[0290] 3. All elements of the two-dimensional (nCbW)×(nCbH) array edgeFlags are initialized to be equal to zero.
[0291] The CTB modification process for SAO is discussed. …
[0292] For all sample positions (xS i ,yS j ) and (xY i ,yY j ) where i = 0..nCtbSw-1 and j = 0..nCtbSh-1, depending on the values of pcm_loop_filter_disabled_flag, pcm_flag[xY i [yY j and cu_transquant_bypass_flag of the coding unit containing the coding block covering recPicture[xSi][ySj], the following applies: ·…
[0293] Modify the highlighted section depending on future decision transformation / quantization bypass. · Otherwise, when SaoTypeIdx[cIdx][rx][ry] is equal to 2, the following ordered steps apply: 1. The values of hPos[k] and vPos[k] for k = 0..1 are defined in Table 8-18 based on SaoEoClass[cIdx][rx][ry]. 2. The variable edgeIdx is derived as follows: · The modified sample position (xS ik' ,yS jk'and (xY ik' , yY jk' ) is derived as follows: (xS ik' , yS jk' ) = (xS i + hPos[k], ySj + vPos[k]) (8 - 1128) (xY ik' , yY jk' ) = (cIdx == 0)? (xS ik' , yS jk' ) : (xS ik' * SubWidthC, yS jk' * SubHeightC) (8 - 1129) ·For all sample positions (xS ik' , yS jk' ) and (xY ik' , yY jk' ) where one or more of the following conditions are true, edgeIdx is set equal to 0. ·The sample at position (xS ik' , yS jk' ) is outside the picture boundary. · The gdr_enabled_flag is equal to 1, the loop_filter_across_refreshed_region_enabled_flag is equal to 0, the refreshed_region_flag of the current tile group is equal to 1, and the refreshed_region_flag of the tile group containing the sample at position (xS ik' , yS jk' ) is equal to 0. ·The sample at position (xS ik' , yS jk' ) belongs to a different tile group and one of the following two conditions is true: ·MinTbAddrZs[xY ik' >> MinTbLog2SizeY][yY jk' >> MinTbLog2SizeY] is less than MinTbAddrZs[xY i >> MinTbLog2SizeY][yY j >> MinTbLog2SizeY], and tile_group_loop_filter_across_tile_groups_enabled_flag in the tile group to which sample recPicture[xS i [yS j belongs is equal to 0. ·MinTbAddrZs[xY i >>MinTbLog2SizeY][yY j >>MinTbLog2SizeY] is MinTbAddrZs[xY ik' >>MinTbLog2SizeY][yY jk' >>MinTbLog2SizeY] is smaller than, and the sample recPicture[xS ik' [yS jk' belongs to a tile group where the tile_group_loop_filter_across_tile_groups_enabled_flag is equal to 0. ·The loop_filter_across_tiles_enabled_flag is equal to 0, and the sample at position (xS ik' ,yS jk' ) belongs to another tile.
[0294] If a tile without a tile group is incorporated, correct the highlighted section. ·Otherwise, edgeIdx is derived as follows: ·The following applies: edgeIdx = 2 + Sign(recPicture[xS i [yS j - recPicture[xS i + hPos[0]][yS j + vPos[0]]) + Sign(recPicture[xS i [yS j - recPicture[xS i + hPos[1]][yS j + vPos[1]]) (8 - 1130) ·If edgeIdx is equal to 0, 1, or 2, edgeIdx is corrected as follows: edgeIdx = (edgeIdx == 2)? 0 : (edgeIdx + 1) (8 - 1131)
[0295] 3. The corrected picture sample array saoPicture[xS i [yS j is derived as follows.
[0296] saoPicture[xS i [yS j =Clip3(0,(1<<bitDepth)-1,recPicture[xS i [yS j + SaoOffsetVal[cIdx][rx][ry][edgeIdx]) (8-1132)
[0297] The coding tree block filtering process for luma samples for ALF is discussed. …
[0298] For the derivation of the filtered reconstructed luma sample alfPictureL[x][y], each reconstructed luma sample within the current luma coding tree block recPictureL[x][y] is filtered as follows for x,y=0..CtbSizeY-1: ·… · For each corresponding luma sample (x,y) in a given array of luma samples recPicture, the position (h x ,v y ) is derived as follows: · When the gdr_enabled_flag is equal to 1, the loop_filter_across_refreshed_region_enabled_flag is equal to 0, and the refreshed_region_flag of the tile group tgA containing the luma sample at position (x, y) is equal to 1, the following applies: · At position (h x ,h y ) is located in another tile group tgB and the refreshed_region_flag of tgB is equal to 0, the variables leftBoundary, rightBoundary, topBoundary, and botBoundary are set equal to TGRefreshedLeftBoundary, TGRefreshedRightBoundary, TGRefreshedTopBoundary, and TGRefreshedBotBoundary, respectively. · Otherwise, the variables leftBoundary, rightBoundary, topBoundary, and botBoundary are set equal to PicRefreshedLeftBoundaryPos, PicRefreshedRightBoundaryPos, PicRefreshedTopBoundaryPos, and PicRefreshedBotBoundaryPos, respectively. h x =Clip3(leftBoundary,rightBoundary,xCtb+x) (8-1140) v y =Clip3(topBoundary,botBoundary,yCtb+y) (8-1141) · Otherwise, the following applies: h x=Clip3(0,pic_width_in_luma_samples-1,xCtb+x) (8-1140) v y =Clip3(0,pic_height_in_luma_samples-1,yCtb+y) (8-1141) ·…
[0299] The ALF transposition for luma samples and the derivation process for filter indices are discussed. …
[0300] For each corresponding luma sample (x, y) in a given array recPicture of luma samples, the position (h x ,v y ) is derived as follows: · When gdr_enabled_flag is equal to 1, loop_filter_across_refreshed_region_enabled_flag is equal to 0, and the refreshed_region_flag of the tile group tgA containing the luma sample at position (x,y) is equal to 1, the following applies: · Position (h x ,h y ) is located in another tile group tgB and the refreshed_region_flag of tgB is equal to 0, the variables leftBoundary, rightBoundary, topBoundary, and botBoundary are set equal to TGRefreshedLeftBoundary, TGRefreshedRightBoundary, TGRefreshedTopBoundary, and TGRefreshedBotBoundary, respectively. · Otherwise, the variables leftBoundary, rightBoundary, topBoundary, and botBoundary are set equal to PicRefreshedLeftBoundaryPos, PicRefreshedRightBoundaryPos, PicRefreshedTopBoundaryPos, and PicRefreshedBotBoundaryPos, respectively. h x =Clip3(leftBoundary,rightBoundary,x) (8-1140) v y =Clip3(topBoundary,botBoundary,y) (8-1141) · Otherwise, the following applies: h x =Clip3(0,pic_width_in_luma_samples-1,x) (8-1145) v y =Clip3(0,pic_height_in_luma_samples-1,y) (8-1146) The coding tree block filtering process for chroma samples is discussed. …
[0301] For the derivation of the filtered and reconstructed chroma sample alfPicture[x][y], each reconstructed chroma sample within the current chroma coding tree block recPicture[x][y] is filtered as follows, with x,y = 0..ctbSizeC-1: · For each position (h x ,v y ) of the corresponding chroma sample (x,y) in a given array recPicture of chroma samples, it is derived as follows: · When gdr_enabled_flag is equal to 1, loop_filter_across_refreshed_region_enabled_flag is equal to 0, and the refreshed_region_flag of tile group tgA containing the luma sample at position (x,y) is equal to 1, the following applies: · The position (h x ,h y ) is located in another tile group tgB and the refreshed_region_flag of tgB is equal to 0, the variables leftBoundary, rightBoundary, topBoundary, and botBoundary are set equal to TGRefreshedLeftBoundary, TGRefreshedRightBoundary, TGRefreshedTopBoundary, and TGRefreshedBotBoundary, respectively. · Otherwise, the variables leftBoundary, rightBoundary, topBoundary, and botBoundary are set equal to PicRefreshedLeftBoundaryPos, PicRefreshedRightBoundaryPos, PicRefreshedTopBoundaryPos, and PicRefreshedBotBoundaryPos, respectively. h x =Clip3(leftBoundary / SubWidthC,rightBoundary / SubWidthC,xCtbC+x) (8-1140) v y =Clip3(topBoundary / SubWidthC,botBoundary / SubWidthC,yCtbC+y) (8-1141) · Otherwise, the following applies: h x =Clip3(0,pic_width_in_luma_samples / SubWidthC-1,xCtb+x) (8-1177) v y =Clip3(0,pic_height_in_luma_samples / SubHeightC-1,yCtb+y) (8-1178)
[0302] FIG. 7 shows a video bitstream 750 configured to implement a progressive decode - refresh (GDR) technique 700 according to an embodiment of the present disclosure. The GDR technique 700 may be similar to the GDR technique 500 of FIG. 5. As used herein, the video bitstream 750 may be referred to as a coded video bitstream, a bitstream, or a variation thereof. As shown in FIG. 7, the bitstream 750 includes a sequence parameter set (SPS) 752, a picture parameter set (PPS) 754, a slice header 756, and picture data 758.
[0303] The SPS 752 includes data that is common to all pictures in a sequence of pictures (SOP). In contrast, the PPS 754 includes data that is common to an entire picture. The slice header 756 includes information about the current slice, such as the slice type, which reference picture is used, etc. The SPS 752 and PPS 754 may generally be referred to as parameter sets. The SPS 752, PPS 754, and slice header 756 are of the type of network abstraction layer (NAL) units. An NAL unit is a syntax structure that includes an indication of the type of data to follow (e.g., coded video data). NAL units are classified into video coding layer (VCL) and non - VCL NAL units. A VCL NAL unit includes data representing the values of samples within a video picture, and a non - VCL NAL unit includes any relevant additional information such as parameter sets (important header data applicable to a number of VCL NAL units) and supplementary enhancement information (data that may enhance the usefulness of the decoded video signal but is not necessary for decoding the timing information and the values of samples within the video picture). Those skilled in the art will understand that the bitstream 750 may include other parameters and information in an actual application.
[0304] The image data 758 in FIG. 7 includes data related to an image or video to be encoded or decoded. The image data 758 may simply be referred to as the payload or data carried within the bitstream 750. In certain embodiments, the image data 758 includes a CVS 708 that includes a GDR picture 702, one or more subsequent pictures 704, and a recovery point picture 706. In certain embodiments, the GDR picture 702, subsequent pictures 704, and recovery point picture 706 may define a GDR period within the CVS 708.
[0305] As shown in FIG. 7, the GDR technique 700 or principle operates on a series of pictures that begin with a GDR picture 702 and end with a recovery point picture 706. The GDR picture 702 includes a refresh / clean area 710 that includes blocks that are all coded using intra prediction (i.e., intra-predicted blocks), and an un-refreshed / dirty area 712 that includes blocks that are all coded using inter prediction (i.e., inter-predicted blocks).
[0306] The subsequent picture 704 immediately adjacent to the GDR picture 702 includes a refresh / clean area 710 having a first portion 710A coded using intra prediction and a second portion 710B coded using inter prediction. The second portion 710B is coded, for example, by referring to the refresh / clean area 710 of a preceding picture within the GDR period of the CVS 708. As shown in the figure, the refresh / clean area 710 of the subsequent picture 704 expands as the coding process moves or progresses in a consistent direction (e.g., from left to right), and correspondingly, the unrefreshed / dirty area 712 is shrunk. Finally, a recovery point picture 706 including only the refresh / clean area 710 is obtained from the coding process. The second portion 710B of the refresh / clean area 710 coded as an inter prediction block may simply refer to the refresh area / clean area 710 in the reference picture.
[0307] As shown in FIG. 7, the GDR picture 702, the subsequent picture 704, and the recovery point picture 706 within the CVS 708 are each included within their own VCL NAL unit 730. The set of VCL NAL units 730 within the CVS 708 may be referred to as an access unit.
[0308] The NAL unit 730 containing the GDR picture 702 within the CVS 708 has a GDR NAL unit type (GDR_NUT). That is, in certain embodiments, the NAL unit 730 containing the GDR picture 702 within the CVS 708 has its own unique NAL unit type with respect to subsequent pictures 704 and recovery point pictures 706. In certain embodiments, the GDR_NUT allows the bitstream to start with a GDR picture 702 rather than requiring the bitstream to start with an IRAP picture. Designating the VCL NAL unit 730 of the GDR picture 702 as the GDR_NUT can indicate to, for example, a decoder that the first VCL NAL unit 730 in the CVS 708 contains the GDR picture 702.
[0309] In certain embodiments, the GDR picture 702 is the first picture in the CVS 708. In certain embodiments, the GDR picture 702 is the first picture in the GDR period. In certain embodiments, the GDR picture 702 has a temporal identifier (ID) equal to zero. The temporal ID is a value or number that identifies the position or order of a picture relative to other pictures. In certain embodiments, an access unit containing the VCL NAL unit 730 having the GDR_NUT is designated as a GDR access unit. In certain embodiments, the GDR picture 702 is a code slice of another (e.g., larger) GDR picture. That is, the GDR picture 702 may be part of a larger GDR picture.
[0310] FIG. 8 is an embodiment of a method 800 for decoding a coded video bitstream implemented by a video decoder (e.g., video decoder 30). Method 800 may be executed after the decoded bitstream has been received directly or indirectly from a video encoder (e.g., video encoder 20). Since this method allows for progressive intra refresh to enable random access without the need to use IRAP pictures, method 800 improves the decoding process. By using GDR pictures instead of IRAP pictures, for example, due to the size of the GDR picture compared to the size of the IRAP picture, a smoother and more consistent bitrate may be achieved, which allows for a reduced end-to-end delay (i.e., latency). Thus, as a practical matter, the codec performance is improved, which leads to a better user experience.
[0311] In block 802, the video decoder determines that the coded video sequence (CVS) of the coded video bitstream includes a video coding layer (VCL) network abstraction layer (NAL) unit having a gradual decoding refresh (GDR) network abstraction layer (NAL) unit type (GDR_NUT). The VCL NAL unit having the GDR_NUT includes a GDR picture.
[0312] In some embodiments, the GDR picture is the first picture in the CVS. In some embodiments, the GDR picture is the first picture in the GDR period. In some embodiments, the GDR picture has a temporal identifier (ID) equal to zero. A zero temporal ID may indicate that the GDR picture is, for example, the first picture in the CVS or GDR period. In some embodiments, an access unit containing a VCL NAL unit with GDR_NUT is designated as a GDR access unit. In some embodiments, the GDR picture is a coded slice of another GDR picture.
[0313] In block 804, the video decoder begins decoding the CVS in the GDR picture.
[0314] In block 806, the video decoder generates an image according to the decoded CVS. The image may then be displayed for a user of an electronic device (e.g., smartphone, tablet, laptop, personal computer, etc.).
[0315] FIG. 9 is an embodiment of a method 900 for encoding a video bitstream implemented by a video encoder (e.g., video encoder 20). Method 900 may be performed when a picture (e.g., from a video) is encoded into a video bitstream and then transmitted towards a video decoder (e.g., video decoder 30). Method 900 allows for progressive intra refresh to enable random access without the need to use IRAP pictures, thus improving the encoding process. By using GDR pictures instead of IRAP pictures, for example, due to the size of the GDR pictures compared to the size of the IRAP pictures, a smoother and more consistent bitrate may be achieved, which allows for a reduced end-to-end delay (i.e., latency). Thus, as a practical matter, the codec performance is improved, which leads to a better user experience.
[0316] In block 902, the video encoder determines a random access point for the video sequence. In block 904, the video encoder encodes a gradual decoding refresh (GDR) picture into a video coding layer (VCL) network abstraction layer (NAL) unit having a gradual decoding refresh (GDR) network abstraction layer (NAL) unit type (GDR_NUT) at the random access point for the video sequence.
[0317] In some embodiments, the GDR picture is the first picture in the CVS. In some embodiments, the GDR picture is the first picture in the GDR period. In some embodiments, the GDR picture has a temporal identifier (ID) equal to zero. A zero temporal ID may indicate that the GDR picture is, for example, the first picture in the CVS or the GDR period. In some embodiments, an access unit containing a VCL NAL unit with GDR_NUT is designated as a GDR access unit. In some embodiments, the GDR picture is a coded slice of another GDR picture.
[0318] In block 906, the video encoder generates a bitstream including the video sequence having the GDR picture in the VCL NAL unit with the GDR_NUT at the random access point.
[0319] In block 908, the bitstream is stored for transmission to a video decoder. The video bitstream is also referred to as a coded video bitstream or an encoded video bitstream. The video encoder can transmit the bitstream to the video decoder. Once received by the video decoder, the encoded video bitstream can be decoded (as described above, for example) to generate or produce an image for display to a user on a display or screen of an electronic device (such as a smartphone, tablet, laptop, personal computer, etc.).
[0320] FIG. 10 is a schematic diagram of a video coding apparatus 1000 (e.g., video encoder 20 or video decoder 30) according to an embodiment of the present disclosure. The video coding apparatus 1000 is suitable for implementing the disclosed embodiments as described herein. The video coding apparatus 1000 includes an input port 1010 and a receiver unit (Rx) 1020 for receiving data, a processor, logic unit, or central processing unit (CPU) 1030 for processing the data, a transmitter unit (Tx) 1040 and an output port 1050 for transmitting the data, and a memory 1060 for storing the data. The video coding apparatus 1000 may also include optical-to-electrical (OE) components and electrical-to-optical (EO) components coupled to the input port 1010, the receiver unit 1020, the transmitter unit 1040, and the output port 1050 for the input and output of optical or electrical signals.
[0321] Processor 1030 is implemented by hardware and software. Processor 1030 may be implemented as one or more CPU chips, cores (e.g., multi-core processors), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and digital signal processors (DSPs). Processor 1030 communicates with an input port 1010, a receiver unit 1020, a transmitter unit 1040, an output port 1050, and a memory 1060. Processor 1030 has a coding module 1070. The coding module 1070 implements the disclosed embodiments described above. For example, the coding module 1070 implements, processes, prepares, or provides various codec functions. Thus, by including the coding module 1070, the functions of the video coding device 1000 are substantially improved, and the video coding device 1000 is converted to different states. Alternatively, the coding module 1070 may be implemented as instructions stored in the memory 1060 and executed by the processor 1030.
[0322] The video coding device 1000 may also include an input and / or output (I / O) device 1080 for communicating data with a user. The I / O device 1080 may include output devices such as a display for displaying video data, a speaker for outputting audio data, etc. The I / O device 1080 may also include input devices such as a keyboard, a mouse, a trackball, etc., and / or corresponding interfaces for interacting with such output devices.
[0323] Memory 1060 includes one or more disks, tape drives, and solid state drives and may be used as an overflow data storage device to store a program when the program is selected for execution and to store instructions and data read during the execution of the program. Memory 1060 may be volatile and / or non-volatile and may be read only memory (ROM), random access memory (RAM), ternary associative memory (TCAM), and / or static random access memory (SRAM).
[0324] FIG. 11 is a schematic diagram of an embodiment of coding means 1100. In one embodiment, coding means 1100 is implemented in a video coding device 1102 (e.g., video encoder 20 or video decoder 30). Video coding device 1102 includes receiving means 1101. Receiving means 1101 is configured to receive a picture to be encoded or a bitstream to be decoded. Video coding device 1102 includes transmitting means 1107 coupled to receiving means 1101. Transmitting means 1107 is configured to transmit a bitstream to a decoder or a decoded picture to a display means (e.g., one of I / O devices 1080).
[0325] The video coding device 1102 includes a storage means 1103. The storage means 1103 is coupled to at least one of the receiving means 1101 or the transmitting means 1107. The storage means 1103 is configured to store instructions. The video coding device 1102 also includes a processing means 1105. The processing means 1105 is coupled to the storage means 1103. The processing means 1105 is configured to execute the instructions stored in the storage means 1103 in order to execute the methods disclosed herein. Also, the steps of the exemplary methods described herein need not necessarily be executed in the order described, and the order of such method steps should be understood to be merely exemplary. Similarly, additional steps may be included in such methods, and certain steps may be omitted or combined in ways consistent with various embodiments of the present disclosure.
[0326] Although several embodiments have been provided in the present disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. The examples herein are considered to be illustrative and not restrictive, and the intention is not to be limited to the details given herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.
[0327] Furthermore, the techniques, systems, subsystems, and methods described and illustrated as discrete or separate in various embodiments may be combined with or integrated into other systems, modules, techniques, or methods without departing from the scope of the present disclosure. Other items shown or discussed as being coupled to, directly coupled to, or communicating with each other may be indirectly coupled or communicate through some interface, device, or intermediate component, either electrically, mechanically, or otherwise. Other examples of changes, substitutions, and modifications will be apparent to those skilled in the art and may be made without departing from the spirit and scope disclosed herein.
Claims
1. 1. An apparatus for storing and transmitting a bitstream, the apparatus comprising: a transmitter and a memory, the memory configured to store the bitstream, the transmitter configured to transmit the bitstream, the bitstream including a coded video sequence (CVS) including a video coding layer (VCL) network abstraction layer (NAL) unit, the VCL NAL unit having a gradual decode refresh (GDR) network abstraction layer (NAL) unit type (GDR_NUT) for the video sequence, the VCL NAL unit having the GDR_NUT including a coded GDR picture; an access unit that includes the plurality of VCL NAL units having the GDR_NUT is designated as a GDR access unit; Device.
2. 2. The apparatus of claim 1, wherein the memory stores the coded video stream prior to transmitting the coded video stream to a video decoder.
3. The apparatus of claim 2 , wherein the GDR picture is the first picture in the CVS.
4. The apparatus of claim 2 , wherein the GDR picture is a first picture in a GDR period.
5. The device according to claim 1 , wherein the GDR picture has a temporal identifier (ID) equal to zero.
6. The apparatus of claim 1 , wherein the GDR_NUT indicates to a video decoder that the VCL NAL units having the GDR_NUT include the coded GDR picture.
7. A system comprising an apparatus according to any one of claims 1 to 6 and a decoder for decoding said bitstream.
8. 1. A method of generating a bitstream, the method comprising: Coding a gradual decode refresh (GDR) picture; performing entropy coding to encode the GDR picture and a plurality of syntax elements into the bitstream, the coded GDR picture being included in a video coding layer (VCL) network abstraction layer (NAL) unit, the VCL NAL unit having a gradual decode refresh (GDR) network abstraction layer (NAL) unit type (GDR_NUT) for the video sequence, and an access unit including the VCL NAL unit having the GDR_NUT being designated a GDR access unit. method.
9. The method of claim 8 , wherein the GDR picture is the first picture in a coded video sequence (CVS).
10. The method of claim 9 , wherein the GDR picture is the first picture in a GDR period.
11. 11. The apparatus according to claim 8, wherein a temporal identifier (ID) associated with said GDR picture is equal to zero.
12. The encoding device according to claim 1 , wherein the GDR_NUT indicates to a video decoder that the VCL NAL unit having the GDR_NUT contains the coded GDR picture.
Citation Information
Patent Citations
Encoders, decoders and corresponding methods
JP7678049B2
Video encoding / decoding device, method, and program
WO2014002385A1
Signaling of regions of interest and gradual decoding refresh in video coding
WO2014051915A1
Gradual decoding refresh with temporal scalability support in video coding
WO2014107723A1