Output of the previous picture for a picture that starts a new coded video sequence in video coding
By employing a flag to empty the DPB before decoding non-IDR random access point pictures, the method addresses buffer overflow issues, ensuring continuous playback and improved user experience in video coding.
Patent Information
- Application Number
- JP2021565968
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-05-06
- Filing Date
- 2020-05-01
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2040-05-01
AI Technical Summary
Existing video coding technologies face challenges in managing decoded picture buffers (DPB) during random access points, leading to potential overflow and disrupted playback, especially with non-instantaneous decoder refresh (IDR) pictures, which affect user experience.
Implementing a method where a flag (e.g., no_output_of_prior_pics_flag) is used to empty the decoded picture buffer (DPB) before decoding a non-IDR random access point picture, such as Clean Random Access (CRA) or Gradual Decoding Refresh (GDR) pictures, ensuring continuous playback by preventing buffer overflow.
This approach enhances user experience by ensuring continuous and efficient playback by preventing DPB overflow and improving codec performance during random access points in video decoding.
Smart Images

Figure 0007708362000002 
Figure 0007708362000003 
Figure 0007708362000004
Abstract
Description
Technical Field
[0001] This patent application claims the benefit of U.S. Provisional Patent Application No. 62 / 843,991, filed May 6, 2019, by Ye-Kui Wang, which is hereby incorporated by reference herein in its entirety.
[0002] In general, the present disclosure describes techniques for supporting the output of previously decoded pictures in video coding. More specifically, the present disclosure enables the output of previously decoded pictures corresponding to random access point pictures that start a coded video sequence (CVS) from a decoded picture buffer (DPB).
Background Art
[0003] The amount of video data required to depict even relatively short videos is substantial, which can be problematic when the data is streamed or communicated over a communication network with limited bandwidth capacity. Thus, video data is generally compressed before being communicated over modern telecommunications networks. Also, the size of the video can be an issue when the video is stored on a storage device, since memory resources may be limited. Video compression devices often use software and / or hardware at the source to code the video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompression device that decodes the video data. Improved compression and decompression techniques that improve the compression ratio without sacrificing much in terms of image quality are desirable, since network resources are limited and the demand for higher video quality is constantly increasing.
Summary of the Invention
[0004] A first aspect relates to a decoding method implemented by a video decoder. This method includes receiving, by the video decoder, a coded video bitstream, where the coded video bitstream includes a hierarchical decoding refresh (GDR) picture and a first flag having a first value; setting, by the video decoder, a second value of a second flag equal to the first value of the first flag; emptying, by the video decoder, a picture from a decoded picture buffer (DPB) based on the second flag having the second value; and decoding, by the video decoder, a current picture after the DPB becomes empty. All that were decoded before GDR pictures This method provides a technique for output of a previous picture (e.g., a previously decoded picture) in a decoded picture buffer (DPB) when a random access point picture other than an instantaneous decoder refresh (IDR) picture (e.g., a clean random access (CRA) picture, a hierarchical random access (GRA) picture, or a hierarchical decoding refresh (GDR) picture, a CVSS picture, etc.) is encountered in decoding order. When reaching a random access point picture, emptying a previously decoded picture from the DPB prevents the DPB from overflowing and promotes more continuous playback. Thus, a coder / decoder (also known as a "codec") in video coding is improved compared to a current codec. In practical terms, an improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed.
[0005] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the GDR picture is not the first picture of the coded video bitstream.
[0006] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the GDR picture is not the first picture of the coded video bitstream.
[0007] Optionally, in any of the foregoing aspects, another implementation form of the aspect provides that the GDR picture is disposed in a video coding layer (VCL) network abstraction layer (NAL) unit having a group of decoded pictures (GDR) network abstraction layer (NAL) unit type (GDR_NUT).
[0008] Optionally, in any of the foregoing aspects, another implementation form of the aspect provides that the first flag is specified as no_output_of_pri o r_pics_flag, and the second flag is specified as NoOutputOfPriorPicsFlag.
[0009] Optionally, in any of the foregoing aspects, another implementation form of the aspect provides that when the first flag is set to a first value, the DPB fullness parameter is set to zero.
[0010] Optionally, in any of the foregoing aspects, another implementation form of the aspect provides that the DPB is emptied after the GDR picture is decoded.
[0011] Optionally, in any of the foregoing aspects, another implementation form of the aspect provides that an image generated based on the current picture is displayed.
[0012] A second aspect relates to an encoding method implemented by a video encoder. The method includes determining, by the video encoder, a random access point of a video sequence; encoding, by the video encoder, a group of decoded pictures (GDR) picture for the video sequence at the random access point; and, by the video encoder, from a decoded picture buffer (DPB) All that were decoded beforeSetting a flag to a first value that instructs a video decoder to clear a picture, and generating, by a video encoder, a video sequence having a GDR picture at a random access point and a video bitstream including the flag, and storing, by the video encoder, the video bitstream for transmission to the video decoder.
[0013] The method provides a technique for output of a previous picture (e.g., a previously decoded picture) in a decoded picture buffer (DPB) when a random access point picture other than an instantaneous decoder refresh (IDR) picture (e.g., a clean random access (CRA) picture, a gradual random access (GRA) picture, or a gradual decoding refresh (GDR) picture, a CVSS picture, etc.) is encountered in decoding order. When reaching a random access point picture, emptying the previously decoded picture from the DPB prevents the DPB from overflowing and promotes more continuous playback. Thus, a coder / decoder (also known as a "codec") in video coding is improved compared to the current codec. Practically, the improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed.
[0014] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the GDR picture is not the first picture of the video bitstream, and the video decoder is instructed to empty the DPB after the GDR picture is decoded.
[0015] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the GDR picture is disposed in a video coding layer (VCL) network abstraction layer (NAL) unit having a gradual decoding refresh (GDR) network abstraction layer (NAL) unit type (GDR_NUT).
[0016] Optionally, in any of the foregoing aspects, another implementation form of the aspect provides to instruct the video decoder to set the DPB fullness parameter to zero when the flag is set to the first value.
[0017] Optionally, in any of the foregoing aspects, another implementation form of the aspect provides that the flag is specified as no_output_of_prior_pics_flag.
[0018] Optionally, in any of the foregoing aspects, another implementation form of the aspect is that the first value of the flag is 1.
[0019] A third aspect relates to a decoding device. The decoding device includes a receiver configured to receive a coded video bitstream, a memory coupled to the receiver, the memory storing instructions, and a processor coupled to the memory. The processor causes the decoding device to receive a coded video bitstream that includes a hierarchical refresh (GDR) picture and a first flag having a first value, set a second value of a second flag equal to the first value of the first flag, empty the picture from a decoded picture buffer (DPB) based on the second flag having the second value, and, after the DPB is emptied, decode a current picture. All that were decoded before The processor is configured to execute instructions to cause the above actions.
[0020] This decoding device provides a technique for output of a previous picture (e.g., a previously decoded picture) in a decoded picture buffer (DPB) when a random access point picture other than an instantaneous decoder refresh (IDR) picture (e.g., a clean random access (CRA) picture, a gradual random access (GRA) picture, or a gradual decoding refresh (GDR) picture, a CVSS picture, etc.) is encountered in decoding order. When reaching a random access point picture, emptying the previously decoded picture from the DPB prevents the DPB from overflowing and promotes more continuous playback. Thus, a coder / decoder (also known as a "codec") in video coding is improved compared to current codecs. As a practical matter, the improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed.
[0021] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the GDR picture is not the first picture of the coded video bitstream.
[0022] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that a first flag is specified as no_output_of_pri o r_pics_flag and a second flag is specified as NoOutputOfPriorPicsFlag.
[0023] Optionally, in any of the foregoing aspects, another implementation of the aspect provides a display configured to display a picture generated based on a current picture.
[0024] A fourth aspect relates to an encoding device. The encoding device includes a memory containing instructions and a processor coupled to the memory. The processor is configured to cause the encoding device to determine a random access point of a video sequence, encode a gradual decode refresh (GDR) picture for the video sequence at the random access point, set a flag to a first value for instructing a video decoder to empty pictures from a decoded picture buffer (DPB), and generate a video bitstream including the video sequence having the GDR picture at the random access point and the flag. All that were decoded before The encoding device includes a processor configured to implement the instructions to perform the above operations, and a transmitter coupled to the processor. The transmitter is configured to transmit the video bitstream towards the video decoder.
[0025] This encoding device provides a technique for output of previous pictures (e.g., previously decoded pictures) in a decoded picture buffer (DPB) when a random access point picture other than an instantaneous decoder refresh (IDR) picture (e.g., a clean random access (CRA) picture, a gradual random access (GRA) picture, or a gradual decode refresh (GDR) picture, a CVSS picture, etc.) is encountered in decoding order. When reaching a random access point picture, emptying the previously decoded pictures from the DPB prevents the DPB from overflowing and promotes more continuous playback. Thus, the coder / decoder (also known as a "codec") in video coding is improved compared to current codecs. Practically, the improved video coding process provides a better user experience for the user when the video is sent, received, and / or viewed.
[0026] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the GDR picture is not the first picture of the coded video bitstream.
[0027] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the flag is specified as the no_output_of_prior_pics_flag.
[0028] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the memory stores the bitstream before the transmitter transmits the bitstream towards the video decoder.
[0029] The fifth aspect relates to a coding apparatus. The coding apparatus includes a receiver configured to receive a picture to be encoded or a bitstream to be decoded, a transmitter coupled to the receiver, where the transmitter is configured to transmit the bitstream to a decoder or transmit the decoded image to a display, at least one memory coupled to at least one of the receiver or the transmitter, where the memory is configured to store instructions, and a processor coupled to the memory, where the processor is configured to execute the instructions stored in the memory for performing any of the methods disclosed herein.
[0030] This coding device provides a technique for output of previous pictures (e.g., previously decoded pictures) in a decoded picture buffer (DPB) when a random access point picture other than an instantaneous decoder refresh (IDR) picture (e.g., clean random access (CRA) picture, gradual random access (GRA) picture, or gradual decode refresh (GDR) picture, CVSS picture, etc.) is encountered in decoding order. When reaching a random access point picture, emptying the previously decoded pictures from the DPB prevents the DPB from overflowing and promotes more continuous playback. Accordingly, a coder / decoder (also known as a “codec”) in video coding is improved compared to current codecs. As a practical matter, the improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed.
[0031] Optionally, in any of the foregoing aspects, another implementation of the aspect provides a display configured to display a picture.
[0032] A sixth aspect relates to a system. The system includes an encoder and a decoder that communicates with the encoder, and the encoder or the decoder includes a decoding device, an encoding device, or a coding device disclosed herein.
[0033] This system provides a technique for output of previous pictures (e.g., previously decoded pictures) in a decoded picture buffer (DPB) when random access point pictures other than instant decoder refresh (IDR) pictures (e.g., clean random access (CRA) pictures, gradual random access (GRA) pictures, or gradual decoder refresh (GDR) pictures, CVSS pictures, etc.) are encountered in decoding order. When reaching a random access point picture, emptying the previously decoded pictures from the DPB prevents the DPB from overflowing and promotes more continuous playback. Accordingly, a coder / decoder (also known as a “codec”) in video coding is improved compared to current codecs. As a practical matter, the improved video coding process provides a better user experience for the user when video is sent, received, and / or viewed.
[0034] A seventh aspect relates to coding means. The coding means includes receiving means configured to receive a picture to be coded or a bitstream to be decoded, transmitting means coupled to the receiving means, the transmitting means being configured to transmit the bitstream to decoding means or the decoded image to display means, storage means coupled to at least one of the receiving means or the transmitting means, the storage means being configured to store instructions, and processing means coupled to the storage means, the processing means being configured to execute instructions stored in the storage means for performing any of the methods disclosed herein.
[0035] The coding means provides a technique for output of the previous picture (e.g., the previously decoded picture) in the decoded picture buffer (DPB) when a random access point picture other than an instantaneous decoder refresh (IDR) picture (e.g., a clean random access (CRA) picture, a gradual random access (GRA) picture, or a gradual decode refresh (GDR) picture, a CVSS picture, etc.) is encountered in decoding order. When reaching a random access point picture, emptying the previously decoded picture from the DPB prevents the DPB from overflowing and promotes more continuous playback. Thus, the coder / decoder (also known as a "codec") in video coding is improved compared to the current codec. As a practical matter, the improved video coding process provides a better user experience for the user when the video is sent, received, and / or viewed.
[0036] For clarity, any one of the foregoing embodiments may be combined with any one or more of the other foregoing embodiments to create new embodiments within the scope of the present disclosure.
[0037] These and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and the claims.
Brief Description of the Drawings
[0038] To more fully understand the present disclosure, reference is made to the following brief description, taken in conjunction with the accompanying drawings and detailed description, in which like reference numerals represent like parts.
[0039]
Figure 1
[0040]
Figure 2
[0041]
Figure 3
[0042]
Figure 4
[0043]
Figure 5
[0044]
Figure 6
[0045]
Figure 7
[0046]
Figure 8
[0047]
Figure 9
[0048]
Figure 10
[0049]
Figure 11
DETAILED DESCRIPTION OF THE INVENTION
[0050] Initially, exemplary implementations of one or more embodiments are provided below, but it should be understood that the disclosed systems and / or methods may be implemented using any number of technologies, whether currently known or to exist. The present disclosure is in no way limited to the exemplary implementations, drawings, and technologies shown and described herein, including the exemplary designs and implementations, but may be modified within the scope of the appended claims, along with the full scope of their equivalents.
[0051] FIG. 1 is a block diagram showing an exemplary coding system 10 that may utilize the video coding techniques described herein. As shown in FIG. 1, the coding system 10 includes a source device 12 that provides encoded video data to be decoded later by a destination device 14. In particular, the source device 12 may provide the video data to the destination device 14 via a computer-readable medium 16. The source device 12 and the destination device 14 may include any of a wide range of devices, including desktop computers, notebook (e.g., laptop) computers, tablet computers, set-top boxes, so-called "smart" phone handsets, so-called "smart" pads, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, and the like. In some cases, the source device 12 and the destination device 14 may be equipped for wireless communication.
[0052] The destination device 14 may receive encoded video data that is decoded via a computer-readable medium 16. The computer-readable medium 16 may include any type of medium or device capable of moving the encoded video data from the source device 12 to the destination device 14. In one example, the computer-readable medium 16 may include a communication medium that enables the source device 12 to directly transmit the encoded video data to the destination device 14 in real time. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and transmitted to the destination device 14. The communication medium may include any wireless or wired communication medium such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or other devices useful for facilitating communication from the source device 12 to the destination device 14.
[0053] In some examples, the encoded data may be output from the output interface 22 to a storage device. Similarly, the encoded data may be accessed from the storage device by an input interface. The storage device may include any of various distributed or locally accessible data storage media such as a hard drive, a Blu-ray disk, a digital video disk (DVD), a compact disk read-only memory (CD-ROM), a flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In a further example, the storage device may correspond to a file server or another intermediate storage device that can store the encoded video generated by the source device 12. The destination device 14 may access the stored video data from the storage device via streaming or downloading. The file server may be any type of server capable of storing the encoded video data and transmitting the encoded video data to the destination device 14. Exemplary file servers include a web server (e.g., for a website), a File Transfer Protocol (FTP) server, a Network Attached Storage (NAS) device, or a local disk drive. The destination device 14 may access the encoded video data via any standard data connection including an Internet connection. This may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a Digital Subscriber Line (DSL), a cable modem, etc.), or a combination of both suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage device may be a streaming transmission, a download transmission, or a combination thereof.
[0054] The technology of the present disclosure is not necessarily limited to wireless applications or settings. This technology may be applied to video coding that supports any of various multimedia applications, such as wireless television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission such as dynamic adaptive streaming over HTTP (DASH), digital video encoded on a data storage medium, decoding of digital video stored on a data storage medium, or other applications. In some examples, the coding system 10 may be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.
[0055] In the example of FIG. 1, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The destination device 14 includes an input interface 28, a video decoder 30, and a display device 32. According to the present disclosure, the video encoder 20 of the source device 12 and / or the video decoder 30 of the destination device 14 may be configured to apply techniques for video coding. In other examples, the source device and the destination device may include other components or arrangements. For example, the source device 12 may receive video data from an external video source such as an external camera. Similarly, the destination device 14 may interface with an external display device rather than including an integrated display device.
[0056] The coding system 10 shown in FIG. 1 is merely an example. The techniques for video coding may be performed by any digital video encoding and / or decoding device. The techniques of the present disclosure are generally performed by a video coding device, but may typically be performed by a video encoder / decoder, commonly referred to as a "CODEC". Further, the techniques of the present disclosure may also be performed by a video preprocessor. The video encoder and / or decoder may be a graphics processing unit (GPU) or a similar device.
[0057] The source device 12 and the destination device 14 are merely examples of such coding devices that generate encoded video data for transmission from the source device 12 to the destination device 14. In some examples, the source device 12 and the destination device 14 may operate in a substantially symmetric manner such that each of the source device 12 and the destination device 14 includes video encoding and decoding components. Thus, the coding system 10 can support one-way or two-way video transmission between the video devices 12, 14, for example, for video streaming, video playback, video broadcast, or video telephony.
[0058] The video source 18 of the source device 12 may include a video capture device such as a video camera, a video archive containing previously captured video, and / or a video supply interface for receiving video from a video content provider. As a further alternative, the video source 18 may generate computer graphics-based data as source video, or as a combination of live video, archived video, and computer-generated video.
[0059] In some cases, when video source 18 is a video camera, source device 12 and destination device 14 may form a so-called camera phone or video phone. However, as described above, the techniques described in this disclosure may generally be applicable to video coding and may also be applicable to wireless and / or wired applications. In each case, the captured, pre-captured, or computer-generated video may be encoded by video encoder 20. The encoded video information may then be output by output interface 22 to computer-readable medium 16.
[0060] Computer-readable medium 16 may include a transient medium such as a wireless broadcast or wired network transmission, or a storage medium such as a hard disk, flash drive, compact disk, digital video disk, Blu-ray disk, or other computer-readable medium. In some examples, a network server (not shown) may receive the encoded video data from source device 12 and provide the encoded video data to destination device 14, for example, via a network transmission. Similarly, a computing device of a media manufacturing facility such as a disk stamping facility may receive the encoded video data from source device 12 and generate a disk containing the encoded video data. Thus, computer-readable medium 16 may be understood to include one or more computer-readable media of various forms in various examples.
[0061] The input interface 28 of the destination device 14 receives information from the computer-readable medium 16. The information on the computer-readable medium 16 may include syntax information defined by the video encoder 20, which is also used by the video decoder 30 and includes syntax elements that describe the characteristics and / or processing of blocks and / or other coded units, such as groups of pictures (GOPs). The display device 32 displays the decoded video data to the user and may include any of various display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.
[0062] The video encoder 20 and the video decoder 30 may operate according to a video coding standard, such as the currently developed High Efficiency Video Coding (HEVC) standard, and may conform to the HEVC Test Model (HM). Alternatively, the video encoder 20 and the video decoder 30 may operate according to other proprietary or industry standards, such as the H.264 standard of the International Telecommunication Union Telecommunication Standardization Sector (ITU-T), also known as Moving Picture Experts Group (MPEG)-4 Part 10, Advanced Video Coding (AVC), H265 / HEVC, or extensions of such standards. However, the technology of the present disclosure is not limited to any particular coding standard. Other examples of video coding standards include MPEG-2 and ITU-T H.263. Although not shown in FIG. 1, in some aspects, the video encoder 20 and the video decoder 30 may each be integrated with an audio encoder and decoder and may include an appropriate multiplexer-demultiplexer unit or other hardware and software to process the coding of both audio and video in a common data stream or separate data streams. When applicable, the MUX-DEMUX unit may conform to other protocols, such as the ITU H.223 multiplexer protocol, or the User Datagram Protocol (UDP).
[0063] Video encoder 20 and video decoder 30 may each be implemented as any of a variety of suitable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is implemented partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Each of video encoder 20 and video decoder 30 may be included in one or more encoders or decoders, and any of these may be integrated as part of an integrated encoder / decoder (CODEC) within their respective devices. Devices including video encoder 20 and / or video decoder 30 may include integrated circuits, microprocessors, and / or wireless communication devices such as cellular telephones.
[0064] FIG. 2 is a block diagram illustrating an example of video encoder 20 that may implement video coding techniques. Video encoder 20 may perform intra-coding and inter-coding of video blocks within a video slice. Intra-coding relies on spatial prediction to reduce or remove spatial redundancy in video within a given video frame or picture. Inter-coding relies on temporal prediction to reduce or remove temporal redundancy in video within adjacent frames or pictures of a video sequence. The intra mode (I mode) may refer to any of several spatial-based coding modes. Inter-modes, such as uni-directional (also known as single prediction) prediction (P mode) or bi-directional (bi-prediction) (B mode), may refer to any of several time-based coding modes.
[0065] As shown in FIG. 2, video encoder 20 receives the current video block within the video frame to be encoded. In the example of FIG. 2, video encoder 20 includes a mode selection unit 40, a reference frame memory 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. Mode selection unit 40, in turn, includes a motion compensation unit 44, a motion estimation unit 42, an intra prediction (also known as intra prediction) unit 46, and a partitioning unit 48. In video block reconstruction, video encoder 20 also includes an inverse quantization unit 58, an inverse transform unit 60, and an adder 62. Also, a deblocking filter (not shown in FIG. 2) may be included to filter the block boundaries in order to remove blocky artifacts from the reconstructed video. If desired, the deblocking filter typically filters the output of adder 62. In addition to the deblocking filter, additional filters (in-loop or post-loop) may also be used. Such filters are not shown for simplicity, but if desired, may filter the output of adder 50 (as an in-loop filter).
[0066] During the encoding process, video encoder 20 receives the video frame or slice to be coded. The frame or slice may be divided into a plurality of video blocks. Motion estimation unit 42 and motion compensation unit 44 perform inter-prediction coding of the received video block with respect to one or more blocks in one or more reference frames to provide temporal prediction. Alternatively, intra prediction unit 46 may perform intra-prediction coding of the received video block with respect to one or more adjacent blocks in the same frame or slice as the block to be coded to provide spatial prediction. Video encoder 20 may execute a plurality of coding passes to select an appropriate coding mode for, for example, each block of video data.
[0067] Furthermore, the partitioning unit 48 may partition blocks of video data into sub-blocks based on an evaluation of previous partitioning schemes in previous coding paths. For example, the partitioning unit 48 may first partition a frame or a slice into the largest coding units (LCUs), and then partition each LCU into sub-coding units based on rate distortion analysis (e.g., rate distortion optimization). Further, the mode selection unit 40 may generate a quad-tree data structure indicating the partitioning of an LCU into sub-CUs. A leaf node CU of the quad-tree may include one or more prediction units (PUs) and one or more transform units (TUs).
[0068] This disclosure uses the term "block" to refer to any of a CU, a PU, or a TU in the context of HEVC, or a similar data structure in the context of other standards (e.g., macroblocks and their sub-blocks in H.264 / AVC). A CU includes a coding node, and PUs and TUs associated with the coding node. The size of a CU corresponds to the size of the coding node and is square. The size of a CU may range from 8×8 pixels to the size of a tree block of 64×64 pixels or more. Each CU may include one or more PUs and one or more TUs. The syntax data associated with a CU may describe, for example, the partitioning of the CU into one or more PUs. The partitioning mode may vary depending on whether the CU is coded in skip or direct mode, intra prediction mode, or inter prediction (also known as inter prediction) mode. A PU may be partitioned into a non-square shape. The syntax data associated with a CU may also describe, for example, the partitioning of the CU into one or more TUs according to a quad-tree. A TU can be square or non-square (e.g., rectangular) in shape.
[0069] The mode selection unit 40 may select, for example, one of the intra or inter - coding modes based on an error result, provide the obtained intra - or inter - coded block to the adder 50, generate residual block data, provide it to the adder 62, and reconstruct the coded block for use as a reference frame. The mode selection unit 40 also provides syntax elements such as motion vectors, intra - mode indicators, partition information, and other such syntax information to the entropy coding unit 56.
[0070] The motion estimation unit 42 and the motion compensation unit 44 may be highly integrated but are shown separately for conceptual purposes. Motion estimation performed by the motion estimation unit 42 is a process of generating a motion vector, and the motion vector estimates the motion of a video block. The motion vector may indicate, for example, the displacement of a prediction block (or other coded unit) in a reference frame with respect to a current video block (or other coded unit) being coded within the current frame, within a current video frame or picture. The prediction block is a block that is found to closely match the block being coded with respect to pixel differences, which can be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or the sum of other difference metrics. In some examples, the video coder 20 may calculate the values of sub - integer pixel positions of a reference picture stored in the reference frame memory 64. For example, the video coder 20 may interpolate the values at 1 / 4 - pixel positions, 1 / 8 - pixel positions, or other fractional - pixel positions of the reference image. Thus, the motion estimation unit 42 may perform a motion search for all pixel positions and fractional - pixel positions and output a motion vector with fractional - pixel accuracy.
[0071] The motion estimation unit 42 calculates the motion vectors of the PUs of the video blocks within an inter-coded slice by comparing the positions of the PUs with the positions of the prediction blocks of the reference picture. The reference picture may be selected from the first reference picture list (list 0) or the second reference picture list (list 1), and each of the lists identifies one or more reference pictures stored in the reference frame memory 64. The motion estimation unit 42 transmits the calculated motion vectors to the entropy coding unit 56 and the motion compensation unit 44.
[0072] Motion compensation performed by the motion compensation unit 44 may involve fetching or generating a prediction block based on the motion vectors determined by the motion estimation unit 42. Also, in some examples, the motion estimation unit 42 and the motion compensation unit 44 may be functionally integrated. Upon receiving the motion vectors for the PUs of the current video block, the motion compensation unit 44 can identify the position of the prediction block pointed to by the motion vectors in one of the reference picture lists. The adder 50 forms a residual video block, as will be described later, by subtracting the pixel values of the prediction block from the pixel values of the encoded current video block, which forms pixel difference values. Generally, the motion estimation unit 42 performs motion estimation for the luma component, and the motion compensation unit 44 uses the motion vectors calculated based on the luma component for both the chroma component and the luma component. The mode selection unit 40 may also generate syntax elements related to the video blocks and the video slice for use by the video decoder 30 when decoding the video blocks of the video slice.
[0073] As described above, the intra prediction unit 46 may perform intra prediction on the current block as an alternative to inter prediction executed by the motion estimation unit 42 and the motion compensation unit 44. In particular, the intra prediction unit 46 may determine an intra prediction mode to be used for encoding the current block. In some examples, the intra prediction unit 46 may encode the current block using various intra prediction modes, for example, during separate encoding paths, and the intra prediction unit 46 (or, in some examples, the mode selection unit 40) may select an appropriate intra prediction mode to use from the tested modes.
[0074] For example, the intra prediction unit 46 may calculate rate distortion values using rate distortion analysis for various tested intra prediction modes, and select an intra prediction mode having the best rate distortion characteristics among the tested modes. Rate distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block that is encoded to generate the encoded block, and the bit rate (i.e., the number of bits) used to generate the encoded block. The intra prediction unit 46 may calculate a ratio from the distortion and rate for various encoded blocks to determine which intra prediction mode exhibits the best rate distortion value for the block.
[0075] Additionally, the intra prediction unit 46 may be configured to code depth blocks of a depth map using a depth modeling mode (DMM). The mode selection unit 40 may determine, for example, using rate distortion optimization (RDO), whether the available DMM modes result in better coding results than the intra prediction mode and other DMM modes. The data of the texture image corresponding to the depth map may be stored in the reference frame memory 64. The motion estimation unit 42 and the motion compensation unit 44 may also be configured to perform inter prediction on the depth blocks of the depth map.
[0076] After selecting an intra prediction mode for a block (e.g., one of a conventional intra prediction mode or a DMM mode), the intra prediction unit 46 may provide information indicating the selected intra prediction mode for the block to the entropy coding unit 56. The entropy coding unit 56 may encode the information indicating the selected intra prediction mode. The video encoder 20 may include a plurality of intra prediction mode index tables and a plurality of modified intra prediction mode index tables (also referred to as codeword mapping tables), definitions of coding contexts for various blocks, and the most likely intra prediction mode, intra prediction mode index table, and modified intra prediction mode index table used for each context in the transmitted bitstream setting data
[0077] The video encoder 20 forms a residual video block by subtracting the prediction data from the mode selection unit 40 from the coded original video block. The adder 50 represents the component or components that perform this subtraction operation.
[0078] The transform processing unit 52 applies a transform such as a discrete cosine transform (DCT) or a conceptually similar transform to the residual block to generate a video block including residual transform coefficient values. The transform processing unit 52 may perform other transforms conceptually similar to the DCT. A wavelet transform, integer transform, subband transform, or other type of transform may also be used.
[0079] The conversion processing unit 52 applies the conversion to the residual block and generates a block of residual conversion coefficients. The conversion may convert the residual information from the pixel value domain to a conversion domain, for example, the frequency domain. The conversion processing unit 52 may send the obtained conversion coefficients to the quantization unit 54. The quantization unit 54 quantizes the conversion coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting the quantization parameter. In some examples, the quantization unit 54 may then perform a scan of the matrix containing the quantized conversion coefficients. Alternatively, the entropy coding unit 56 may perform the scan.
[0080] After quantization, the entropy coding unit 56 entropy-codes the quantized conversion coefficients. For example, the entropy coding unit 56 may perform context-adaptive variable-length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding technique. In the case of context-based entropy coding, the context may be based on adjacent blocks. Following the entropy coding by the entropy coding unit 56, the encoded bitstream may be sent to another device (e.g., the video decoder 30) or archived for later transmission or retrieval.
[0081] The inverse quantization unit 58 and the inverse transform unit 60 each apply inverse quantization and inverse transform to reconstruct, for example, a residual block in the pixel domain for later use as a reference block. The motion compensation unit 44 may calculate a reference block by adding the residual block to a prediction block of one of the frames in the reference frame memory 64. Also, the motion compensation unit 44 may apply one or more interpolation filters to the reconstructed residual block to calculate sub-integer pixel values for use in motion estimation. The adder 62 adds the reconstructed residual block to the motion compensation prediction block generated by the motion compensation unit 44 to generate a reconstructed video block for storage in the reference frame memory 64. The reconstructed video block may be used by the motion estimation unit 42 and the motion compensation unit 44 as a reference block for inter-coding blocks in subsequent video frames.
[0082] Figure 3 is a block diagram showing an example of a video decoder 30 that can implement video coding techniques. In the example of Figure 3, the video decoder 30 includes an entropy decoding unit 70, a motion compensation unit 72, an intra prediction unit 74, an inverse quantization unit 76, an inverse transform unit 78, a reference frame memory 82, and an adder 80. The video decoder 30 may perform a decoding path that is generally inverse to the encoding path described with respect to the video encoder 20 (Figure 2) in some examples. The motion compensation unit 72 can generate prediction data based on the motion vectors received from the entropy decoding unit 70, while the intra prediction unit 74 may generate prediction data based on the intra prediction mode indicator received from the entropy decoding unit 70.
[0083] During decoding processing, video decoder 30 receives, from video encoder 20, an encoded video bitstream representing video blocks of an encoded video slice and associated syntax elements. Entropy decoding unit 70 of video decoder 30 decodes the bitstream and generates quantization coefficients, motion vectors or intra prediction mode indicators, and other syntax elements. Entropy decoding unit 70 transfers the motion vectors and other syntax elements to motion compensation unit 72. Video decoder 30 may receive syntax elements at the video slice level and / or the video block level.
[0084] When the video slice is coded as an intra-coded (I) slice, intra prediction unit 74 may generate prediction data for the video blocks of the current video slice based on the signaled intra prediction mode and data from previously decoded blocks of the current frame or picture. When the video frame is coded as an inter-coded (e.g., B, P, or GPB) slice, motion compensation unit 72 generates a predicted block for the video block of the current video slice based on the motion vectors and other syntax elements received from entropy decoding unit 70. The predicted block may be generated from one of the reference pictures within one of the reference picture lists. Video decoder 30 may configure reference frame lists, list 0 and list 1, using default configuration techniques based on the reference pictures stored in reference frame memory 82.
[0085] The motion compensation unit 72 determines prediction information for video blocks of the current video slice by analyzing motion vectors and other syntax elements, and generates a prediction block for the current video block to be decoded using the prediction information. For example, the motion compensation unit 72 uses some of the received syntax elements to determine a prediction mode (e.g., intra or inter prediction) used to code video blocks of the video slice, an inter prediction slice type (e.g., B slice, P slice, or GPB slice), configuration information for one or more of the reference picture lists for the slice, a motion vector for each inter-coded video block of the slice, an inter prediction status for each inter-coded video block of the slice, and other information for decoding video blocks in the current video slice.
[0086] The motion compensation unit 72 may also perform interpolation based on an interpolation filter. The motion compensation unit 72 may use the interpolation filter used by the video encoder 20 during the encoding of the video block to calculate interpolation values for sub-pixel of the reference block. In this case, the motion compensation unit 72 determines the interpolation filter used by the video encoder 20 from the received syntax elements and may use the interpolation filter to generate the prediction block.
[0087] Data of the texture image corresponding to the depth map may be stored in the reference frame memory 82. The motion compensation unit 72 may also be configured to inter-predict depth blocks of the depth map.
[0088] In one embodiment, video decoder 30 includes a user interface (UI) 84. User interface 84 is configured to receive input from a user (e.g., a network administrator) of video decoder 30. Via user interface 84, the user can manage or change settings on video decoder 30. For example, the user can input or otherwise provide the value of a parameter (e.g., a flag) to control the settings and / or operation of video decoder 30 according to the user's preferences. User interface 84 may be a graphical user interface (GUI) that enables the user to interact with video decoder 30 via, for example, graphical icons, drop-down menus, check boxes, etc. In some cases, user interface 84 may receive information from the user via a keyboard, a mouse, or other peripheral devices. In one embodiment, the user can access user interface 84 via a smartphone, a tablet device, a personal computer located remotely from video decoder 30, etc. As used herein, user interface 84 may sometimes be referred to as an external input or an external means.
[0089] Taking the above into consideration, video compression techniques perform spatial (intra-picture) prediction and / or temporal (inter-picture) prediction in order to reduce or remove redundancy inherent in a video sequence. For block-based video coding, a video slice (i.e., a video picture or a portion of a video picture) may be partitioned into video blocks, which may be referred to as coding tree blocks (CTBs), coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks within an intra-coded (I) slice of a picture are coded using spatial prediction with respect to reference samples in adjacent blocks within the same picture. Video blocks within an inter-coded (P or B) slice of a picture may use spatial prediction with respect to reference samples in adjacent blocks within the same picture or temporal prediction with respect to reference samples in other reference pictures. A picture may be referred to as a frame, and a reference picture may be referred to as a reference frame.
[0090] Spatial or temporal prediction results in a predicted block of the block to be coded. Residual data represents the pixel difference between the original block to be coded and the predicted block. An inter-coded block is coded according to a motion vector that points to a block of reference samples forming the predicted block, and the residual data indicates the difference between the coded block and the predicted block. An intra-coded block is coded according to an intra-coding mode and residual data. For further compression, the residual data may be transformed from the pixel domain to a transform domain, which results in residual transform coefficients, and the residual transform coefficients may then be quantized. The quantized transform coefficients are initially arranged in a two-dimensional array and scanned to generate a one-dimensional vector of transform coefficients, and entropy coding may be applied to achieve more compression.
[0091] Image and video compression have experienced rapid growth, leading to various coding standards. Such video coding standards include ITU-T H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) MPEG-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, Advanced Video Coding (AVC), also known as ITU-T H.264 or ISO / IEC MPEG-4 Part 10, and High Efficiency Video Coding (HEVC), also known as ITU-T H.265 or MPEG-H Part 2. AVC includes extensions such as Scalable Video Coding (SVC), Multi-View Video Coding (MVC) and Multi-View Video Coding Plus Depth (MVC+D), and 3D AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multi-View HEVC (MV-HEVC), 3D HEVC (3D-HEVC).
[0092] There is also a new video coding standard named Versatile Video Coding (VVC) developed by the Joint Video Experts Team (JVET) of ITU-T and ISO / IEC. The VVC standard has several working drafts, and in particular, one working draft of VVC, namely "Versatile Video Coding (Draft 5)" by B. Bross, J. Chen, and S. Liu, JVET-N1001-v3, 13th JVET Meeting, March 27, 2019 (VVC5 Draft5) is referred to herein.
[0093] The description of the technology disclosed herein is based on the Versatile Video Coding of the video coding standard under development by the Joint Video Experts Team of ITU-T and ISO / IEC. However, this technology is also applicable to other video codec specifications.
[0094] Figure 4 is a representation 400 of the relationship between an Intra Random Access (IRAP) picture 402 with respect to a leading picture 404 and subsequent pictures 406 in decoding order 408 and presentation order 410. In one embodiment, the IRAP picture 402 is called an Instantaneous Decoder Refresh (IDR) picture having a Clean Random Access (CRA) picture or a Random Access Decodable (RADL) picture. In HEVC, IDR pictures, CRA pictures, and Broken Link Access (BLA) pictures are all considered IRAP pictures 402. In VVC, at the 12th JVET meeting in October 2018, it was agreed to have both IDR pictures and CRA pictures as IRAP pictures. In one embodiment, Broken Link Access (BLA) and Gradual Decoder Refresh (GDR) pictures may also be considered IRAP pictures. The decoding process of an encoded video sequence always starts with an IRAP.
[0095] A CRA picture is an IRAP picture in which each Video Coding Layer (VCL) Network Abstraction Layer (NAL) unit has a nal_unit_type equal to CRA_NUT. A CRA picture does not reference pictures other than itself for inter prediction in its decoding process and may be the first picture in the bitstream in decoding order or may appear later in the bitstream. A CRA picture may have a related RADL or Random Access Skip Leading (RASL) picture. When the CRA picture has a NoOutputBeforeRecoveryFlag equal to 1, the related RASL picture may not be decodable because it may contain references to pictures that do not exist in the bitstream and thus is not output by the decoder.
[0096] As shown in FIG. 4, the leading pictures 404 (e.g., pictures 2 and 3) are after the IRAP picture 402 in the decoding order 408, but precede the IRAP picture 402 in the presentation order 410. The subsequent pictures 406 are after the IRAP picture 402 in both the decoding order 408 and the presentation order 410. Although two leading pictures 404 and one subsequent picture 406 are shown in FIG. 4, those skilled in the art will understand that a greater or lesser number of leading pictures 404 and / or subsequent pictures 406 may be in the decoding order 408 and the presentation order 410 in an actual application.
[0097] The leading pictures 404 in FIG. 4 are divided into two types: Random Access Skip Leading (RASL) and RADL. When decoding starts with the IRAP picture 402 (e.g., picture 1), the RADL picture (e.g., picture 3) can be decoded properly, but the RASL picture (e.g., picture 2) cannot be decoded properly. Therefore, the RASL picture is discarded. In light of the distinction between the RADL picture and the RASL picture, the type of the leading picture 404 related to the IRAP picture 402 should be identified as either RADL or RASL for efficient and proper coding. In HEVC, when there are RASL and RADL pictures, for the RASL and RADL pictures related to the same IRAP picture 402, it is restricted that the RASL picture precedes the RADL picture in the presentation order 410.
[0098] The IRAP picture 402 provides the following two important functions / advantages. First, the presence of the IRAP picture 402 indicates that the decoding process can start from that picture. This function enables the decoding process to have a random access feature that starts from its position in the bitstream, not necessarily at the beginning of the bitstream, as long as the IRAP picture 402 is present at that position. Second, the presence of the IRAP picture 402 refreshes the decoding process so that the coded pictures starting with the IRAP picture 402, excluding the RASL picture, are coded without referring to the previous picture. As a result, when the IRAP picture 402 exists in the bitstream, any errors that may occur during the decoding of the pictures coded before the IRAP picture 402 will be stopped from propagating to the IRAP picture 402 and those pictures following the IRAP picture 402 in the decoding order 408.
[0099] Although the IRAP picture 402 provides important functions, there is a penalty for compression efficiency. The presence of the IRAP picture 402 causes a bitrate surge. This penalty for compression efficiency is due to two reasons. First, since the IRAP picture 402 is an intra-predicted picture, the picture itself requires relatively more bits to represent compared to other pictures that are inter-predicted pictures (e.g., the leading picture 404, the subsequent picture 406). Second, the presence of the IRAP picture 402 disrupts temporal prediction (this is because the decoder refreshes the decoding process, and one of the actions of the decoding process is to remove the previous reference pictures in the decoded picture buffer (DPB)), and the IRAP picture 402 makes the coding of the pictures following the IRAP picture 402 in the decoding order 408 less efficient (i.e., requires more bits to represent). This is because there are fewer reference pictures for their inter-prediction coding.
[0100] Among the picture types regarded as IRAP picture 402, IDR pictures in HEVC have different signaling and derivation when compared with other picture types. Some of the differences are as follows.
[0101] In the signaling and derivation of the picture order count (POC) value of an IDR picture, the most significant bit (MSB) part of the POC is not derived from the previous important picture and is simply set equal to 0.
[0102] In the signaling information required for reference picture management, the slice header of an IDR picture does not contain the information that needs to be signaled to assist in reference picture management. For other picture types (i.e., CRA, trailing, temporal sublayer access, etc.), information such as the reference picture set or other forms of similar information (e.g., reference picture list) described below is required for the reference picture marking process (i.e., the process of determining the state of the reference pictures in the decoded picture buffer (DPB) regardless of whether they are used for reference or not). However, for IDR pictures, since the presence of an IDR indicates that the decoding process simply marks all the reference pictures in the DPB as not used for reference, such information does not need to be signaled.
[0103] In HEVC and VVC, the IRAP picture 402 and the leading picture 404 may each be included within a single Network Abstraction Layer (NAL) unit. A set of NAL units may be referred to as an access unit. The IRAP picture 402 and the leading picture 404 are given different NAL unit types so that they can be easily identified by system-level applications. For example, a video splicer needs to understand the coded picture type without the need to understand the excessive details of the syntax elements in the coded bitstream. In particular, it needs to identify the IRAP picture 402 from non-IRAP pictures and identify the leading picture 404 from subsequent pictures 406, including the determination of RASL and RADL pictures. The subsequent picture 406 is a picture related to the IRAP picture 402 and following the IRAP picture 402 in the presentation order 410. A picture can be after a particular IRAP picture 402 in the decoding order 408 and can precede other IRAP pictures 402 in the decoding order 408. For this purpose, giving the IRAP picture 402 and the leading picture 404 their own NAL unit types helps such applications.
[0104] In HEVC, the NAL unit types for IRAP pictures include the following: BLA with leading picture (BLA_W_LP): The NAL unit of a broken link access (BLA) picture, where one or more leading pictures may follow in the decoding order. BLA with RADL (BLA_W_RADL): The NAL unit of a BLA picture, where one or more RADL pictures may follow in the decoding order, but no RASL pictures follow. BLA without leading picture (BLA_N_LP): The NAL unit of a BLA picture, where no leading picture follows in the decoding order. IDR with RADL (IDR_W_RADL): The NAL unit of an IDR picture, where one or more RADL pictures may follow in the decoding order, but no RASL pictures follow. IDR without leading picture (IDR_N_LP): A NAL unit of an IDR picture where no leading picture follows in decoding order. CRA: A NAL unit of a Clean Random Access (CRA) picture where a leading picture (RASL picture, RADL picture, or both) may follow. RADL: A NAL unit of a RADL picture. RASL: A NAL unit of a RASL picture.
[0105] In VVC, the NAL unit types of IRAP pictures 402 and leading pictures 404 are as follows. IDR with RADL (IDR_W_RADL): A NAL unit of an IDR picture where one or more RADL pictures may follow in decoding order, but no RASL picture follows. IDR without leading picture (IDR_N_LP): A NAL unit of an IDR picture where no leading picture follows in decoding order. CRA: A NAL unit of a Clean Random Access (CRA) picture where a leading picture (RASL picture, RADL picture, or both) may follow. RADL: A NAL unit of a RADL picture. RASL: A NAL unit of a RASL picture.
[0106] FIG. 5 shows a video bitstream 550 configured to implement a Gradual Decoding Refresh (GDR) technique 500. As used herein, video bitstream 550 may also be referred to as a coded video bitstream, a bitstream, or a variation thereof. As shown in FIG. 5, bitstream 550 includes a Sequence Parameter Set (SPS) 552, a Picture Parameter Set (PPS) 554, a slice header 556, and picture data 558.
[0107] The SPS 552 contains data that is common to all pictures in a series of pictures (SOP). In contrast, the PPS 554 contains data that is common to the entire picture. The slice header 556 contains information about the current slice, such as, for example, the slice type, which reference picture is used, etc. The SPS 552 and PPS 554 are sometimes generally referred to as parameter sets. The SPS 552, PPS 554, and slice header 556 are types of network abstraction layer (NAL) units. An NAL unit is a syntax structure that contains an indication of the type of data to follow (e.g., coded video data). NAL units are classified into video coding layer (VCL) and non-VCL NAL units. A VCL NAL unit contains data representing the values of samples within a video picture, and a non-VCL NAL unit contains parameter sets (important header data applicable to a number of VCL NAL units) and any associated additional information such as supplementary enhancement information (timing information and other supplementary data that is not necessary to decode the values of samples within a video picture but may enhance the usefulness of the decoded video signal). Those skilled in the art will understand that the bitstream 550 may contain other parameters and information in an actual application.
[0108] The image data 558 in FIG. 5 includes data related to an image or video to be encoded or decoded. The image data 558 may simply be referred to as the payload or data carried in the bitstream 550. In one embodiment, the image data 558 includes a CVS 508 (or CLVS) that includes a GDR picture 502, one or more subsequent pictures 504, and a recovery point picture 506. In one embodiment, the GDR picture 502 is referred to as the CVS start (CVSS) picture. The CVS 508 is an encoded video sequence for each coded layer video sequence (CLVS) in the video bitstream 550. Notably, when the video bitstream 550 includes a single layer, the CVS and the CLVS are the same. The CVS and the CLVS are different only when the video bitstream 550 includes multiple layers. In one embodiment, since the subsequent picture 504 precedes the recovery point picture 506 during the GDR period, it may be regarded as being in the form of a GDR picture.
[0109] In one embodiment, the GDR picture 502, the subsequent picture 504, and the recovery point picture 506 may define a GDR period within the CVS 508. In one embodiment, the decoding order starts with the GDR picture 502, continues with the subsequent picture 504, and then proceeds to the recovery picture 506.
[0110] The CVS 508 is a series of pictures (or a partial portion thereof) starting with the GDR picture 502 and including all pictures (or a portion thereof) up to but not including the next GDR picture, or up to the end of the bitstream. The GDR period is a series of pictures starting with the GDR picture 502 and including all pictures up to and including the recovery point picture 506. The decoding process of the CVS 508 always starts with the GDR picture 502.
[0111] As shown in FIG. 5, the GDR technique 500 or principle operates on a series of pictures that start with a GDR picture 502 and end with a recovery point picture 506. The GDR picture 502 includes a refresh / clean area 510 that contains all coded blocks using intra prediction (i.e., intra-predicted blocks), and an unrefreshed / dirty area 512 that contains all coded blocks using inter prediction (i.e., inter-predicted blocks).
[0112] A subsequent picture 504 that is immediately adjacent to the GDR picture 502 includes a refresh / clean area 510 that has a first portion 510A coded using intra prediction and a second portion 510B coded using inter prediction. The second portion 510B is coded, for example, by referring to the refresh / clean area 510 of a leading picture within the GDR period of the CVS 508. As shown, the refresh / clean area 510 of the subsequent picture 504 expands as the coding process moves or progresses in a consistent direction (e.g., from left to right), and accordingly, the unrefreshed / dirty area 512 is shrunk. Finally, a recovery point picture 506 that includes only the refresh / clean area 510 is obtained from the coding process. In particular, as further discussed below, the second portion 510B of the refresh / clean area 510 coded as an inter prediction block may refer only to the refresh / clean area 510 of the reference picture.
[0113] As shown in FIG. 5, the GDR picture 502, subsequent picture 504, and recovery point picture 506 in the CVS 508 are each contained within their own VCL NAL unit 530. A set of NAL units is sometimes referred to as an access unit.
[0114] In one embodiment, a VCL NAL unit 530 including a GDR picture 502 in CVS 508 has a GDR NAL unit type (GDR_NUT). That is, in the embodiment, the VCL NAL unit 530 including the GDR picture 502 in CVS 508 has its own unique NAL unit type with respect to subsequent pictures 504 and recovery point pictures 506. In one embodiment, the GDR_NUT allows the bitstream 550 to start with a GDR picture 502 instead of having to start with an IRAP picture. Designating the VCL NAL unit 530 of the GDR picture 502 as the GDR_NUT may, for example, indicate to the decoder that the initial VCL NAL unit 530 in CVS 508 includes the GDR picture 502. In one embodiment, the GDR picture 502 is the initial picture in CVS 508. In one embodiment, the GDR picture 502 is the initial picture in the GDR period.
[0115] FIG. 6 is a schematic diagram showing an undesirable motion search 600 when using encoder constraints to support GDR. As shown, the motion search 600 shows a current picture 602 and a reference picture 604. The current picture 602 and the reference picture 604 each include a refresh region 606 coded by intra prediction, a refresh region 608 coded by inter prediction, and an unrefreshed region 608. The refresh region 604, the refresh region 606, and the unrefreshed region 608 are similar to the first portion 510A of the refresh / clean region 510, the second portion 510B of the refresh / clean region 510, and the unrefreshed / dirty region 512 of FIG. 5.
[0116] During motion search, the encoder is constrained or prevented from selecting any motion vector 610 that results in a portion of the samples of the reference block 612 located outside the refresh region 606. This occurs even when the reference block 612 provides the best rate-distortion cost criterion when predicting the current block 614 within the current picture 602. Thus, FIG. 6 shows the reasons why it is not optimal in motion search 600 when using encoder restrictions to support GDR.
[0117] FIG. 7 shows a video bitstream 750 configured to implement the hierarchical decoding refresh (GDR) technique 700. As used herein, the video bitstream 750 may also be referred to as a coded video bitstream, a bitstream, or a variation thereof. As shown in FIG. 7, the bitstream 750 includes a sequence parameter set (SPS) 752, a picture parameter set (PPS) 754, a slice header 756, and picture data 758. The bitstream 750, SPS 752, PPS 754, and slice header 756 in FIG. 7 are similar to the bitstream 550, SPS 552, PPS 554, and slice header 556 in FIG. 5. Therefore, for the sake of brevity, the description of these elements will not be repeated.
[0118] The picture data 758 in FIG. 7 includes data related to the picture or video to be encoded or decoded. The picture data 758 may simply be referred to as the payload or data carried in the bitstream 750. In one embodiment, the picture data 758 includes a CVS 708 (or CLVS) that includes a GDR picture 702, one or more subsequent pictures 704, and an end picture 706 of the sequence picture. In one embodiment, the GDR picture 702 is called a CVSS picture. The decoding process of the CVS 508 always starts with the GDR picture 702.
[0119] As shown in FIG. 7, the GDR picture 702, subsequent picture 704, and sequence end picture 706 in CVS708 are each included within their own VCL NAL unit 730. The set of NAL units 730 in CVS708 may be referred to as an access unit.
[0120] In the latest draft specification of VVC, the output of the picture before the IRAP picture is defined as follows. The previous picture for the IRAP picture (e.g., the previously decoded picture) is 1) decoded earlier than the IRAP picture, 2) an output is indicated, 3) present in the decoded picture buffer (DPB) at the start of decoding of the IRAP picture, and 4) refers to a picture that has not been output at the start of decoding of the IRAP picture. As used herein, a conventional picture may sometimes be referred to as a previously decoded picture.
[0121] The slice header syntax includes the syntax element no_output_of_prior_pics_flag for IDR and CRA pictures. The meaning is as follows.
[0122] The no_output_of_prior_pics_flag affects the output of the previously decoded pictures in the decoded picture buffer after decoding of an IDR picture that is not the first picture of the bitstream as defined in Annex C of VVC draft 5.
[0123] Section C.3.2 of VVC draft 5 (Deletion of pictures from the DPB before decoding the current picture) includes the following text.
[0124] When the current picture is an IRAP picture with NoIncorrectPicOutputFlag equal to 1 and not picture 0, the following ordered procedure applies.
[0125] 1. The variable NoOutputOfPriorPicsFlag is derived as follows for the decoder under test.
[0126] If the current picture is a CRA picture, NoOutputOfPriorPicsFlag is set to 1 (regardless of the value of no_output_of_prior_pics_flag).
[0127] Otherwise, if pic_width_in_luma_samples, pic_height_in_luma_samples, croma_format_idc, separate_colour_plane_flag, bit_depth_luma_minus8, bit_depth_chroma_minus8, or sps_max_dec_pic_buffering_minus1[HighestTid] derived from the active SPS is different from the value of pic_width_in_luma_samples, pic_height_in_luma_samples, chroma_format_idc, separate_colour_plane_flag, bit_depth_luma_minus8, bit_depth_chroma_minus8, or sps_max_dec_pic_buffering_minus1[HighestTid] derived from the active SPS for the previous picture, then the decoder under test may (but should not) set NoOutputOfPriorPicsFlag to 1, regardless of the value of no_output_of_pics_flag.
[0128] Note - Under these conditions, it is preferred to set NoOutputOfPriorPicsFlag equal to no_output_of_prior_pics_flag, but in this case, the decoder under test is allowed to set NoOutputOfPriorPicsFlag to 1.
[0129] Otherwise, NoOutputOfPriorPicsFlag is set equal to no_output_of_prior_pics_flag.
[0130] The value of NoOutputOfPriorPicsFlag derived for the decoder under test is applied to the Hypothetical Reference Decoder (HRD). When the value of NoOutputOfPriorPicsFlag is equal to 1, all picture memory buffers within the DPB are emptied without output of the pictures they contain, and the fullness of the DPB is set to 0.
[0131] Section C.5.2.2 (Output and deletion of pictures from the DPB) of VVC Draft 5 contains the following text.
[0132] When the current picture is an IRAP picture with NoIncorrectPicOutputFlag equal to 1 and not Picture 0, the following ordered procedure applies.
[0133] 1. The variable NoOutputOfPriorPicsFlag is derived for the decoder under test as follows.
[0134] If the current picture is a CRA picture, NoOutputOfPriorPicsFlag is set to 1 (regardless of the value of no_output_of_prior_pics_flag).
[0135] Otherwise, if pic_width_in_luma_samples, pic_height_in_luma_samples, croma_format_idc, separate_colour_plane_flag, bit_depth_luma_minus8, bit_depth_chroma_minus8, or sps_max_dec_pic_buffering_minus1[HighestTid] derived from the active SPS is different from the value of pic_width_in_luma_samples, pic_height_in_luma_samples, chroma_format_idc, separate_colour_plane_flag, bit_depth_luma_minus8, bit_depth_chroma_minus8, or sps_max_dec_pic_buffering_minus1[HighestTid] derived from the SPS active for the previous picture, respectively, the decoder under test may (but should not) set NoOutputOfPriorPicsFlag to 1, regardless of the value of no_output_of_pics_flag.
[0136] Note - Under these conditions, it is preferred to set NoOutputOfPriorPicsFlag equal to no_output_of_prior_pics_flag, but in this case, the decoder under test is allowed to set NoOutputOfPriorPicsFlag to 1.
[0137] Otherwise, NoOutputOfPriorPicsFlag is set equal to no_output_of_prior_pics_flag.
[0138] 2. The value of NoOutputOfPriorPicsFlag derived for the decoder under test is applied to the HRD as follows:
[0139] When NoOutputOfPriorPicsFlag is equal to 1, all picture memory buffers in the DPB are emptied without output of the pictures they contain, and the DPB fullness is set to 0.
[0140] Otherwise (NoOutputOfPriorPicsFlag is equal to 0), all picture memory buffers including pictures marked as "not needed for output" and "not used for reference" are emptied (without output), and all non-empty picture memory buffers in the DPB are emptied by repeatedly calling the "bumping" process specified in Section C.5.2.4, and the DPB fullness is set equal to 0.
[0141] The problems of existing designs have been discussed.
[0142] In the latest draft specification of VVC, for a CRA picture where NoIncorrectPicOutputFlag is equal to 1 (i.e., a CRA picture that starts a new CVS), the value of NoOutputOfPriorPicsFlag is set equal to 1 regardless of the value of no_output_of_prior_pics_flag, so the value of no_output_of_prior_pics_flag is not used. That is, the pictures before each CRA picture that starts a CVS are not output. However, similar to the case of IDR pictures, the output / display of the previous pictures can provide more continuous playback, and thus, a better user experience can be provided as long as the DPB does not overflow when decoding the picture that starts a new CVS and the subsequent pictures in decoding order.
[0143] To solve the problems discussed above, the present disclosure provides the following inventive aspects. The value of no_output_of_prior_pics_flag is used in the specification of the output of the pictures before each CRA picture that starts a new CVS and is not the first picture in the bitstream. This enables more continuous playback and thus a better user experience.
[0144] Also, this disclosure applies to other types of pictures that start a new CVS, such as the hierarchical random access (GRA) pictures currently specified in the latest VVC draft specification. In one embodiment, the GRA picture may refer to a GDR picture or be synonymous with a GDR picture.
[0145] As an example, when decoding a video bitstream, a flag corresponding to a clean random access (CRA) picture is signaled in the bitstream. The flag specifies whether a decoded picture in the decoded picture buffer that is decoded earlier than the CRA picture is output when the CRA picture starts a newly coded video sequence. That is, when the value of the flag indicates that the previous picture is to be output (e.g., when the value is equal to 0), the previous picture is output. In one embodiment, the flag is specified as the no_output_of_prior_pics_flag.
[0146] As another example, when decoding a video bitstream, a flag corresponding to a hierarchical random access (GRA) picture is signaled in the bitstream. The flag specifies whether a decoded picture in the decoded picture buffer that is decoded earlier than the GRA picture is output when the GRA picture starts a newly coded video sequence. That is, when the value of the flag indicates that the previous picture is to be output (e.g., when the value is equal to 0), the previous picture is output. In one embodiment, the flag is specified as the no_output_of_prior_pics_flag.
[0147] Disclosed herein is a technique for output of previous pictures (e.g., previously decoded pictures) in a decoded picture buffer (DPB) when random access point pictures other than instant decoder refresh (IDR) pictures (e.g., clean random access (CRA) pictures, gradual random access (GRA) pictures, or gradual decode refresh (GDR) pictures, CVSS pictures, etc.) are encountered in decoding order. When reaching a random access point picture, emptying the previously decoded pictures from the DPB prevents the DPB from overflowing and promotes more continuous playback. Accordingly, a coder / decoder (also known as a “codec”) in video coding is improved compared to current codecs. In practical terms, an improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed.
[0148] FIG. 8 is one embodiment of a method 800 for decoding a coded video bitstream implemented by a video decoder (e.g., video decoder 30). Method 800 may be performed after the decoded bitstream is received directly or indirectly from a video encoder (e.g., video encoder 20). Method 800 improves the decoding process by emptying the DPB before the current picture is decoded when a random access point picture is encountered. Method 800 prevents overflow of the DPB and promotes more continuous playback. Accordingly, in practical terms, the performance of the codec is improved, which leads to a better user experience.
[0149] In block 802, the video decoder receives a coded video bitstream (e.g., bitstream 750). The coded video bitstream includes a hierarchical decoding refresh (GDR) picture and a first flag having a first value. In one embodiment, the GDR picture is not the first picture of the video bitstream. In one embodiment, the first flag is designated as no_output_of_prior_pics_flag. In one embodiment, the GDR picture is disposed in a video coding layer (VCL) network abstraction layer (NAL) unit having a hierarchical decoding refresh (GDR) network abstraction layer (NAL) unit type (GDR_NUT).
[0150] In block 804, the video decoder sets a second value of a second flag equal to the first value of the first flag. In one embodiment, the second flag is designated as NoOutputOfPriorPicsFlag. In one embodiment, the second flag is internal to the decoder.
[0151] In block 806, the video decoder empties the previously decoded picture corresponding to the GDR picture from the DPB based on the second flag having the second value. In one embodiment, All the previously decoded picture is emptied from the DPB. That is, the video decoder removes the previously decoded picture from the picture storage buffer in the DPB. In one embodiment, the previously decoded picture is not output or displayed when the previously decoded picture is removed from the DPB. In one embodiment, the DPB fullness parameter is set to zero when the first flag is set to the first value. The "DPB fullness" parameter indicates how many pictures are held in the DPB. Setting the DPB fullness parameter to zero represents that the DPB is empty. GDR pictures
[0152] At block 808, the video decoder decodes the current picture after the DPB becomes empty. In one embodiment, the current picture is from the same CVS as the CRA picture and is obtained or encountered after the CRA in the decoding order. In one embodiment, an image generated based on the current picture is displayed to a user of an electronic device (such as a smartphone, tablet, laptop, personal computer, etc.).
[0153] FIG. 9 is an embodiment of a method 900 for encoding a video bitstream implemented by a video encoder (such as video encoder 20). Method 900 may be executed when a picture (such as from a video) is encoded into a video bitstream and sent towards a video decoder (such as video decoder 30). Method 900 improves the encoding process by instructing the video decoder to empty the DPB before the current picture is decoded when a random access point picture is encountered. Method 900 prevents DPB overflow and promotes more continuous playback. Thus, as a practical matter, the codec performance is improved, which leads to a better user experience.
[0154] At block 902, the video encoder determines a random access point of the video sequence. At block 904, the video encoder encodes a hierarchical decoding refresh (GDR) picture for the video sequence at the random access point. In one embodiment, the GDR picture is not the first picture of the video bitstream. In one embodiment, the GDR picture is arranged in a video coding layer (VCL) network abstraction layer (NAL) unit having a hierarchical decoding refresh (GDR) network abstraction layer (NAL) unit type (GDR_NUT).
[0155] At block 906, the video encoder removes a previously decoded picture from the decoded picture buffer (DPB). AllSet a flag to a first value that instructs the video decoder to empty the picture. In one embodiment, the video decoder is instructed from the DPB to All that were decoded before GDR pictures empty the picture. In one embodiment, the flag is specified as no_output_of_prior_pics_flag. In one embodiment, the video encoder instructs the video decoder to set the DPB fullness parameter to zero when the flag is set to the first value. In one embodiment, the first value of the flag is 1.
[0156] At block 908, the video encoder generates a video sequence having a GDR picture at a random access point and a video bitstream including the flag. At block 910, the video encoder stores the video bitstream for transmission to the video decoder.
[0157] The following syntax and semantics may be used to implement the embodiments disclosed herein. The following description is with respect to the base text of the latest VVC draft specification. In other words, only the deltas are described, but the text of the base text not mentioned below is applied as is. The text added to the base text is shown in bold, and the deleted text is shown in italics (hereinafter, the bold part may be represented by the part sandwiched between (bold start) and (bold end), and the italic part may be represented by the part sandwiched between (italic start) and (italic end)).
[0158] General slice header syntax (7.3.5.1 of VVC).
Table 1
[0159] Meaning of the general slice header (7.4.6.1 of VVC).
[0160] If present, the values of the slice header syntax elements slice_pic_parameter_set_id, slice_pic_order_cnt_lsb, no_output_of_prior_pics_flag, and slice_temporal_mvp_enabled_flag shall be the same for all slice headers of a coded picture.
[0161] ···
[0162] no_output_of_prior_pics_flag affects the output of previously decoded pictures in the decoded picture buffer after decoding of an (italic start)IDR picture(italic end)(bold start)CVSS picture(bold end) that is not the first picture in the bitstream as specified in Annex C.
[0163] ···
[0164] Deletion of pictures from the DPB before decoding of the current picture (VVC C.3.2).
[0165] ···
[0166] When the current picture is a (italic start)CVSS picture other than picture 0 that is an IRAP picture with NoIncorrectPicOutputFlag equal to 1(italic end), the following ordered procedure applies.
[0167] 1. The variable NoOutputOfPriorPicsFlag is derived as follows for the decoder under test.
[0168] (italic start)If the current picture is a CRA picture, NoOutputOfPriorPicsFlag is set to 1 (regardless of the value of no_output_of_prior_pics_flag).(italic end)
[0169] Otherwise, if pic_width_in_luma_samples, pic_height_in_luma_samples, croma_format_idc, separate_colour_plane_flag, bit_depth_luma_minus8, bit_depth_chroma_minus8, or sps_max_dec_pic_buffering_minus1[HighestTid] derived from the active SPS is different from the values of pic_width_in_luma_samples, pic_height_in_luma_samples, chroma_format_idc, separate_colour_plane_flag, bit_depth_luma_minus8, bit_depth_chroma_minus8, or sps_max_dec_pic_buffering_minus1[HighestTid] derived from the active SPS for the previous picture respectively, the NoOutputOfPriorPicsFlag may be set to 1 (but should not be) by the decoder under test, regardless of the value of no_output_of_pics_flag.
[0170] Note - Under these conditions, it is preferred to set NoOutputOfPriorPicsFlag equal to no_output_of_prior_pics_flag, but in this case, the decoder under test is allowed to set NoOutputOfPriorPicsFlag to 1.
[0171] Otherwise, NoOutputOfPriorPicsFlag is set equal to no_output_of_prior_pics_flag.
[0172] 2. The value of NoOutputOfPriorPicsFlag derived for the decoder during testing is applied to the HRD, and when the value of NoOutputOfPriorPicsFlag is equal to 1, all picture memory buffers in the DPB are emptied without output of the pictures they contain, and the fullness of the DPB is set to 0.
[0173] FIG. 10 is a schematic diagram of a video coding device 1000 (e.g., video encoder 20 or video decoder 30) according to an embodiment of the present disclosure. The video coding device 1000 is suitable for implementing the disclosed embodiments described herein. The video coding device 1000 includes an input port 1010 and a receiver unit (Rx) 1020 for receiving data, a processor, logic unit, or central processing unit (CPU) 1030 for processing data, a transmitter unit (Tx) 1040 and an output port 1050 for transmitting data, and a memory 1060 for storing data. The video coding device 1000 may also include optical to electrical (OE) components and electrical to optical (EO) components coupled to the input port 1010, the receiver unit 1020, the transmitter unit 1040, and the output port 1050 for the ingress and egress of optical or electrical signals.
[0174] Processor 1030 is implemented by hardware and software. Processor 1030 may be implemented as one or more CPU chips, cores (e.g., multi-core processors), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and digital signal processors (DSPs). Processor 1030 communicates with an input port 1010, a receiver unit 1020, a transmitter unit 1040, an output port 1050, and a memory 1060. Processor 1030 includes a coding module 1070. Coding module 1070 implements the disclosed embodiments described above. For example, coding module 1070 implements, processes, prepares, or provides various codec functions. Thus, including the encoding module 1070 provides a substantial improvement to the functionality of the video coding device 1000 and results in a transformation of the video coding device 1000 to different states. Alternatively, coding module 1070 is implemented as instructions stored in memory 1060 and executed by processor 1030.
[0175] Video coding device 1000 may also include an input and / or output (I / O) device 1080 for communicating data with a user. I / O device 1080 may include output devices such as a display for displaying video data, a speaker for outputting audio data, etc. I / O device 1080 may also include input devices such as a keyboard, a mouse, a trackball, etc., and / or corresponding interfaces for interacting with such output devices.
[0176] Memory 1060 may include one or more disks, tape drives, and solid state drives, and is used as an overflow data storage device, stores programs when such programs are selected for execution, and stores instructions and data read during program execution. Memory 1060 may be volatile and / or non-volatile, and may be read-only memory (ROM), random access memory (RAM), ternary content addressable memory (TCAM), and / or static random access memory (RSAM).
[0177] FIG. 11 is a schematic diagram of an embodiment of coding means 1100. In one embodiment, coding means 1100 is implemented in a video coding device 1102 (e.g., video encoder 20 or video decoder 30). Video coding device 1102 includes receiving means 1101. Receiving means 1101 is configured to receive a picture to be coded or a bitstream to be decoded. Video coding device 1102 includes transmitting means 1107 coupled to receiving means 1101. Transmitting means 1107 is configured to transmit a bitstream to a decoder or a decoded image to a display means (e.g., one of I / O devices 1080).
[0178] Video coding device 1102 includes storage means 1103. Storage means 1103 is coupled to at least one of receiving means 1101 or transmitting means 1107. Storage means 1103 is configured to store instructions. Video coding device 1102 also includes processing means 1105. Processing means 1105 is coupled to storage means 1103. Processing means 1105 is configured to execute instructions stored in storage means 1103 to perform the methods disclosed herein.
[0179] Also, it should be understood that the steps of the exemplary methods described herein need not necessarily be executed in the order described, and that the order of such method steps is merely exemplary. Similarly, additional steps may be included in such methods, or certain steps may be omitted or combined in ways consistent with various embodiments of the present disclosure.
[0180] Although multiple embodiments are provided in the present disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. This example is illustrative and not restrictive, and its intention is not limited to the details given herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.
[0181] Additionally, the techniques, systems, subsystems, and methods described and illustrated individually or separately in various embodiments may be combined or integrated with other systems, modules, techniques, or methods without departing from the scope of the present disclosure. Other items shown or discussed as being coupled, directly coupled, or communicating with each other may be indirectly coupled or communicated through some interface, device, or intermediate component, electrically, mechanically, or otherwise. Other examples of changes, substitutions, and modifications are ascertainable by those skilled in the art and may be made without departing from the spirit and scope disclosed herein.
Claims
1. A decoding method implemented by a video decoder, comprising: receiving, by the video decoder, a coded video bitstream, the coded video bitstream including a coded hierarchical refresh (GDR) picture and a first flag having a first value, the coded GDR picture not being the first picture of the coded video bitstream, and the first flag being present in a slice header corresponding to the coded GDR picture in the video bitstream; setting, by the video decoder, a second value of a second flag equal to the first value of the first flag, the first flag being designated as no_output_of_prior_pics_flag and the second flag being designated as NoOutputOfPriorPicsFlag; emptying, by the video decoder, all pictures decoded prior to the GDR picture from a decoded picture buffer (DPB) based on the second flag having the second value; decoding, by the video decoder, a current picture after the DPB becomes empty.
2. The method according to claim 1, wherein the coded GDR picture is arranged in a video coding layer (VCL) NAL unit having a GDR network abstraction layer (NAL) unit type (GDR_NUT).
3. The method according to claim 1 or 2, further comprising setting a DPB fullness parameter to zero when the first flag is set to the first value.
4. The method according to any one of claims 1 to 3, wherein the DPB is emptied after the coded GDR picture is decoded.
5. The method according to any one of claims 1 to 4, wherein the first value of the first flag is 1.
6. A decoding device, comprising: a receiver configured to receive a coded video bitstream; a memory coupled to the receiver, the memory storing instructions. a processor coupled to the memory, wherein the processor causes the decoding device to receive the coded video bitstream, the coded video bitstream including a coded hierarchical refresh (GDR) picture and a first flag having a first value, the coded GDR picture not being the first picture of the coded video bitstream and the first flag being present in a slice header corresponding to the coded GDR picture in the video bitstream; set a second value of a second flag equal to the first value of the first flag, the first flag being designated as no_output_of_prior_pics_flag and the second flag being designated as NoOutputOfPriorPicsFlag; empty all pictures decoded prior to the GDR picture from a decoded picture buffer (DPB) based on the second flag having the second value; decode a current picture after the DPB is emptied, the decoding device being configured to execute the instruction that causes the above. **Claim 7** The decoding device according to claim 6, further comprising a display configured to display a picture generated based on the current picture. **Claim 8** An encoding device, comprising: a receiver configured to receive a picture to be encoded or a bitstream to be decoded; a transmitter coupled to the receiver, the transmitter being configured to transmit the bitstream to a decoder or transmit a decoded image to a display; a memory coupled to at least one of the receiver or the transmitter, the memory being configured to store instructions; a processor coupled to the memory, the processor being configured to execute the instructions stored in the memory for performing the method according to any one of claims 1 to 5, the encoding device comprising the processor. **Claim 9** The coding device according to claim 8, further including a display configured to display an image.
10. Coding means, comprising: Receiving means configured to receive a picture to be coded or a bitstream to be decoded; Transmitting means coupled to the receiving means, the transmitting means being configured to transmit the bitstream to decoding means or the decoded image to display means; Storage means coupled to at least one of the receiving means or the transmitting means, the storage means being configured to store instructions; Processing means coupled to the storage means, the processing means being configured to execute the instructions stored in the storage means for performing the method according to any one of claims 1 to 5.
11. A computer-readable storage medium storing a computer program executable by a processor, wherein when the computer program is executed by the processor, the processor performs the method according to any one of claims 1 to 5.
12. A program including program code for performing the method according to any one of claims 1 to 5 when executed on a computer or a processor.
13. An encoder including a processing circuit for performing the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Updating parameter sets in video coding
JP2015515239A
Signaling of dpb parameters and dpb behavior in the vps extension
JP2016518041A
Improved Inference of NoOutputOfPriorPicsFlag in Video Coding
JP2017507542A