Encoder, decoder, and corresponding method
The GDR method enables random access in video coding without IRAPs by controlling flag settings, enhancing efficiency and user experience in low-latency applications.
Patent Information
- Application Number
- JP2023101903
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-07-05
- Filing Date
- 2023-06-21
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2040-03-11
AI Technical Summary
Existing video coding methods require intra random access points (IRAPs) for random access, which increase data size and reduce compression efficiency, especially in low-latency applications like multiparty video conferencing.
A method for gradual decoding refresh (GDR) that enables random access without IRAPs by setting flags to prevent output of potentially dirty data, allowing intra refresh and sequential decoding.
Improves video coding efficiency by reducing data size and latency, providing a better user experience in low-latency applications by avoiding the need for IRAPs and optimizing data output.
Smart Images

Figure 0007704810000005 
Figure 0007704810000006 
Figure 0007704810000007
Abstract
Description
[Technical field]
[0002] In general, this disclosure provides a method for supporting gradual decoding refresh in video coding. More specifically, this disclosure describes a technique for implementing intra-random access points (IRAPs). Sequential intra refresh is possible without the need for intra random access point (IRAP) pictures. This allows random access to be enabled. [Background technology]
[0003] The amount of video data required to render even a relatively short video can be substantial. This means that data can be streamed across communication networks with limited bandwidth capacity. This may pose difficulties when the information is to be communicated or otherwise transmitted. Video data is typically compressed before being communicated across modern telecommunications networks. Since memory resources may be limited, videos are stored on storage devices. Even when the size of the video is important, video compression devices often Use software and / or hardware to decode video prior to transmission or storage. It codes the audio data needed to represent a digital video image. The compressed data is then passed to a video processor that decodes the video data. The data is received at the destination by the recovery device. With the ever-increasing demand for higher video quality, Improved compression and decompression techniques that improve compression ratios are desirable. Summary of the Invention [Means for solving the problem]
[0004] The first aspect relates to a method of decoding a coded video bit stream, which is performed by a video decoder. The method includes steps of: determining by the video decoder whether a value for a first flag is provided by an external input; when the value for the first flag is provided by the external input, setting, by the video decoder, the first flag to be equal to the value provided by the external input and setting a second flag to be equal to the first flag, to prevent a gradual decoding refresh (GDR) picture in the coded video bit stream from being output; decoding, by the video decoder, the GDR picture; and storing, by the video decoder, the GDR picture in a decoded picture buffer (DPB). The method provides a technique that enables intra refresh to enable random access without the need to use an intra random access point (IRAP) picture. When the value for the first flag is provided by the external input, the first flag is set to be equal to the value provided by the external input and the second flag is set to be equal to the first flag, to prevent a gradual decoding refresh (GDR) picture from being output. The external input may be an input received from a user (e.g., a network administrator) via, for example, a graphic user interface (GUI) of the video decoder. The first aspect relates to a method of decoding a coded video bit stream, which is performed by a video decoder. The method includes steps of: determining by the video decoder whether a value for a first flag is provided by an external input; when the value for the first flag is provided by the external input, setting, by the video decoder, the first flag to be equal to the value provided by the external input and setting a second flag to be equal to the first flag, to prevent a gradual decoding refresh (GDR) picture in the coded video bit stream from being output; decoding, by the video decoder, the GDR picture; and storing, by the video decoder, the GDR picture in a decoded picture buffer (DPB). The method provides a technique that enables intra refresh to enable random access without the need to use an intra random access point (IRAP) picture. When the value for the first flag is provided by the external input, the first flag is set to be equal to the value provided by the external input and the second flag is set to be equal to the first flag, to prevent a gradual decoding refresh (GDR) picture from being output. The external input may be an input received from a user (e.g., a network administrator) via, for example, a graphic user interface (GUI) of the video decoder. The first aspect relates to a method of decoding a coded video bit stream, which is performed by a video decoder. The method includes steps of: determining by the video decoder whether a value for a first flag is provided by an external input; when the value for the first flag is provided by the external input, setting, by the video decoder, the first flag to be equal to the value provided by the external input and setting a second flag to be equal to the first flag, to prevent a gradual decoding refresh (GDR) picture in the coded video bit stream from being output; decoding, by the video decoder, the GDR picture; and storing, by the video decoder, the GDR picture in a decoded picture buffer (DPB). The method provides a technique that enables intra refresh to enable random access without the need to use an intra random access point (IRAP) picture. When the value for the first flag is provided by the external input, the first flag is set to be equal to the value provided by the external input and the second flag is set to be equal to the first flag, to prevent a gradual decoding refresh (GDR) picture from being output. The external input may be an input received from a user (e.g., a network administrator) via, for example, a graphic user interface (GUI) of the video decoder. The first aspect relates to a method of decoding a coded video bit stream, which is performed by a video decoder. The method includes steps of: determining by the video decoder whether a value for a first flag is provided by an external input; when the value for the first flag is provided by the external input, setting, by the video decoder, the first flag to be equal to the value provided by the external input and setting a second flag to be equal to the first flag, to prevent a gradual decoding refresh (GDR) picture in the coded video bit stream from being output; decoding, by the video decoder, the GDR picture; and storing, by the video decoder, the GDR picture in a decoded picture buffer (DPB). The method provides a technique that enables intra refresh to enable random access without the need to use an intra random access point (IRAP) picture. When the value for the first flag is provided by the external input, the first flag is set to be equal to the value provided by the external input and the second flag is set to be equal to the first flag, to prevent a gradual decoding refresh (GDR) picture from being output. The external input may be an input received from a user (e.g., a network administrator) via, for example, a graphic user interface (GUI) of the video decoder. The first aspect relates to a method of decoding a coded video bit stream, which is performed by a video decoder. The method includes steps of: determining by the video decoder whether a value for a first flag is provided by an external input; when the value for the first flag is provided by the external input, setting, by the video decoder, the first flag to be equal to the value provided by the external input and setting a second flag to be equal to the first flag, to prevent a gradual decoding refresh (GDR) picture in the coded video bit stream from being output; decoding, by the video decoder, the GDR picture; and storing, by the video decoder, the GDR picture in a decoded picture buffer (DPB). The method provides a technique that enables intra refresh to enable random access without the need to use an intra random access point (IRAP) picture. When the value for the first flag is provided by the external input, the first flag is set to be equal to the value provided by the external input and the second flag is set to be equal to the first flag, to prevent a gradual decoding refresh (GDR) picture from being output. The external input may be an input received from a user (e.g., a network administrator) via, for example, a graphic user interface (GUI) of the video decoder.
[0005] The method provides a technique that enables intra refresh to enable random access without the need to use an intra random access point (IRAP) picture. When the value for the first flag is provided by the external input, the first flag is set to be equal to the value provided by the external input and the second flag is set to be equal to the first flag, to prevent a gradual decoding refresh (GDR) picture from being output. The external input may be an input received from a user (e.g., a network administrator) via, for example, a graphic user interface (GUI) of the video decoder. The method provides a technique that enables intra refresh to enable random access without the need to use an intra random access point (IRAP) picture. When the value for the first flag is provided by the external input, the first flag is set to be equal to the value provided by the external input and the second flag is set to be equal to the first flag, to prevent a gradual decoding refresh (GDR) picture from being output. The external input may be an input received from a user (e.g., a network administrator) via, for example, a graphic user interface (GUI) of the video decoder. The method provides a technique that enables intra refresh to enable random access without the need to use an intra random access point (IRAP) picture. When the value for the first flag is provided by the external input, the first flag is set to be equal to the value provided by the external input and the second flag is set to be equal to the first flag, to prevent a gradual decoding refresh (GDR) picture from being output. The external input may be an input received from a user (e.g., a network administrator) via, for example, a graphic user interface (GUI) of the video decoder. The method provides a technique that enables intra refresh to enable random access without the need to use an intra random access point (IRAP) picture. When the value for the first flag is provided by the external input, the first flag is set to be equal to the value provided by the external input and the second flag is set to be equal to the first flag, to prevent a gradual decoding refresh (GDR) picture from being output. The external input may be an input received from a user (e.g., a network administrator) via, for example, a graphic user interface (GUI) of the video decoder. The method provides a technique that enables intra refresh to enable random access without the need to use an intra random access point (IRAP) picture. When the value for the first flag is provided by the external input, the first flag is set to be equal to the value provided by the external input and the second flag is set to be equal to the first flag, to prevent a gradual decoding refresh (GDR) picture from being output. The external input may be an input received from a user (e.g., a network administrator) via, for example, a graphic user interface (GUI) of the video decoder. The method provides a technique that enables intra refresh to enable random access without the need to use an intra random access point (IRAP) picture. When the value for the first flag is provided by the external input, the first flag is set to be equal to the value provided by the external input and the second flag is set to be equal to the first flag, to prevent a gradual decoding refresh (GDR) picture from being output. The external input may be an input received from a user (e.g., a network administrator) via, for example, a graphic user interface (GUI) of the video decoder. The method provides a technique that enables intra refresh to enable random access without the need to use an intra random access point (IRAP) picture. When the value for the first flag is provided by the external input, the first flag is set to be equal to the value provided by the external input and the second flag is set to be equal to the first flag, to prevent a gradual decoding refresh (GDR) picture from being output. The external input may be an input received from a user (e.g., a network administrator) via, for example, a graphic user interface (GUI) of the video decoder. Setting the first and second flags in this way prevents potentially dirty data from being output to the display. That is, the values of the first and second flags control whether potentially dirty data from the GDR picture is output, or whether the video decoder starts displaying data waiting for full synchronization. By having the ability to limit the output of dirty data, the coder / decoder (also called "codec") in video coding is improved compared to the current codec. Practically speaking, an improved video coding process gives the user a better user experience when the video is sent, received, and / or viewed. In a first implementation of the method according to the first aspect itself, when the value for the first flag is provided by an external input, to prevent the progressive decode refresh (GDR) picture and any trailing pictures between the GDR picture and the recovery point picture in the output order from being output, the first flag is equal to the value provided by the external input, and the second flag is set equal to the first flag. In a second implementation of the method according to the first aspect itself or any preceding implementation of the first aspect, the external input is in the graphical user interface (GUI) of the video decoder, and the value of the first flag is provided by the user of the video decoder using the external input. In a third implementation of the method according to the first aspect itself or any preceding implementation of the first aspect,
[0006]
[0007]
[0008]
[0007] In a second implementation of the method according to the first aspect itself or any preceding implementation of the first aspect, the external input is in the graphical user interface (GUI) of the video decoder, and the value of the first flag is provided by the user of the video decoder using the external input. In a second implementation of the method according to the first aspect itself or any preceding implementation of the first aspect, the external input is in the graphical user interface (GUI) of the video decoder, and the value of the first flag is provided by the user of the video decoder using the external input. In a second implementation of the method according to the first aspect itself or any preceding implementation of the first aspect, the external input is in the graphical user interface (GUI) of the video decoder, and the value of the first flag is provided by the user of the video decoder using the external input. In a second implementation of the method according to the first aspect itself or any preceding implementation of the first aspect, the external input is in the graphical user interface (GUI) of the video decoder, and the value of the first flag is provided by the user of the video decoder using the external input.
[0008] In a third implementation of the method according to the first aspect itself or any preceding implementation of the first aspect, In a form, the first flag is designated as HandleGdrAsCvsStartFlag.
[0009] A fourth implementation of a method according to the first aspect itself or any preceding implementation form of the first aspect In a form, to prevent any trailing pictures between the GDR picture and the recovery point picture from being output in the GDR picture and output order, the values of the first flag and the second flag are set to 1.
[0010] A fifth implementation of a method according to the first aspect itself or any preceding implementation form of the first aspect In a form, when the value for the first flag is not provided by an external input, the value of the first flag is set to 0.
[0011] The second aspect relates to a decoding device. The decoding device includes a receiver configured to receive a coded video bit stream, a memory coupled to the receiver for storing instructions, and a processor coupled to the memory, and the processor is configured to execute instructions to determine for the decoding device whether the value for the first flag is provided by an external input, and when the value for the first flag is provided by an external input, to set the first flag equal to the value provided by the external input and the second flag equal to the first flag, to prevent a progressive decoding refresh (GDR) picture from being output, to decode the GDR picture, and to store the GDR picture in a decoded picture buffer (DPB).
[0012] The decoding device does not need to use an intra random access point (IRAP) picture Provided is a technique that enables sequential intra refresh to enable random access. When a value for a first flag is provided by an external input, in order to prevent a progressive decoded refresh (GDR) picture from being output, the first flag is set equal to the value provided by the external input, and a second flag is set equal to the first flag. The external input may be, for example, an input received from a user (such as a network administrator) via a graphic user interface (GUI) of a video decoder. Setting the first and second flags in this way potentially prevents dirty data from being output to the display. That is, the values of the first and second flags control whether potentially dirty data from the GDR picture is output or whether the video decoder waits for full synchronization before starting to display the data. By having the ability to limit the output of dirty data, the coder / decoder (also referred to as a “codec”) in video coding is improved compared to the current codec. In practice, the improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed. In a first implementation form of a decoding device according to a second aspect itself, when a value for a first flag is provided by an external input, a progressive decoded refresh (GDR) picture and any trailing pictures between the GDR picture and a recovery point picture in output order are prevented from being output. The first flag is set equal to the value provided by the external input, and a second flag is set equal to the first flag. The external input may be, for example, an input received from a user (such as a network administrator) via a graphic user interface (GUI) of a video decoder. Setting the first and second flags in this way potentially prevents dirty data from being output to the display. That is, the values of the first and second flags control whether potentially dirty data from the GDR picture is output or whether the video decoder waits for full synchronization before starting to display the data. By having the ability to limit the output of dirty data, the coder / decoder (also referred to as a “codec”) in video coding is improved compared to the current codec. In practice, the improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed. In a first implementation form of a decoding device according to a second aspect itself, when a value for a first flag is provided by an external input, a progressive decoded refresh (GDR) picture and any trailing pictures between the GDR picture and a recovery point picture in output order are prevented from being output. The first flag is set equal to the value provided by the external input, and a second flag is set equal to the first flag. The external input may be, for example, an input received from a user (such as a network administrator) via a graphic user interface (GUI) of a video decoder.
[0013] In a first implementation form of a decoding device according to a second aspect itself, when a value for a first flag is provided by an external input, a progressive decoded refresh (GDR) picture and any trailing pictures between the GDR picture and a recovery point picture in output order are prevented from being output. To prevent it from being output, the first flag is equal to the value provided by an external input, and the second flag is set equal to the first flag.
[0014] In a second implementation form of the decoding device according to the second aspect itself, the external input is the graphic user interface (GUI) of the video decoder, and the value of the first flag is provided by the user of the video decoder using the external input.
[0015] In a third implementation form of the decoding device according to the second aspect itself, the first flag is designated as HandleGdr AsCvsStartFlag.
[0016] In a fourth implementation form of the decoding device according to the second aspect itself, to prevent any trailing pictures between the GDR picture and the recovery point picture from being output in the GDR picture and output order, the values of the first flag and the second flag are set to 1.
[0017] In a fifth implementation form of the decoding device according to the second aspect itself, when the value for the first flag is not provided by an external input, the value of the first flag is set to 0.
[0018] The third aspect relates to an encoding device. The encoding device includes a receiver configured to receive a picture to be encoded or a bitstream to be decoded, a transmitter coupled to the receiver and configured to transmit the bitstream to a decoder or transmit a decoded image to a display, and a memory coupled to at least one of the receiver or the transmitter and configured to store instructions, A memory and a processor coupled to the memory, the processor being configured to execute instructions stored in the memory to perform any of the methods disclosed herein including the processor.
[0019] The coding device provides a technique that enables sequential intra-refresh to enable random access without the need to use an intra-random access point (IRAP) picture. When a value for a first flag is provided by an external input, the first flag is set equal to the value provided by the external input and the second flag is set equal to the first flag to prevent a progressive decode refresh (GDR) picture from being output. The external input may be, for example, an input received from a user (e.g., a network administrator) via a graphic user interface (GUI) of a video decoder. Setting the first and second flags in this way prevents potentially dirty data from being output to the display. That is, the values of the first and second flags control whether potentially dirty data from the GDR picture is output or whether the video decoder waits for full synchronization before starting to display the data. By having the ability to limit the output of dirty data, the coder / decoder (also referred to as a "codec") in the video coding is improved compared to the current codec. In practice, the improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed.
[0020] In a first implementation form of the coding device according to the third aspect itself, it further includes a display configured to display an image.
[0021] The fourth aspect relates to a system. The system includes an encoder and a decoder communicating with the encoder, and the encoder or the decoder includes a decoding device, an encoding device, or a coding device disclosed herein.
[0022] The system provides a technique that enables sequential intra refresh to enable random access without the need to use an intra random access point (IRAP) picture. When a value for a first flag is provided by an external input, to prevent a gradual decoding refresh (GDR) picture from being output, the first flag is set equal to the value provided by the external input, and the second flag is set equal to the first flag. The external input may be, for example, an input received from a user (such as a network administrator) via a graphic user interface (GUI) of a video decoder. Setting the first and second flags in this way potentially prevents dirty data from being output to the display. That is, the values of the first and second flags control whether potentially dirty data from a GDR picture is output or whether the video decoder starts displaying data waiting for complete synchronization. By having the ability to limit the output of dirty data, a coder / decoder (also called a "codec") in video coding is improved compared to the current codec. This occurs. In practice, an improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed. This occurs. In practice, an improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed. This occurs. In practice, an improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed.
[0023] A fifth aspect relates to means for coding. The means for coding comprises receiving means configured to receive a picture to be coded or a bitstream to be decoded, transmitting means coupled to the receiving means and configured to transmit the bitstream to a decoding means or the decoded picture to a display means, storage means coupled to at least one of the receiving means or the transmitting means and configured to store instructions, and processing means coupled to the storage means and configured to execute instructions stored in the storage means to perform any of the methods disclosed herein. This occurs. In practice, an improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed. This occurs. In practice, an improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed. This occurs. In practice, an improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed. This occurs. In practice, an improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed. This occurs. In practice, an improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed. This occurs. In practice, an improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed. This occurs. In practice, an improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed.
[0024] The means for coding provides a technique that enables sequential intra refresh to enable random access without the need to use an intra random access point (IRAP) picture. When a value for a first flag is provided by an external input, the first flag is set equal to the value provided by the external input and a second flag is set equal to the first flag to prevent a gradual decoding refresh (GDR) picture from being output. The external input is received, for example, from a user (e.g., a network administrator) via a graphical user interface (GUI) of a video decoder. This occurs. In practice, an improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed. This occurs. In practice, an improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed. This occurs. In practice, an improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed. This occurs. In practice, an improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed. This occurs. In practice, an improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed. This occurs. In practice, an improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed. It may be an input. Setting the first and second flags in this way potentially prevents dirty data from being output to the display. That is, the values of the first and second flags control whether potentially dirty data from the GDR picture is output, or whether the video decoder starts displaying data waiting for complete synchronization. By having the ability to limit the output of dirty data, the coder / decoder (also called "codec") in video coding is improved compared to the current codec. As a practical matter, the improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed. For a more complete understanding of the present disclosure, reference is now made to the following brief description taken in conjunction with the accompanying drawings and the detailed description of the invention, in which like reference
[0025] numerals represent like parts.
Brief Description of the Drawings
Brief Description of the Drawings
[0026]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
DETAILED DESCRIPTION OF THE INVENTION
[0027] FIG. 1 is a block diagram showing an exemplary coding system 10 that may utilize video coding techniques as described herein. As shown in FIG. 1, the coding system 10 includes a source device 12 that provides encoded video data to be decoded later by a destination device 14. Specifically, the source device 12 may provide video data to the destination device 14 via a computer-readable medium 16. The source device 12 and the destination device 14 may include any of a wide range of devices, such as a desktop computer, a notebook (e.g., laptop) computer, a tablet computer, a set-top box, a so-called "smart" phone handset, a so-called "smart" pad, a television, a camera, a display device, a digital media player, a video game console, a video streaming device, etc. In some cases, the source device 12 and the destination device 14 may be equipped for wireless communication.
[0028] The destination device 14 may receive the encoded video data to be decoded via a computer-readable medium 16. The computer-readable medium 16 may comprise any type of medium or device capable of moving the encoded video data from the source device 12 to the destination device 14. In one example, the computer-readable medium 16 may comprise a communication medium to enable the source device 12 to directly transmit the encoded video data to the destination device 14 in real time. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and transmitted to the destination device 14. The communication medium may comprise any wireless or wired communication medium such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include a router, a switch, a base station, or any other device that may be useful to facilitate communication from the source device 12 to the destination device 14. In some examples, the encoded data may be output from the output interface 22 to a storage device. Similarly, the encoded data may be accessed from a storage device by an input interface. The storage device may be a hard drive, a Blu-ray disc, a digital video disc (DVD), a compact disc read-only memory (CD-ROM), a flash memory,
[0029] volatile or non-volatile memory, or any other medium or device capable of storing the encoded video data. Likewise, the encoded data may be accessed from a storage device by an input interface. The storage device may be a hard drive, a Blu-ray disc, a digital video disc (DVD), a compact disc read-only memory (CD-ROM), a flash memory, volatile or non-volatile memory, or any other medium or device capable of storing the encoded video data. for storing the encoded video data. Any other suitable digital storage media, such as distributed or locally accessible among various data storage media. In a further example, the storage device may correspond to a file server or another intermediate storage device that stores the encoded video generated by the source device 12. The destination device 14 may access the stored video data from the storage device via streaming or download. The file server may be any type of server capable of storing the encoded video data and transmitting the encoded video data to the destination device 14. Exemplary file servers include web servers (e.g., for websites), File Transfer Protocol (FTP) servers, Network Attached Storage (NAS) devices, or local disk drives. The destination device 14 may access the encoded video data through any standard data connection including an Internet connection. This may include a wireless channel (e.g., Wi-Fi connection), a wired connection (e.g., Digital Subscriber Line (DSL), cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage device may be by streaming, download, or a combination thereof. The techniques of the present disclosure are not necessarily limited to wireless
[0030] applications or settings. The techniques are applicable to over-the-air television broadcasting, cable television transmission, satellite television transmission, dynamic Internet streaming video transmission such as HTTP streaming (DASH), data - Digital video encoded on a storage medium, digital Video decoded from video stored on a storage medium, or any of various It may be applied to video coding that supports multimedia applications such as. In some examples, the coding System 10 may be configured to support one-way or two-way video Transmission to support applications such as streaming, video playback, video broadcasting, And / or video telephony.
[0031] In the example of FIG. 1, the source device 12 includes a video source 18, a video encoder 20, and an output Interface 22. The destination device 14 includes an input interface 28, a video decoder 30, and a display device 32. According to the present disclosure, the video encoder 20 of the source device 12 And / or the video decoder 30 of the destination device 14 may be configured to apply techniques for video coding. In other examples, the source device and The destination device may include other components or configurations. For example, the source device 12 may receive video data from an external video source such as an external camera. Similarly The destination device 14 may interface with an external display device rather than including an integrated display device.
[0032] The illustrated coding system 10 of FIG. 1 is merely an example. Techniques for video coding Are performed by any digital video encoding and / or decoding device This is also possible. The techniques of the present disclosure are generally performed by a video coding device, but the techniques may also be performed by a video encoder / decoder, often referred to as a "codec". Moreover, the techniques of the present disclosure may also be performed by a video preprocessor. The video encoder and / or decoder may be a graphics processing unit (GPU) or a similar device.
[0033] Source device 12 and destination device 14 are merely examples of such coding devices that generate encoded video data for transmission to destination device 14 by source device 12. In some examples, source device 12 and destination device 14 may operate substantially symmetrically such that each of source device 12 and destination device 14 includes video encoding and decoding components. Accordingly, coding system 10 may support one-way or two-way video transmission between video devices 12, 14, for example, for video streaming, video playback, video broadcasting, or video telephony. The video source 18 of source device 12 may include a video capture device such as a video camera, a video archive containing previously captured video, and / or a video feed interface for receiving video from a video content provider. As a further alternative, video source 18 may generate a combination of computer graphics-based data, live video, archived video, and computer-generated video as source video. For example, for video streaming, video playback, video broadcasting, or video telephony. Support one-way or two-way video transmission between video devices 12, 14. It may also be supported.
[0034] The video source 18 of source device 12 may include a video capture device such as a video camera, a video archive containing previously captured video, and / or a video feed interface for receiving video from a video content provider. Including a video archive containing previously captured video, and / or a video feed interface for receiving video from a video content provider. It may also include a video capture device such as a video camera. As a further alternative, video source 18 may generate a combination of computer graphics-based data, live video, archived video, and computer-generated video as source video. Or a combination of live video, archived video, and computer-generated video. It may also be generated.
[0035] In some cases, when the video source 18 is a video camera, the source device 12 and the destination device 14 may form a so-called camera phone or video phone. However as described above, the techniques described in this disclosure are generally applicable to video coding and may also be applicable to wireless and / or wired applications. In each case captured, pre-captured, or computer-generated video may be encoded by the video encoder 20. The encoded video information may then be output onto the computer-readable medium 16 by the output i nterface 22.
[0036] The computer-readable medium 16 may include a temporary medium such as a wireless broadcast or wired network tra nsmission, or a storage medium (i.e., a non-temporary storage medium) such as a hard disk, flash drive, compact disk , digital video disk, Blu-ray disk, or other computer-readable medium. In some examples, a network work server (not shown) may receive encoded video data from the source device 12, for example, via a network transmission, and provide the encoded video data to the destination device 14. Similarly, a computing device of a media manufacturing facility such as a disk stamping facility may receive encoded video data from the source device 12 and manufacture a disk containing the encoded video data. Thus, in various examples, the computer-readable medium 16 may be understood to include one or more computer-readable media in various forms . .
[0037] The input interface 28 of the destination device 14 receives information from the computer-readable medium 16. The information on the computer-readable medium 16 may include syntax information defined by the video encoder 20, including blocks and other coded units, for example, characteristics of a group of pictures (GOP) and / or syntax elements that describe processing. The syntax information may be used by the video decoder 30. The display device 32 displays the decoded video data to the user and may include various display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.
[0038] The video encoder 20 and the video decoder 30 may operate according to a video coding standard such as the currently developing High Efficiency Video Coding (HEVC) standard and may conform to the HEVC Test Model (HM). Alternatively, the video encoder 20 and the video decoder 30 may operate according to other proprietary or industry standards, such as the International Telecommunication Union Telecommunication Standardization Sector (ITU-T) H.264 standard, also known as Moving Picture Experts Group (MPEG)-4 Part 10, Advanced Video Coding (AVC), H.265 / HEVC, or extensions of such standards. However, the techniques of the present disclosure are not limited to any particular coding standard. Other examples of video coding standards include MPEG-2 and ITU- It includes H.263. Although not shown in FIG. 1, in some embodiments, video encoder 20 and video decoder 30 may each be integrated with an audio encoder and decoder, and a suitable multiplexer-demultiplexer (MUX-DEMUX) unit, or other hardware and software, for processing the encoding of both audio and video in a common data stream or separate data streams may be included. When applicable, the MUX-DEMUX unit may conform to the ITU H.223 multiplexer protocol, or other protocols such as the User Datagram Protocol (UDP).
[0039] Video encoder 20 and video decoder 30 may each be implemented as any of a variety of suitable encoder circuit configurations, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the techniques herein are implemented partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable medium and may execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Each of video encoder 20 and video decoder 30 may be included within one or more encoders or decoders, and any of them may be integrated within their respective devices as part of a combined encoder / decoder (codec). Devices It may be provided with a road, a microprocessor, and / or a wireless communication device such as a cellular phone. It may be provided.
[0040] Figure 2 shows a block diagram of an example of a video encoder 20 that may implement video coding techniques. The video encoder 20 may perform intra and inter coding of video blocks within a video slice. Intra coding relies on spatial prediction to reduce or remove spatial redundancy in the video within a given video frame or picture. Inter coding relies on temporal prediction to reduce or remove temporal redundancy in the video within adjacent frames or pictures in a video sequence. The intra mode (I mode) may refer to any of several spatial-based coding modes. The inter mode, such as unidirectional prediction (also called single prediction) (P mode) or bi-prediction (also called bi prediction) (B mode), may refer to any of several time-based coding modes. It may be executed. To reduce or remove the spatial redundancy in the video within a given video frame or picture. It relies on spatial prediction. To reduce or remove the temporal redundancy in the video within adjacent frames or pictures in a video sequence. It relies on temporal prediction. The intra mode (I mode) may refer to any of several spatial-based coding modes. Unidirectional prediction (also called single prediction) (P mode) or bi-prediction (also called bi prediction) (B mode). The inter mode, such as unidirectional prediction (also called single prediction) (P mode) or bi-prediction (also called bi prediction) (B mode), may refer to any of several time-based coding modes. It may refer to any of them.
[0041] As shown in Figure 2, the video encoder 20 receives the current video block within the video frame to be encoded. In the example of Figure 2, the video encoder 20 includes a mode selection unit 40, a reference frame memory 64, an adder 50, a conversion processing unit 52, a quantization unit 54, and an entropy coding unit 56. Next, the mode selection unit 40 includes a motion compensation unit 44, a motion estimation unit 42, and an intra-prediction unit. In the example of Figure 2, the video encoder 20 includes a mode selection unit 40. A reference frame memory 64, an adder 50, a conversion processing unit 52, a quantization unit 54, and an entropy coding unit 56. Next, the mode selection unit 40 includes a motion compensation unit 44. A motion estimation unit 42, and an intra-prediction unit. a unit 46 (also called a prediction), and a partitioning unit 48. A video block For reconstruction, the video encoder 20 also includes an inverse quantization unit 58, an inverse transform unit 60, and an adder 62. To filter block boundaries to remove blocking artifacts from the reconstructed video, a deblocking filter (not shown in FIG. 2) may also be included as well. If desired, the deblocking filter will typically filter the output of the adder 62 In addition to the deblocking filter, additional filters (either within or after the loop) may also be used. Such filters are not shown for simplicity but, if desired, may filter the output of the adder 50 (as a loop filter) as well.
[0042] During the encoding process, the video encoder 20 receives a video frame or slice to be encoded The frame or slice may be divided into a plurality of video blocks The motion estimation unit 42 and the motion compensation unit 44 perform inter-prediction coding of the received video block with respect to one or more blocks in one or more reference frames to perform temporal prediction. The intra prediction unit 46 may alternatively perform intra-prediction coding of the received video block with respect to one or more adjacent blocks in the same frame or slice as the block to be encoded to perform spatial prediction. The video encoder 20 may execute a plurality of coding paths to select an appropriate coding mode for each block of the video data for example.
[0043] In addition, the partitioning unit 48 may partition a block of video data into sub-blocks based on the evaluation of the previous partitioning method in the previous coding path. For example, the partitioning unit 48 may first partition a frame or a slice into largest coding units (LCUs), and then partition each LCU into sub-coding units (sub-CUs) based on rate distortion analysis (e.g., rate distortion optimization). The mode selection unit 40 may further generate a quadtree data structure indicating the partitioning of the LCU into sub-CUs. The leaf node CUs of the quadtree may include one or more prediction units (PUs) and one or more transform units (TUs). The present disclosure uses the term "block" to refer to any of a CU, a PU, or a TU in the context of HEVC, or a similar data structure (e.g., macroblocks and their sub-blocks in H.264 / AVC) in the context of other standards. A CU includes a coding node, a PU, and a TU associated with the coding node. The size of the CU corresponds to the size of the coding node and is square in shape. The size of the CU may range from 8×8 pixels to a tree block size with a maximum value of 64×64 pixels or more. Each CU may include one or more PUs and one or more TUs. The syntax data associated with the CU may describe, for example, the partitioning of the CU into one or more PUs. Whether the CU is coded in skip mode or direct mode, or coded in intra prediction mode Furthermore, the partitioning unit 48 may partition a block of video data into sub-blocks based on the evaluation of the previous partitioning method in the previous coding path. For example, the partitioning unit 48 may first partition a frame or a slice into largest coding units (LCUs), and then partition each LCU into sub-coding units (sub-CUs) based on rate distortion analysis (e.g., rate distortion optimization). The mode selection unit 40 may further generate a quadtree data structure indicating the partitioning of the LCU into sub-CUs. The leaf node CUs of the quadtree may include one or more prediction units (PUs) and one or more transform units (TUs). Furthermore, the partitioning unit 48 may partition a block of video data into sub-blocks based on the evaluation of the previous partitioning method in the previous coding path. For example, the partitioning unit 48 may first partition a frame or a slice into largest coding units (LCUs), and then partition each LCU into sub-coding units (sub-CUs) based on rate distortion analysis (e.g., rate distortion optimization). The mode selection unit 40 may further generate a quadtree data structure indicating the partitioning of the LCU into sub-CUs. The leaf node CUs of the quadtree may include one or more prediction units (PUs) and one or more transform units (TUs). Furthermore, the partitioning unit 48 may partition a block of video data into sub-blocks based on the evaluation of the previous partitioning method in the previous coding path. For example, the partitioning unit 48 may first partition a frame or a slice into largest coding units (LCUs), and then partition each LCU into sub-coding units (sub-CUs) based on rate distortion analysis (e.g., rate distortion optimization). The mode selection unit 40 may further generate a quadtree data structure indicating the partitioning of the LCU into sub-CUs. The leaf node CUs of the quadtree may include one or more prediction units (PUs) and one or more transform units (TUs). Furthermore, the partitioning unit 48 may partition a block of video data into sub-blocks based on the evaluation of the previous partitioning method in the previous coding path. For example, the partitioning unit 48 may first partition a frame or a slice into largest coding units (LCUs), and then partition each LCU into sub-coding units (sub-CUs) based on rate distortion analysis (e.g., rate distortion optimization). The mode selection unit 40 may further generate a quadtree data structure indicating the partitioning of the LCU into sub-CUs. The leaf node CUs of the quadtree may include one or more prediction units (PUs) and one or more transform units (TUs). Furthermore, the partitioning unit 48 may partition a block of video data into sub-blocks based on the evaluation of the previous partitioning method in the previous coding path. For example, the partitioning unit 48 may first partition a frame or a slice into largest coding units (LCUs), and then partition each LCU into sub-coding units (sub-CUs) based on rate distortion analysis (e.g., rate distortion optimization). The mode selection unit 40 may further generate a quadtree data structure indicating the partitioning of the LCU into sub-CUs. The leaf node CUs of the quadtree may include one or more prediction units (PUs) and one or more transform units (TUs). Furthermore, the partitioning unit 48 may partition a block of video data into sub-blocks based on the evaluation of the previous partitioning method in the previous coding path. For example, the partitioning unit 48 may first partition a frame or a slice into largest coding units (LCUs), and then partition each LCU into sub-coding units (sub-CUs) based on rate distortion analysis (e.g., rate distortion optimization). The mode selection unit 40 may further generate a quadtree data structure indicating the partitioning of the LCU into sub-CUs. The leaf node CUs of the quadtree may include one or more prediction units (PUs) and one or more transform units (TUs).
[0044] The present disclosure uses the term "block" to refer to any of a CU, a PU, or a TU in the context of HEVC, or a similar data structure (e.g., macroblocks and their sub-blocks in H.264 / AVC) in the context of other standards. A CU includes a coding node, a PU, and a TU associated with the coding node. The size of the CU corresponds to the size of the coding node and is square in shape. The size of the CU may range from 8×8 pixels to a tree block size with a maximum value of 64×64 pixels or more. Each CU may include one or more PUs and one or more TUs. The syntax data associated with the CU may describe, for example, the partitioning of the CU into one or more PUs. Whether the CU is coded in skip mode or direct mode, or coded in intra prediction mode The present disclosure uses the term "block" to refer to any of a CU, a PU, or a TU in the context of HEVC, or a similar data structure (e.g., macroblocks and their sub-blocks in H.264 / AVC) in the context of other standards. A CU includes a coding node, a PU, and a TU associated with the coding node. The size of the CU corresponds to the size of the coding node and is square in shape. The size of the CU may range from 8×8 pixels to a tree block size with a maximum value of 64×64 pixels or more. Each CU may include one or more PUs and one or more TUs. The syntax data associated with the CU may describe, for example, the partitioning of the CU into one or more PUs. Whether the CU is coded in skip mode or direct mode, or coded in intra prediction mode The present disclosure uses the term "block" to refer to any of a CU, a PU, or a TU in the context of HEVC, or a similar data structure (e.g., macroblocks and their sub-blocks in H.264 / AVC) in the context of other standards. A CU includes a coding node, a PU, and a TU associated with the coding node. The size of the CU corresponds to the size of the coding node and is square in shape. The size of the CU may range from 8×8 pixels to a tree block size with a maximum value of 64×64 pixels or more. Each CU may include one or more PUs and one or more TUs. The syntax data associated with the CU may describe, for example, the partitioning of the CU into one or more PUs. Whether the CU is coded in skip mode or direct mode, or coded in intra prediction mode The present disclosure uses the term "block" to refer to any of a CU, a PU, or a TU in the context of HEVC, or a similar data structure (e.g., macroblocks and their sub-blocks in H.264 / AVC) in the context of other standards. A CU includes a coding node, a PU, and a TU associated with the coding node. The size of the CU corresponds to the size of the coding node and is square in shape. The size of the CU may range from 8×8 pixels to a tree block size with a maximum value of 64×64 pixels or more. Each CU may include one or more PUs and one or more TUs. The syntax data associated with the CU may describe, for example, the partitioning of the CU into one or more PUs. Whether the CU is coded in skip mode or direct mode, or coded in intra prediction mode The present disclosure uses the term "block" to refer to any of a CU, a PU, or a TU in the context of HEVC, or a similar data structure (e.g., macroblocks and their sub-blocks in H.264 / AVC) in the context of other standards. A CU includes a coding node, a PU, and a TU associated with the coding node. The size of the CU corresponds to the size of the coding node and is square in shape. The size of the CU may range from 8×8 pixels to a tree block size with a maximum value of 64×64 pixels or more. Each CU may include one or more PUs and one or more TUs. The syntax data associated with the CU may describe, for example, the partitioning of the CU into one or more PUs. Whether the CU is coded in skip mode or direct mode, or coded in intra prediction mode The present disclosure uses the term "block" to refer to any of a CU, a PU, or a TU in the context of HEVC, or a similar data structure (e.g., macroblocks and their sub-blocks in H.264 / AVC) in the context of other standards. A CU includes a coding node, a PU, and a TU associated with the coding node. The size of the CU corresponds to the size of the coding node and is square in shape. The size of the CU may range from 8×8 pixels to a tree block size with a maximum value of 64×64 pixels or more. Each CU may include one or more PUs and one or more TUs. The syntax data associated with the CU may describe, for example, the partitioning of the CU into one or more PUs. Whether the CU is coded in skip mode or direct mode, or coded in intra prediction mode The present disclosure uses the term "block" to refer to any of a CU, a PU, or a TU in the context of HEVC, or a similar data structure (e.g., macroblocks and their sub-blocks in H.264 / AVC) in the context of other standards. A CU includes a coding node, a PU, and a TU associated with the coding node. The size of the CU corresponds to the size of the coding node and is square in shape. The size of the CU may range from 8×8 pixels to a tree block size with a maximum value of 64×64 pixels or more. Each CU may include one or more PUs and one or more TUs. The syntax data associated with the CU may describe, for example, the partitioning of the CU into one or more PUs. Whether the CU is coded in skip mode or direct mode, or coded in intra prediction mode The present disclosure uses the term "block" to refer to any of a CU, a PU, or a TU in the context of HEVC, or a similar data structure (e.g., macroblocks and their sub-blocks in H.264 / AVC) in the context of other standards. A CU includes a coding node, a PU, and a TU associated with the coding node. The size of the CU corresponds to the size of the coding node and is square in shape. The size of the CU may range from 8×8 pixels to a tree block size with a maximum value of 64×64 pixels or more. Each CU may include one or more PUs and one or more TUs. The syntax data associated with the CU may describe, for example, the partitioning of the CU into one or more PUs. Whether the CU is coded in skip mode or direct mode, or coded in intra prediction mode The present disclosure uses the term "block" to refer to any of a CU, a PU, or a TU in the context of HEVC, or a similar data structure (e.g., macroblocks and their sub-blocks in H.264 / AVC) in the context of other standards. A CU includes a coding node, a PU, and a TU associated with the coding node. The size of the CU corresponds to the size of the coding node and is square in shape. The size of the CU may range from 8×8 pixels to a tree block size with a maximum value of 64×64 pixels or more. Each CU may include one or more PUs and one or more TUs. The syntax data associated with the CU may describe, for example, the partitioning of the CU into one or more PUs. Whether the CU is coded in skip mode or direct mode, or coded in intra prediction mode whether it is coded in the intra-prediction (also called inter prediction) mode The partition mode may be different between whether it is coded in the intra-prediction (also called inter prediction) mode. The PU may be partitioned such that the shape becomes non-square. The syntax data related to the CU may also describe, for example, the partitioning of the CU into one or more TUs by a quadtree. The TU may have a square or non-square (e.g., rectangular) shape.
[0045] The mode selection unit 40 may select, for example, one of the coding modes, i.e., the intra mode or the inter mode, based on the error result, and provide the obtained intra-coded or inter-coded block to the adder 50 to generate residual block data, and to the adder 62 to reconstruct the coded block for use as a reference frame. The mode selection unit 40 also provides syntax elements such as motion vectors, intra mode indicators, partition information, and other such syntax information to the entropy coding unit 56. The motion estimation unit 42 and the motion compensation unit 44 may be highly integrated but are shown separately for conceptual purposes. The motion estimation performed by the motion estimation unit 42 is a process of generating a motion vector that estimates the motion of a video block. The motion vector is, for example, within the current frame (or other coded unit) during coding of the current block, compared to a predicted block within the reference frame (or other coded unit), within the current video frame or current picture
[0046] The motion estimation unit 42 and the motion compensation unit 44 may be highly integrated but are shown separately for conceptual purposes. The motion estimation performed by the motion estimation unit 42 is a process of generating a motion vector that estimates the motion of a video block. The motion vector is, for example, within the current frame (or other coded unit) during coding of the current block, compared to a predicted block within the reference frame (or other coded unit), within the current video frame or current picture during coding of the current block in the current video frame or current picture, relative to the predicted block within the reference frame (or other coded unit) within the current video frame or current picture It may also show the displacement of the PU of the video block. A prediction block is a block found to closely match the block to be coded from the perspective of pixel difference, where the pixel difference may be determined by the sum of absolute differences (SAD), the sum of square differences (SSD), or other difference metrics. In some examples, the video encoder 20 may calculate the values for the sub-pixel positions of the reference picture stored in the reference frame memory 64. For example, the video encoder 20 may interpolate the values at the 1 / 4 pixel position, 1 / 8 pixel position, or other fractional pixel positions of the reference picture. Therefore, the motion estimation unit 42 may perform motion search for full pixel positions and fractional pixel positions, and output a motion vector with fractional pixel accuracy. The motion estimation unit 42 compares the position of the PU with the position of the prediction block of the reference picture to calculate the motion vector for the PU of the video block in the inter-coded slice. The reference picture may be selected from the first reference picture list (list 0) or the second
[0047] reference picture list (list 1), each of which identifies one or more reference pictures stored in the reference frame memory 64. The motion estimation unit 42 sends the calculated motion vector to the entropy coding unit 56 and the motion compensation unit 44. The motion compensation performed by the motion compensation unit 44 is determined by the motion estimation unit 42. reference picture list (list 1). The motion estimation unit 42 sends the calculated motion vector to the entropy coding unit 56 and the motion compensation unit 44. The motion vector is sent to the entropy coding unit 56 and the motion compensation unit 44.
[0048] The motion compensation performed by the motion compensation unit 44 is determined by the motion estimation unit 42. fetching or generating a prediction block based on the obtained motion vector may also be involved be. Again, the motion estimation unit 42 and the motion compensation unit 44 may be functionally integrated in some examples. Upon receiving a motion vector for the current video block's PU, the motion compensation unit 44 may identify the position of the prediction block pointed to by the motion vector within one of the reference picture lists. The adder 50 forms a residual video block by subtracting the pixel values of the prediction block from the pixel values of the current video block being coded, and forms pixel difference values as will be described below. Generally, the motion estimation unit 42 performs motion estimation on the luma component, and the motion compensation unit 44 uses the motion vector calculated based on the luma component for both the chroma and luma components. The mode selection unit 40 may also generate syntax elements related to the video block and the video slice for use by the video decoder 30 when decoding the video blocks of the video slice.
[0049] The intra prediction unit 46 may intra-predict the current block as an alternative to the inter prediction performed by the motion estimation unit 42 and the motion compensation unit 44 as described above. Specifically, the intra prediction unit 46 may determine the intra prediction mode to be used for encoding the current block. In some examples, the intra prediction unit 46 may encode the current block using various intra prediction modes, for example, during separate coding passes, and the intra prediction unit 46 (or, in some examples, the mode The mode selection unit 40 may select an appropriate intra prediction mode to be used from among the tested modes. For example, the intra prediction unit 46 may calculate rate distortion values for various tested intra prediction modes using rate distortion analysis, and select the intra prediction mode with the best rate distortion characteristics from among the tested modes. Rate distortion analysis generally determines the amount of distortion (i.e., error) between an encoded block and the original block that was encoded to generate the encoded block, as well as the bit rate (i.e., the number of bits) used to generate the encoded block.
[0050] The intra prediction unit 46 may calculate a ratio from the distortion and rate for various encoded blocks to determine which intra prediction mode presents the best rate distortion value for the block. using rate distortion analysis, and select the intra prediction mode with the best rate distortion characteristics from among the tested modes. Rate distortion analysis generally determines the amount of distortion (i.e., error) between an encoded block and the original block that was encoded to generate the encoded block, as well as the bit rate (i.e., the number of bits) used to generate the encoded block. from among the tested modes. Rate distortion analysis generally determines the amount of distortion (i.e., error) between an encoded block and the original block that was encoded to generate the encoded block, as well as the bit rate (i.e., the number of bits) used to generate the encoded block. The intra prediction unit 46 may calculate a ratio from the distortion and rate for various encoded blocks to determine which intra prediction mode presents the best rate distortion value for the block. analysis generally determines the amount of distortion (i.e., error) between an encoded block and the original block that was encoded to generate the encoded block, as well as the bit rate (i.e., the number of bits) used to generate the encoded block. The intra prediction unit 46 may calculate a ratio from the distortion and rate for various encoded blocks to determine which intra prediction mode presents the best rate distortion value for the block. from the distortion and rate for various encoded blocks to determine which intra prediction mode presents the best rate distortion value for the block. In addition, the intra prediction unit 46 may be configured to code depth blocks of a depth map using a depth modeling mode (DMM). The mode selection unit 40 may determine whether the available DMM mode generates better coding results than the intra prediction mode and other DMM modes, for example, using rate distortion optimization (RDO). Data for the texture image corresponding to the depth map may be stored in the reference frame memory 64. The motion estimation unit 42 and the motion compensation unit 44 may also inter
[0051] In addition, the intra prediction unit 46 may be configured to code depth blocks of a depth map using a depth modeling mode (DMM:depth modeling mode ) The mode selection unit 40 may determine whether the available DMM mode generates better coding results than the intra prediction mode and other DMM modes, for example, using rate distortion optimization (RDO). Data for the texture image corresponding to the depth map may be stored in the reference frame memory 64. The motion estimation unit 42 and the motion compensation unit 44 may also inter The mode selection unit 40 may determine whether the available DMM mode generates better coding results than the intra prediction mode and other DMM modes, for example, using rate distortion optimization (RDO). Data for the texture image corresponding to the depth map may be stored in the reference frame memory 64. The motion estimation unit 42 and the motion compensation unit 44 may also inter using rate distortion optimization (RDO). Data for the texture image corresponding to the depth map may be stored in the reference frame memory 64. The motion estimation unit 42 and the motion compensation unit 44 may also inter generate better coding results than the intra prediction mode and other DMM modes, for example, using rate distortion optimization (RDO). Data for the texture image corresponding to the depth map may be stored in the reference frame memory 64. The motion estimation unit 42 and the motion compensation unit 44 may also inter Data for the texture image corresponding to the depth map may be stored in the reference frame memory 64. The motion estimation unit 42 and the motion compensation unit 44 may also inter The motion estimation unit 42 and the motion compensation unit 44 may also inter It may be configured to predict.
[0052] After selecting an intra prediction mode for a block (e.g., one of a conventional intra prediction mode or a DMM mode), the intra prediction unit 46 may provide information indicating the selected intra prediction mode for the block to the entropy coding unit 56. After selecting an intra prediction mode for a block (e.g., one of a conventional intra prediction mode or a DMM mode), the intra prediction unit 46 may provide information indicating the selected intra prediction mode for the block to the entropy coding unit 56. After selecting an intra prediction mode for a block (e.g., one of a conventional intra prediction mode or a DMM mode), the intra prediction unit 46 may provide information indicating the selected intra prediction mode for the block to the entropy coding unit 56. The entropy coding unit 56 may encode the information indicating the selected intra prediction mode. The video encoder 20 may include a plurality of intra prediction mode index tables and a plurality of modified intra prediction mode index tables (also referred to as codeword mapping tables), definitions of coding contexts for various blocks, as well as the most accurate intra prediction mode to be used for each context, the intra prediction mode index table, and the display of the modified intra prediction mode index table in the transmitted bitstream. The entropy coding unit 56 may encode the information indicating the selected intra prediction mode. The video encoder 20 may include a plurality of intra prediction mode index tables and a plurality of modified intra prediction mode index tables (also referred to as codeword mapping tables), definitions of coding contexts for various blocks, as well as the most accurate intra prediction mode to be used for each context, the intra prediction mode index table, and the display of the modified intra prediction mode index table in the transmitted bitstream. The entropy coding unit 56 may encode the information indicating the selected intra prediction mode. The video encoder 20 may include a plurality of intra prediction mode index tables and a plurality of modified intra prediction mode index tables (also referred to as codeword mapping tables), definitions of coding contexts for various blocks, as well as the most accurate intra prediction mode to be used for each context, the intra prediction mode index table, and the display of the modified intra prediction mode index table in the transmitted bitstream. The entropy coding unit 56 may encode the information indicating the selected intra prediction mode. The video encoder 20 may include a plurality of intra prediction mode index tables and a plurality of modified intra prediction mode index tables (also referred to as codeword mapping tables), definitions of coding contexts for various blocks, as well as the most accurate intra prediction mode to be used for each context, the intra prediction mode index table, and the display of the modified intra prediction mode index table in the transmitted bitstream. The entropy coding unit 56 may encode the information indicating the selected intra prediction mode. The video encoder 20 may include a plurality of intra prediction mode index tables and a plurality of modified intra prediction mode index tables (also referred to as codeword mapping tables), definitions of coding contexts for various blocks, as well as the most accurate intra prediction mode to be used for each context, the intra prediction mode index table, and the display of the modified intra prediction mode index table in the transmitted bitstream. The entropy coding unit 56 may encode the information indicating the selected intra prediction mode. The video encoder 20 may include a plurality of intra prediction mode index tables and a plurality of modified intra prediction mode index tables (also referred to as codeword mapping tables), definitions of coding contexts for various blocks, as well as the most accurate intra prediction mode to be used for each context, the intra prediction mode index table, and the display of the modified intra prediction mode index table in the transmitted bitstream. The entropy coding unit 56 may encode the information indicating the selected intra prediction mode. The video encoder 20 may include a plurality of intra prediction mode index tables and a plurality of modified intra prediction mode index tables (also referred to as codeword mapping tables), definitions of coding contexts for various blocks, as well as the most accurate intra prediction mode to be used for each context, the intra prediction mode index table, and the display of the modified intra prediction mode index table in the transmitted bitstream. The entropy coding unit 56 may encode the information indicating the selected intra prediction mode. The video encoder 20 may include a plurality of intra prediction mode index tables and a plurality of modified intra prediction mode index tables (also referred to as codeword mapping tables), definitions of coding contexts for various blocks, as well as the most accurate intra prediction mode to be used for each context, the intra prediction mode index table, and the display of the modified intra prediction mode index table in the transmitted bitstream.
[0053] The video encoder 20 forms a residual video block by subtracting the prediction data from the mode selection unit 40 from the original video block during coding. The adder 50 represents one or more components that perform this subtraction operation. The video encoder 20 forms a residual video block by subtracting the prediction data from the mode selection unit 40 from the original video block during coding. The adder 50 represents one or more components that perform this subtraction operation. The video encoder 20 forms a residual video block by subtracting the prediction data from the mode selection unit 40 from the original video block during coding. The adder 50 represents one or more components that perform this subtraction operation.
[0054] The transform processing unit 52 applies a transform such as a discrete cosine transform (DCT) or a conceptually similar transform to the residual block to generate a video block with residual transform coefficient values. The transform processing unit 52 may perform other transforms conceptually similar to the DCT. Wavelet transform, integer transform, subband transform, or other types of transforms may also be used. The transform processing unit 52 applies a transform such as a discrete cosine transform (DCT) or a conceptually similar transform to the residual block to generate a video block with residual transform coefficient values. The transform processing unit 52 may perform other transforms conceptually similar to the DCT. Wavelet transform, integer transform, subband transform, or other types of transforms may also be used. The transform processing unit 52 applies a transform such as a discrete cosine transform (DCT) or a conceptually similar transform to the residual block to generate a video block with residual transform coefficient values. The transform processing unit 52 may perform other transforms conceptually similar to the DCT. Wavelet transform, integer transform, subband transform, or other types of transforms may also be used. The transform processing unit 52 applies a transform such as a discrete cosine transform (DCT) or a conceptually similar transform to the residual block to generate a video block with residual transform coefficient values. The transform processing unit 52 may perform other transforms conceptually similar to the DCT. Wavelet transform, integer transform, subband transform, or other types of transforms may also be used.
[0055] The conversion processing unit 52 applies a conversion to the residual block to generate a block of residual conversion coefficients. The conversion may convert the residual information from the pixel value domain to a conversion domain such as the frequency domain. The conversion processing unit 52 may send the obtained conversion coefficients to the quantization unit 54. The quantization unit 54 quantizes the conversion coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting the quantization parameter. In some examples, the quantization unit 54 may then perform a scan of the matrix including the quantized conversion coefficients. Alternatively, the entropy encoding unit 56 may perform the scan. After quantization, the entropy coding unit 56 entropy-codes the quantized conversion coefficients. For example, the entropy coding unit 56 may perform context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding technique. In the case of context-based entropy coding, the context may be based on adjacent blocks. After the entropy coding by the entropy coding unit 56, the encoded bit stream may be sent to another device (e.g., the video decoder 30), or may be archived for later transmission or retrieval.
[0056]
[0057] The inverse quantization unit 58 and the inverse transform unit 60 respectively apply inverse quantization and inverse transform to reconstruct, for example, the residual block in the pixel region for later use as a reference block. The motion compensation unit 44 may calculate the reference block by adding the residual block to one of the prediction blocks in the frame of the reference frame memory 64. The motion compensation unit 44 may also apply one or more interpolation filters to the reconstructed residual block to calculate sub-pixel values for use in motion estimation. The adder 62 adds the reconstructed residual block to the motion-compensated prediction block generated by the motion compensation unit 44 to generate a reconstructed video block for storage in the reference frame memory 64. The reconstructed video block may be used by the motion estimation unit 42 and the motion compensation unit 44 as a reference block for inter-coding blocks in subsequent video frames. For example, for later use as a reference block, the inverse quantization unit 58 and the inverse transform unit 60 respectively apply inverse quantization and inverse transform to reconstruct the residual block in the pixel region. The motion compensation unit 44 may calculate the reference block by adding the residual block to one of the prediction blocks in the frame of the reference frame memory 64. For example, for later use as a reference block, the inverse quantization unit 58 and the inverse transform unit 60 respectively apply inverse quantization and inverse transform to reconstruct the residual block in the pixel region. The motion compensation unit 44 may also apply one or more interpolation filters to the reconstructed residual block to calculate sub-pixel values for use in motion estimation. For example, for later use as a reference block, the inverse quantization unit 58 and the inverse transform unit 60 respectively apply inverse quantization and inverse transform to reconstruct the residual block in the pixel region. The adder 62 adds the reconstructed residual block to the motion-compensated prediction block generated by the motion compensation unit 44 to generate a reconstructed video block for storage in the reference frame memory 64. For example, for later use as a reference block, the inverse quantization unit 58 and the inverse transform unit 60 respectively apply inverse quantization and inverse transform to reconstruct the residual block in the pixel region. The reconstructed video block may be used by the motion estimation unit 42 and the motion compensation unit 44 as a reference block for inter-coding blocks in subsequent video frames. For example, for later use as a reference block, the inverse quantization unit 58 and the inverse transform unit 60 respectively apply inverse quantization and inverse transform to reconstruct the residual block in the pixel region. The reconstructed video block may be used by the motion estimation unit 42 and the motion compensation unit 44 as a reference block for inter-coding blocks in subsequent video frames.
[0058] Figure 3 is a block diagram showing an example of a video decoder 30 that may implement video coding techniques. In the example of Figure 3, the video decoder 30 includes an entropy decoding unit 70, a motion compensation unit 72, an intra prediction unit 74, an inverse quantization unit 76, an inverse transform unit 78, a reference frame memory 82, and an adder 80. In some examples, the video decoder 30 may perform a decoding path that is generally opposite to the encoding path described with respect to the video encoder 20 (Figure 2). The motion compensation unit 72 may generate prediction data based on the motion vector received from the entropy decoding unit 70, while the intra prediction unit 74 may perform intra prediction. Figure 3 is a block diagram showing an example of a video decoder 30 that may implement video coding techniques. In the example of Figure 3, the video decoder 30 includes an entropy decoding unit 70, a motion compensation unit 72, an intra prediction unit 74, an inverse quantization unit 76, an inverse transform unit 78, a reference frame memory 82, and an adder 80. In some examples, the video decoder 30 may perform a decoding path that is generally opposite to the encoding path described with respect to the video encoder 20 (Figure 2). In some examples, the video decoder 30 may perform a decoding path that is generally opposite to the encoding path described with respect to the video encoder 20 (Figure 2). The motion compensation unit 72 may generate prediction data based on the motion vector received from the entropy decoding unit 70, while the intra prediction unit 74 may perform intra prediction. The motion compensation unit 72 may generate prediction data based on the motion vector received from the entropy decoding unit 70, while the intra prediction unit 74 may perform intra prediction. Prediction data may be generated based on the intra prediction mode indicator received from the copy decoding unit 70.
[0059] During the decoding process, the video decoder 30 receives an encoded video bitstream representing the video blocks of the encoded video slice and the associated syntax elements from the video encoder 20. The entropy decoding unit 70 of the video decoder 30 entropy decodes the bitstream to generate quantization coefficients, motion vectors or intra prediction mode indicators, and other syntax elements. The entropy decoding unit 70 transfers the motion vectors and other syntax elements to the motion compensation unit 72. The video decoder 30 may receive syntax elements at the video slice level and / or the video block level.
[0060] When the video slice is coded as an intra-coded (I) slice, the intra prediction unit 74 may generate prediction data for the video blocks of the current video slice based on the signaled intra prediction mode and the data from the previously decoded blocks of the current frame or current picture. When the video frame is coded as an inter-coded (e.g., B, P, or GPB) slice, the motion compensation unit 72 generates a prediction block for the video blocks of the current video slice based on the motion vectors and other syntax elements received from the entropy decoding unit 70. The prediction block is one of the reference picture lists. It may be generated from one of the reference pictures in. The video decoder 30 uses a default configuration technique based on the reference picture stored in the reference frame memory 82 to construct a reference frame list, that is, list 0 and list 1.
[0061] The motion compensation unit 72 parses the motion vector and other syntax elements to determine the prediction information for the video block of the current video slice, and uses the prediction information to generate a prediction block for the current video block being decoded. For example, the motion compensation unit 72 uses some of the received syntax elements to determine the prediction mode (e.g., intra prediction or inter prediction) used to code the video block of the video slice, the inter prediction slice type (e.g., B slice , P slice, or GPB slice), the configuration information for one or more of the reference picture lists for the slice, the motion vector for each inter-coded video block of the slice, the inter prediction status for each inter-coded video block of the slice, and other information for decoding the video block in the current video slice.
[0062] The motion compensation unit 72 may also perform interpolation based on an interpolation filter. The motion compensation unit 72 may use an interpolation filter as used by the video encoder 20 during the encoding of the video block to calculate the interpolation value for the sub-integer pixels of the reference block. In this case, the motion compensation unit 72 may determine the interpolation filter used by the video encoder 20 from the received syntax elements and use the interpolation filter to predict. A measurement block may be generated.
[0063] Data for the texture image corresponding to the depth map may be stored in the reference frame memory 82. The motion compensation unit 72 may also be configured to interpolate the depth blocks of the depth map. In one embodiment, the video decoder 30 includes a user interface (UI) 84. The user interface 84 is configured to receive input from a user (e.g., a network administrator) of the video decoder 30. Through the user interface 84, the user can manage or change the settings on the video decoder 30. For example, the user can input or otherwise provide values for parameters (e.g., flags) to control the configuration and / or operation of the video decoder 30 according to the user's preferences. The user interface 84 may be a graphical user interface (GUI) that enables the user to interact with the video decoder 30 through, for example, graphical icons, drop-down menus, check boxes, etc. In some cases, the user interface 84 may receive information from the user via a keyboard, mouse, or other peripheral devices. In one embodiment, the user can access the user interface 84 via a smartphone, tablet device, personal computer, etc., located far away from the video decoder 30. The user interface 84 used herein may also be referred to as an external input or external means. It may be configured to perform interpolation.
[0064] In one embodiment, the video decoder 30 includes a user interface (UI) 84. The user interface 84 is configured to receive input from a user (e.g., a network administrator) of the video decoder 30. Through the user interface 84, the user can manage or change the settings on the video decoder 30. For example, the user can input or otherwise provide values for parameters (e.g., flags) to control the configuration and / or operation of the video decoder 30 according to the user's preferences. The user interface 84 may be a graphical user interface (GUI) that enables the user to interact with the video decoder 30 through, for example, graphical icons, drop-down menus, check boxes, etc. In some cases, the user interface 84 may receive information from the user via a keyboard, mouse, or other peripheral devices. In one embodiment, the user can access the user interface 84 via a smartphone, tablet device, personal computer, etc., located far away from the video decoder 30. The user interface 84 used herein may also be referred to as an external input or external means. The user interface 84 is configured to receive input from a user (e.g., a network administrator) of the video decoder 30. From the user (e.g., a network administrator) of the video decoder 30. Through the user interface 84, the user can manage or change the settings on the video decoder 30. For example, the user can input or otherwise provide values for parameters (e.g., flags) to control the configuration and / or operation of the video decoder 30 according to the user's preferences. To control the configuration and / or operation of the video decoder 30 according to the user's preferences, values for parameters (e.g., flags) can be input or provided in another way. The user interface 84 may be, for example, a graphical user interface (GUI) that enables the user to interact with the video decoder 30 through graphical icons, drop-down menus, check boxes, etc. It may be possible for the user to interact with the video decoder 30 through, for example, graphical icons, drop-down menus, check boxes, etc. It may be a graphical user interface (GUI). In some cases, The user interface 84 may receive information from the user via a keyboard, mouse, or other peripheral devices. In one embodiment, the user can access the user interface 84 via a smartphone, tablet device, personal computer, etc., located far away from the video decoder 30. Such as a smartphone, tablet device, personal computer, etc. The user can access the user interface 84 via a smartphone, tablet device, personal computer, etc., located far away from the video decoder 30. The user interface 84 used herein may also be referred to as an external input or external means. The user interface 84 used herein may also be referred to as an external input or external means.
[0065] With the above in mind, video compression techniques include spatial (intra-picture) prediction and / or or perform temporal (inter-picture) prediction to reduce redundancy inherent in video sequences In the case of block-based video coding, the video slice (i.e. A video picture (or a part of a video picture) is a tree block, a coding coding tree block (CTB), coding tree unit (CTU) ree unit), coding unit (CU), and / or coding node The picture may be divided into video blocks, which may also be referred to as intra-coding. Video blocks in an iterated (I) slice are divided into adjacent blocks in the same picture. The image is coded using spatial prediction with respect to the reference samples of the other picture. Video blocks in a segmented (P or B) slice are not related to adjacent blocks in the same picture. Spatial prediction with respect to reference samples in the same picture, or with respect to reference samples in other reference pictures. A picture may be called a frame, and a reference picture may be used for temporal prediction. may be referred to as a reference frame.
[0066] Spatial prediction or temporal prediction is used to obtain a prediction block for a block to be coded. The residual data is the difference between the original block to be coded and the predicted block. Represents pixel difference. Inter-coded blocks form the prediction block A motion vector pointing to a block of reference samples and a coded block The difference between the predicted block and the residual data is coded according to the residual data. The partitioned blocks are encoded according to the intra-coding mode and the residual data. For further compression, the residual data may be transformed from the pixel domain to the transform domain to result in residual transform coefficients, which may then be quantized. The quantized transform coefficients, which are initially arranged in a two-dimensional array, may be scanned to generate a one-dimensional vector of transform coefficients, and entropy coding may be applied to achieve further compression.
[0067] Image and video compression has experienced rapid growth and has led to various coding standards. Such video coding standards include ITU-T H.261, ISO / IEC MPEG-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H.264 or Advanced Video Coding (AVC) also known as ISO / IEC MPEG-4 Part 10, and High Efficiency Video Coding (HEVC) also known as ITU-T H.265 or MPEG-H Part 2. AVC includes extensions such as Scalable Video Coding (SVC), Multi-View Video Coding (MVC) and Multi-View Video Coding Plus Depth (MVC+D), and 3D AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multi-View HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). There is also a new video coding standard called Versatile Video Coding (VVC) being developed by the
[0068] Joint Video Experts Team (JVET) of ITU-T and ISO / IEC. The VVC standard has several has a working draft, and in particular, one working draft (WD) of VVC, namely, B. Bross, J. Chen, and S. Liu, "Versatile Video Coding (Draft 4)", JVET-M1001-v 5, the 13th JVET meeting, January 2019 (VVC draft 4) is referenced herein.
[0069] The description of the techniques disclosed herein is based on the video coding standard under development, namely, the Versatile Video Coding (VVC) by the Joint Video Expert Team (JVET) of ITU- T and ISO / IEC. However, the techniques are also applicable to other video codec specifications.
[0070] Figure 4 is a depiction 400 of the relationship between the IR AP picture 402 and the trailing picture 406 with respect to the leading picture 404 in decoding order 408 and presentation order 410. In one embodiment the IRAP picture 402 is called a clean random access (CRA) picture, or an instantaneous decoder refresh (IDR) picture with a random access decodable (RADL) picture. In HEV C, IDR pictures, CRA pictures, and broken link access (BLA) pictures are all considered IRAP pictures 402. In the case of VVC, during the 12th JVET meeting in October 2018, it was agreed to have both IDR pictures and CRA pictures as IRAP pictures.
[0071] As shown in FIG. 4, the leading pictures 404 (e.g., pictures 2 and 3) follow the IRAP picture 402 in the decoding order 408 but precede the IRAP picture 402 in the presentation order 410. The trailing picture 406 follows the IRAP picture 402 in both the decoding order 408 and the presentation order 410 . Two leading pictures 404 and one trailing picture 406 are shown in FIG. 4 , but in actual applications, more or fewer leading pictures 404 and / or trailing pictures 406 may be present in the decoding order 408 and the presentation order 410, which should be understood by those skilled in the art.
[0072] The leading pictures 404 in FIG. 4 are divided into two types, namely, random access skipped leading (RASL) and RADL. When decoding starts with the IRAP picture 402 (e.g., picture 1), the RADL picture (e.g., picture 3) can be correctly decoded, but the RASL picture (e.g., picture 2) cannot be correctly decoded. Therefore, the RASL picture is discarded. In view of the difference between the RADL picture and the RASL picture, the type of the leading picture related to the IRAP picture should be identified as either RADL or RASL for efficient and proper coding. In HEVC, when RASL pictures and RADL pictures exist, for the RASL pictures and RADL pictures related to the same IRAP picture, it is restricted that the RASL picture must precede the RAD L picture in the presentation order 410.
[0073] IRAP Picture 402 provides two important features / benefits: The presence of 02 indicates that the decoding process can start from that picture. As long as an IRAP picture 402 is present at that position, it does not necessarily have to be at the beginning of the bitstream. Random access, where the decoding process starts at that position in the bitstream. Second, the existence of the IRAP picture 402 enables the following feature: A coded picture, excluding RASL pictures, does not contain any reference to previous pictures. This refreshes the decoding process so that the image is coded without any additional overhead. Having the IRAP picture 402 present in the bit stream means that the IRAP picture 402 and The IRAP picture 402 propagates to pictures that follow it in the decoding order 408. Before that, any errors that may occur during the decoding of the coded picture are stopped. This will result in...
[0074] The IRAP picture 402 provides important functionality, but comes with a penalty in compression efficiency. The presence of P pictures 402 causes a spike in bit rate. This penalty to compression efficiency is This is due to two reasons. First, the IRAP picture 402 is an intra-predicted picture. When the other picture is an inter-predicted picture (e.g., the leading picture), In comparison with the previous picture 404 and the trailing picture 406, the picture itself is compared to depict Second, the presence of the IRAP picture 402 makes temporal prediction difficult. (Because the decoder refreshes the decoding process, this One of the actions of the decoding process for this is to remove the previous reference picture in the decoding picture buffer (DPB). ), since IRAP picture 402 has fewer reference pictures for inter-prediction coding, the decoding order 408 does not make the coding of pictures following IRAP picture 402 very efficient (i.e., it requires more bits to describe).
[0075] Among the picture types regarded as IRAP picture 402, IDR pictures in HEVC have different signaling and derivation compared to other picture types. Some of the differences are as follows.
[0076] For the signaling and derivation of the picture order count (POC) value of IDR pictures, the most significant bit (MSB) part of the POC is not derived from the previous key picture and is simply set to be equal to 0.
[0077] In the case of the signaling information required for reference picture management, the slice header of IDR pictures does not need to be signaled with information to assist in reference picture management . In the case of other picture types (i.e., CRA, trailing, temporal sub-layer access (TSA), etc.), information such as the reference picture set (RPS) described below, or other forms of similar information (e.g., reference picture list) is used in the reference picture marking process (i.e., the decoding picture buffer The status of reference pictures in the DPB, i.e., whether used or unused for reference, is required for the process of determining) However, in the case of IDR pictures, the presence of IDR indicates that the decoding process must simply mark all reference pictures in the DPB as unused for reference, so such information need not be signaled. In HEVC and VVC, the IRAP picture 402 and the leading picture 404 may each be included within a single network abstraction layer (NAL) unit. A set of NAL units may be referred to as an access unit. The IRAP picture 402 and the leading
[0078] picture 404 are given different NAL unit types so that they can be easily identified by system-level applications. For example, a video splicer needs to understand the coded picture type without having to understand too many details of the syntax elements in the coded bitstream, including, in particular, determining the IRAP picture 402 from non-IRAP pictures and determining the leading picture 404 from trailing pictures 406 to RASL pictures and RADL pictures. The trailing picture 406 is a picture related to the IRAP picture 402 and follows the IRAP picture 402 in the presentation order 410. A picture may follow a particular IRAP picture 4 02 in the decoding order 408 or may precede any other IRAP picture 402 in the decoding order 408. Including, in particular, determining the RASL picture and the RADL picture from the trailing picture 406 Identifying the IRAP picture 402 from non-IRAP pictures and identifying the leading picture 404 There is a need to understand the coded picture type without having to understand too many details of the syntax elements in the coded bitstream. The trailing picture 406 is a picture related to the IRAP picture 402 and follows the IRAP picture 402 in the presentation order 410. A picture may follow a particular IRAP picture 4 02 in the decoding order 408 or may precede any other IRAP picture 402 in the decoding order 408. The picture may follow a particular IRAP picture 4 02 in the decoding order 408 or may precede any other IRAP picture 402 in the decoding order 408. Therefore, giving their own NAL unit types to the IRAP picture 402 and the leading picture 404 is helpful for such applications.
[0079] In the case of HEVC, the NAL unit types for the IRAP picture include the following. BLA with leading pictures (BLA_W_LP): NAL unit of a broken link access (BLA) picture where one or more leading pictures may follow in decoding order. BLA with RADL (BLA_W_RADL): NAL unit of a BLA picture where one or more RADL pictures may follow in decoding order but no RA SL picture may follow. BLA without leading pictures (BLA_N_LP): NAL unit of a BLA picture where no leading picture follows in decoding order. IDR with RADL (IDR_W_RADL): NAL unit of an IDR picture where one or more RADL pictures may follow in decoding order but no RA SL picture may follow. IDR without leading pictures (IDR_N_LP): NAL unit of an IDR picture where no leading picture follows in decoding order. CRA: NAL unit of a clean random access (CRA) picture where a leading picture may follow (i.e., either a RASL picture or a RADL picture, or both). RADL: NAL unit of a RADL picture. RASL: NAL unit of a RASL picture.
[0080] In the case of VVC, the NAL unit types for the IRAP picture 402 and the leading picture 404 are as follows. IDR with RADL (IDR_W_RADL): An NAL unit of an IDR picture where one or more RADL pictures may follow in decoding order, but no SL picture may follow. An NAL unit of an IDR picture where no SL picture may follow. IDR without a leading picture (IDR_N_LP): An NAL unit of an IDR picture where no leading picture follows in decoding order. An NAL unit of an IDR picture where no leading picture follows. CRA: A clean random access (CRA) picture where a leading picture may follow. An NAL unit of the CRA picture (i.e., either a RASL picture or a RADL picture, or both). An NAL unit of the CRA picture (i.e., either a RASL picture or a RADL picture, or both). Both). RADL: An NAL unit of a RADL picture. RASL: An NAL unit of a RASL picture.
[0081] Sequential intra refresh / progressive decoding refresh is described below.
[0082] For low - latency applications, compared to non - IRAP pictures (i.e., P pictures / B pictures), due to the relatively large bit rate requirements of IRAP pictures, which are relatively large and thus cause a larger latency / delay, it is desirable to avoid coding a picture as an IRAP picture (e.g., IRAP picture 402). However, completely avoiding the use of IRAP may not be possible in all low - latency applications. For example, in the case of conversational applications such as multiparty video conferencing, it is necessary to provide regular points where new users can join the video conferencing. For low - latency applications, compared to non - IRAP pictures (i.e., P pictures / B pictures), due to the relatively large bit rate requirements of IRAP pictures, which are relatively large and thus cause a larger latency / delay, it is desirable to avoid coding a picture as an IRAP picture (e.g., IRAP picture 402). However, completely avoiding the use of IRAP may not be possible in all low - latency applications. For example, in the case of conversational applications such as multiparty video conferencing, it is necessary to provide regular points where new users can join the video conferencing. For low - latency applications, compared to non - IRAP pictures (i.e., P pictures / B pictures), due to the relatively large bit rate requirements of IRAP pictures, which are relatively large and thus cause a larger latency / delay, it is desirable to avoid coding a picture as an IRAP picture (e.g., IRAP picture 402). However, completely avoiding the use of IRAP may not be possible in all low - latency applications. For example, in the case of conversational applications such as multiparty video conferencing, it is necessary to provide regular points where new users can join the video conferencing. For low - latency applications, compared to non - IRAP pictures (i.e., P pictures / B pictures), due to the relatively large bit rate requirements of IRAP pictures, which are relatively large and thus cause a larger latency / delay, it is desirable to avoid coding a picture as an IRAP picture (e.g., IRAP picture 402). However, completely avoiding the use of IRAP may not be possible in all low - latency applications. For example, in the case of conversational applications such as multiparty video conferencing, it is necessary to provide regular points where new users can join the video conferencing. For low - latency applications, compared to non - IRAP pictures (i.e., P pictures / B pictures), due to the relatively large bit rate requirements of IRAP pictures, which are relatively large and thus cause a larger latency / delay, it is desirable to avoid coding a picture as an IRAP picture (e.g., IRAP picture 402). However, completely avoiding the use of IRAP may not be possible in all low - latency applications. For example, in the case of conversational applications such as multiparty video conferencing, it is necessary to provide regular points where new users can join the video conferencing. For low - latency applications, compared to non - IRAP pictures (i.e., P pictures / B pictures), due to the relatively large bit rate requirements of IRAP pictures, which are relatively large and thus cause a larger latency / delay, it is desirable to avoid coding a picture as an IRAP picture (e.g., IRAP picture 402). However, completely avoiding the use of IRAP may not be possible in all low - latency applications. For example, in the case of conversational applications such as multiparty video conferencing, it is necessary to provide regular points where new users can join the video conferencing. For low - latency applications, compared to non - IRAP pictures (i.e., P pictures / B pictures), due to the relatively large bit rate requirements of IRAP pictures, which are relatively large and thus cause a larger latency / delay, it is desirable to avoid coding a picture as an IRAP picture (e.g., IRAP picture 402). However, completely avoiding the use of IRAP may not be possible in all low - latency applications. For example, in the case of conversational applications such as multiparty video conferencing, it is necessary to provide regular points where new users can join the video conferencing.
[0083] To enable new users to join a multiparty video conferencing application, one possible approach is to provide access to the bitstream To enable new users to join a multiparty video conferencing application, one possible approach is to provide access to the bitstream Rather than using an IRAP picture, in order to avoid having a - slice, a progressive intra refresh (PIR) technique is used instead. PIR may also be referred to as progressive decoding refresh (GDR). The terms PIR and GDR may be used interchangeably in the present disclosure. Figure 5 shows a progressive decoding refresh (GDR) technique 500. As shown, the GDR technique 500 is shown using GDR pictures 502, one or more trailing pictures 504, and recovery point pictures 506 within a coded video sequence 508 of a bitstream. In one embodiment, the GDR pictures 502, trailing pictures 504, and recovery point pictures 506 may define a GDR period within the CVS 508. The CVS 508 is a series of pictures (or portions thereof) that start with a GDR picture 502 and include all pictures (or portions thereof) up to, but not including, the next GDR picture or up to the end of the bitstream.
[0084] The GDR period is a series of pictures that start with a GDR picture 502 and include all pictures up to and including the recovery point picture 506. As shown in Figure 5, the GDR technique 500 or principle functions over a series of pictures that start with a GDR picture 502 and end with a recovery point picture 506. The GDR picture 502 includes a refreshed / clean area 510 that contains blocks that are all coded using intra - prediction (i.e., intra - prediction blocks), and all are inter - predicted blocks. The trailing pictures 504 are coded using inter - prediction with reference to the GDR picture 502 and other previous pictures within the CVS 508. The recovery point picture 506 is a picture that marks a recovery point within the CVS 508 and is used to assist in the recovery of the video sequence in case of errors or losses. In one embodiment, the GDR picture 502 may be coded using only intra - prediction, and the trailing pictures 504 may be coded using inter - prediction with reference to the GDR picture 502 and other previous pictures within the CVS 508. The recovery point picture 506 may be coded using a combination of intra - prediction and inter - prediction, or only intra - prediction. The GDR period may be adjusted based on various factors such as the complexity of the video sequence, the available bandwidth, and the error - resilience requirements.
[0085] As shown in Figure 5, the GDR technique 500 or principle functions over a series of pictures that start with a GDR picture 502 and end with a recovery point picture 506. The GDR picture 502 includes a refreshed / clean area 510 that contains blocks that are all coded using intra - prediction (i.e., intra - prediction blocks), and all are inter - predicted blocks. The trailing pictures 504 are coded using inter - prediction with reference to the GDR picture 502 and other previous pictures within the CVS 508. contains a block coded using it (i.e., an inter-prediction block), a non-refreshed / dirty area 512.
[0086] The trailing picture 504 that is directly adjacent to the GDR picture 502 is coded using intra-prediction for a first portion 510A and coded using inter-prediction for a second portion 510B, and contains a refreshed / clean area 510. The second portion 510B is coded by referring to a refreshed / clean area 510 of a previous picture within the GDR period of the CVS 508, for example. As shown, the refreshed / clean area 510 of the trailing picture 504 expands as the coding process moves or progresses in a consistent direction (e.g., from left to right), and correspondingly, shrinks the non-refreshed / dirty area 512. Finally, a recovery point picture 506 that contains only the refreshed / clean area 510 is obtained from the coding process. In particular, as will be further explained below, the second portion 510B of the refreshed / clean area 510 that is coded as an inter-prediction block only needs to refer to the refreshed area / clean area
[0087] In HEVC, the GDR technique 500 of FIG. 5 is a recovery point supplemental enhancement information (SEI: Supplemental Enhancement Information) are non-normatively supported using Gee. These two SEI messages do not specify how the GDR is to be executed. Rather, the two SEI messages simply provide a mechanism for indicating the first picture and the last picture within the GDR period (i.e., provided by the recovery point SEI message) as well as the regions being refreshed (i.e., provided by the region refresh information SEI message).
[0088] In fact, GDR technique 500 is implemented by using two techniques together. Those two techniques are constraint intra prediction (CIP) and encoder constraints on motion vectors. CIP can be used for GDR purposes to code regions that are coded only as intra prediction blocks (e.g., the first portion 510A of the refreshed / clean region 510), since CIP allows regions that do not use samples from non-refreshed regions (e.g., non-refreshed / dirty regions 512) to be used for reference. However, the constraints on intra blocks must be applied not only to intra blocks within the refreshed region but also to all intra blocks within the picture, and the use of CIP causes significant coding performance degradation. The encoder constraints on motion vectors limit the encoder from using any samples in the reference picture located outside the refreshed region. Such constraints cause sub-optimal motion search.
[0089] Figure 6 shows undesirable motion when using encoder constraints to support GDR It is a schematic diagram showing search 600. As shown in the figure, motion search 600 shows the current picture 602 and the reference picture 604. The current picture 602 and the reference picture 604 each have a refreshed area 606 coded using intra prediction, a refreshed area 608 coded using inter prediction, and an area 608 that has not been refreshed. The refreshed area 604, the refreshed area 606, and the area 608 that has not been refreshed correspond to the first part 510A of the refreshed / clean area 510, the second part 510B of the refreshed / clean area 510, and the non-refreshed / dirty area 512 in Figure 5, respectively. During the motion search process, the encoder is restricted or prevented from selecting any motion vector 610 that results in some of the samples of the reference block 612 located outside the refreshed area 606. This is done even when the reference block 612 gives the best rate-distortion cost criterion when predicting the current block 614 in the current picture 602. Therefore, Figure 6 shows the reason for the non-optimality in motion search 600 when using encoder constraints to support GDR.
[0090]
[0091] JVT submissions JVET-K0212 and JVET-L0160 describe the implementation forms of GDR based on the use of CIP and encoder constraint techniques. The implementation forms can be summarized as follows. That is, the columns For each coding unit, the intra prediction mode is forced, and to ensure the reconstruction of the intra CU, constrained intra prediction is enabled. The motion vector takes into account an additional margin to avoid errors (e.g., 6 pixels) spread due to filtering, and is constrained to point within the refreshed area without removing past reference pictures when re-looping the intra column. Among the problems associated with the existing GDR design, some are described. JVET contribution JVET-M0529 proposed a method to canonically indicate that a picture is the first and last picture within a GDR period. The proposed idea functions as follows. A new NAL unit with NAL unit type recovery point indication is defined as a non-video coding layer (VCL) NAL unit. The payload of the NAL unit contains syntax elements to specify information that can be used to derive the POC value of the last picture within the GDR period. An access unit containing a non-VCL NAL unit with type recovery point indication is called a recovery point begin (RBP) access unit (AU), and the pictures within the RBP access unit are called RBP pictures. The decoding process can start from the RBP AU. When decoding starts from the RBP AU, all pictures within the GDR period except the last picture are not output.
[0092]
[0093]
[0094]
[0095] Existing designs / methods for supporting GDR have at least the following problems.
[0096] The method for formally defining GDR in JVET-M0529 has the following problems. Proposal The proposed method does not explain how GDR is executed. Instead, the proposed method only provides some signaling for indicating the first picture and the last picture during the GDR period. To indicate the first picture and the last picture during the GDR period, new non-VCL NAL units are required. Since the information contained in the recovery point indication (RPI) NAL unit can simply be included in the tile group header of the first picture during the GDR period, this is redundant. Moreover, the proposed method cannot represent which regions in the pictures during the GDR period are refreshed regions and which are non-refreshed regions.
[0097] The GDR techniques described in JVET-K0212 and JVET-L0160 have the following problems. First, the use of CIP. To prevent any samples from non-refreshed regions from being used for spatial reference, it is necessary to code the refreshed regions using intra prediction with some constraints. When CIP is used, the coding is picture-based, which means that all intra blocks in the picture must be coded as CIP intra blocks. Therefore, this causes performance degradation. Furthermore, the reference blocks related to motion vectors The sample may not be entirely within the refreshed region in the reference picture When, the use of encoder constraints to limit motion search prevents the encoder from selecting the best motion vector. Also, the refreshed region coded using only intra prediction is not CTU-sized. Instead, the refreshed region can be made smaller than the CTU size, down to the minimum CU size. This may require block-level display, making the implementation unnecessarily complex
[0098] Techniques for supporting Gradual Decoding Refresh (GDR) in video coding are disclosed herein. The disclosed techniques enable sequential intra refresh to enable random access without the need to use intra random access point (IRAP) pictures. When the value for a first flag is provided by an external input, the first flag is set equal to the value provided by the external input and a second flag is set equal to the first flag to prevent the output of Gradual Decoding Refresh (GDR) pictures and any trailing pictures between the GDR pictures and recovery point pictures in output order. The external input may be an input received from a user (e.g., a network administrator) via a graphical user interface (GUI) of the video decoder 30. Setting the first and second flags in this way potentially prevents dirty data from being output to the display. That is, the values of the first and second flags determine whether potentially dirty data from the GDR picture is output Control whether the video decoder waits for full synchronization before starting to display data By having the ability to limit the output of dirty data, a coder / decoder (also called a "codec") in video coding is improved compared to the current codec. In practice, an improved video coding process provides a better user experience to the user when the video is sent, received, and / or viewed.
[0099] To solve one or more of the problems described above, the present disclosure discloses the following aspects. Each aspect can be applied individually, and some of them can be applied in combination.
[0100] 1) A VCL NAL unit having type GDR_NUT is defined.
[0101] a. A picture having NAL unit type GDR_NUT is called a GDR picture, i.e., the first picture in the GDR period.
[0102] b. The GDR picture has a temporalID equal to 0.
[0103] c. An access unit containing a GDR picture is called a GDR access unit. As described above, an access unit is a set of NAL units. Each NAL unit may contain a single picture.
[0104] 2) The coded video sequence (CVS) may start with a GDR access unit.
[0105] 3) When one of the following is true, the GDR access unit is the first access unit in the CVS.
[0106] a. The GDR access unit is the first access unit in the bitstream.
[0107] b. The GDR access unit comes immediately after the end-of-sequence (EOS) access unit.
[0108] c. The GDR access unit comes immediately after the end-of-bitstream (EOB) access unit.
[0109] d. The decoder flag, so-called NoIncorrectPicOutputFlag, is associated with the GDR picture and the value of the flag is set to 1 (i.e., true) by an entity outside the decoder.
[0110] 4) When the GDR picture is the first access unit in the CVS, the following applies.
[0111] a. All reference pictures in the DPB are marked as "unused for reference".
[0112] b. The POC MSB of the picture is set to be equal to 0.
[0113] c. Excluding the GDR picture and the last picture in the GDR period, all pictures following the GDR picture in output order up to the last picture in the GDR period are not output (i.e., marked as "unnecessary for output").
[0114] 5) A flag for specifying whether the GDR is enabled is signaled in the sequence level parameter set (e.g., in the SPS).
[0115] a. The flag may be designated as gdr_enabled_flag.
[0116] b. When the flag is equal to 1, the GDR picture may be present in the CVS. Otherwise, when the flag is equal to 0, the GDR is not enabled such that the GDR picture is not present in the CVS.
[0117] 6) Information that can be used to derive the POC value of the last picture in the GDR period is signaled in the tile group header of the GDR picture.
[0118] a. The information is signaled as the delta POC between the last picture in the GDR period and the GDR picture. The information can be signaled using the syntax element designated as recovery_point_cnt.
[0119] b. The presence of the syntax element recovery_point_cnt in the tile group header may be conditioned on the value of the gdr_enabled flag and the NAL unit type of the picture, i.e., the flag is present only when the gdr_enabled_flag is equal to 1 and the nal_unit_type of the NAL unit containing the tile group is GDR_NUT.
[0120] 7) A flag for specifying whether a tile group is part of a refreshed area is signaled in the tile group header.
[0121] a. The flag may be designated as the refreshed_region_flag.
[0122] b. The existence of the flag may be conditioned on the value of the gdr_enabled_flag and whether the picture containing the tile group is within the GDR period. Thus, the flag exists only when all of the following are true.
[0123] i. The value of the gdr_enabled_flag is equal to 1.
[0124] ii. The POC of the current picture is greater than or equal to the POC value of the last GDR picture (when the current picture is a GDR picture, the last GDR picture is the current picture), and less than the POC of the last picture within the GDR period.
[0125] c. When the flag does not exist in the tile group header, the value of the flag is presumed to be equal to 1.
[0126] 8) All tile groups for which the refreshed_region_flag is equal to 1 cover the contiguous region. Similarly, all tile groups for which the refreshed_region_flag is equal to 0 also cover the contiguous region.
[0127] 9) Tile groups having the refreshed_region_flag can be of type I (i.e., intra tile groups) or B or P (i.e., inter tile groups).
[0128] 10) Each picture starting from the GDR picture up to the last picture within the GDR period includes at least one tile group for which the refreshed_region_flag is equal to 1.
[0129] 11) The GDR picture includes at least one tile group for which the refreshed_region_flag is equal to 1 and the tile_group_type is equal to I (i.e., the intra-tile group).
[0130] 12) When the gdr_enabled_flag is equal to 1, it is allowed for the information of the rectangular tile group, i.e., the number of tile groups and their addresses, to be signaled either in the picture parameter set (PPS: picture parameter set) or in the tile group header. To do this, a flag is signaled in the PPS to specify whether the rectangular tile group information exists in the PPS. This flag may be called rect_tile_group_info_in_pps_flag. This flag may be constrained to be equal to 1 when the gdr_enabled_flag is equal to 1.
[0131] a. In an alternative form, instead of signaling whether the rectangular tile group information exists in the PPS, a more general flag may be signaled in the PPS to specify whether the tile group information (i.e., any type of tile group such as a rectangular tile group, a raster scan tile group, etc.) exists in the PPS.
[0132] 13) When tile group information does not exist in the PPS, it may be further restricted that there is no signaling of explicit tile group identifier ( ID) information. The explicit tile group ID information includes signaled_tile_group_id_flag, signaled_tile_group_id_length_min us1, and tile_group_id[ i ].
[0133] 14) A flag is signaled to specify whether loop filter processing operations that cross the boundary between the refreshed and non - refreshed regions in a picture are allowed.
[0134] a. This flag may be signaled in the PPS and may be called loop_filter_across_refreshed _region_enabled_flag.
[0135] b. The presence of loop_filter_across_refreshed_region_enabled_flag may be conditioned on the value of loop_filter_acros s_tile_enabled_flag. When loop_filter_across_tile_enabled_fla g is equal to 0, loop_filter_across_refreshed_region_enabled_flag may not exist and its value is presumed to be equal to 0.
[0136] c. In an alternative form, the flag may be signaled in the tile group header and its presence may be conditioned on the value of refreshed_region_flag, i.e., refres The flag exists only when the value of shed_region_flag is equal to 1.
[0137] 15) When it is indicated that the tile group is a refreshed region and it is shown that a loop filter crossing the refreshed region is not allowed, the following applies. When it is shown that a loop filter crossing the refreshed region is not allowed while it is indicated that the tile group is a refreshed region, the following applies. Applies.
[0138] a. When adjacent tile groups sharing an edge are non - refreshed tile groups, de - blocking of the edge at the tile group boundary is not performed. When adjacent tile groups sharing an edge are non - refreshed tile groups, de - blocking of the edge at the tile group boundary is not performed.
[0139] b. The sample adaptive offset (SAO) process for blocks at the tile group boundary does not use any samples from outside the boundary of the refreshed region. The sample adaptive offset (SAO) process for blocks at the tile group boundary does not use any samples from outside the boundary of the refreshed region. Does not use any samples from outside the boundary of the refreshed region.
[0140] c. The adaptive loop filtering (ALF) process for blocks at the tile group boundary does not use any samples from outside the boundary of the refreshed region. The adaptive loop filtering (ALF) process for blocks at the tile group boundary does not use any samples from outside the boundary of the refreshed region. Does not use any samples from outside the boundary of the refreshed region.
[0141] 16) When gdr_enabled_flag is equal to 1, each picture is associated with variables for determining the boundaries of the refreshed regions within the picture. These variables may be called as follows. When gdr_enabled_flag is equal to 1, each picture is associated with variables for determining the boundaries of the refreshed regions within the picture. These variables may be called as follows. May be called as follows.
[0142] a. PicRefreshedLeftBoundaryPos for the left boundary position of the refreshed region within the picture. PicRefreshedLeftBoundaryPos for the left boundary position of the refreshed region within the picture.
[0143] b. PicRefreshedRig htBoundaryPos with respect to the right boundary position of the refreshed region in the picture. htBoundaryPos.
[0144] c. PicRefreshedTop BoundaryPos with respect to the upper boundary position of the refreshed region in the picture. BoundaryPos.
[0145] d. PicRefreshedBot BoundaryPos with respect to the lower boundary position of the refreshed region in the picture. BoundaryPos.
[0146] 17) The boundaries of the refreshed region in the picture may be derived. The boundaries of the refreshed region of the picture are updated by the decoder after the tile group header is parsed, and the value of the refreshed_region_flag of the tile group is equal to 1. The boundaries of the refreshed region of the picture are updated by the decoder after the tile group header is parsed, and the value of the refreshed_region_flag of the tile group is equal to 1.
[0147] 18) In an alternative form of solution 17), the boundaries of the refreshed region in the picture are explicitly signaled within each tile group of the picture. The boundaries of the refreshed region in the picture are explicitly signaled within each tile group of the picture.
[0148] a. A flag may be signaled to indicate whether the picture to which the tile group belongs contains an unrefreshed region. When it is specified that the picture does not contain an unrefreshed region, the refreshed boundary information is not signaled and cannot simply be assumed to be equal to the picture boundary. A flag may be signaled to indicate whether the picture to which the tile group belongs contains an unrefreshed region. When it is specified that the picture does not contain an unrefreshed region, the refreshed boundary information is not signaled and cannot simply be assumed to be equal to the picture boundary. When it is specified that the picture does not contain an unrefreshed region, the refreshed boundary information is not signaled and cannot simply be assumed to be equal to the picture boundary.
[0149] 19) For the current picture, the boundaries of the refreshed region are used in the in-loop filter process as follows. For the current picture, the boundaries of the refreshed region are used in the in-loop filter process as follows.
[0150] a. In the case of the deblocking process, determine whether an edge needs to be deblocked by determining the edge of the refreshed region. To determine whether an edge needs to be deblocked in the deblocking process, the edge of the refreshed region is determined.
[0151] b. In the case of the SAO process, if a loop filter that crosses the refresh region is not allowed, determine the boundary of the refreshed region so that a clipping process can be applied to avoid using samples from the non-refreshed region. If a loop filter that crosses the refresh region is not allowed in the SAO process, to avoid using samples from the non-refreshed region, the boundary of the refreshed region is determined so that a clipping process can be applied. To determine the boundary of the refreshed region so that a clipping process can be applied to avoid using samples from the non-refreshed region when a loop filter that crosses the refresh region is not allowed in the SAO process.
[0152] c. In the case of the ALF process, when a loop filter that crosses the refreshed region is not allowed, determine the boundary of the refresh region so that a clipping process can be applied to avoid using samples from the non-refreshed region. When a loop filter that crosses the refreshed region is not allowed in the ALF process, to avoid using samples from the non-refreshed region, the boundary of the refresh region is determined so that a clipping process can be applied. To determine the boundary of the refresh region so that a clipping process can be applied to avoid using samples from the non-refreshed region when a loop filter that crosses the refreshed region is not allowed in the ALF process.
[0153] 20) For the motion compensation process, information about the boundary of the refreshed region, particularly the boundary of the refreshed region in the reference picture, is used as follows. That is, when the current block in the current picture is in a tile group where the refreshed_region_flag is equal to 1 and the reference block is in a reference picture that includes a non-refreshed region, the following applies. That is, when the current block in the current picture is in a tile group where the refreshed_region_flag is equal to 1 and the reference block is in a reference picture that includes a non-refreshed region, the following applies. when the current block in the current picture is in a tile group where the refreshed_region_flag is equal to 1 and the reference block is in a reference picture that includes a non-refreshed region, the following applies.
[0154] a. The motion vector from the current block to its reference picture is clipped by the boundary of the refreshed region in that reference picture. For the fractional interpolation filter for samples in that reference picture, such motion vectors are clipped by the boundary of the refreshed region in that reference picture.
[0155] b. For the fractional interpolation filter for samples in that reference picture, such motion vectors are clipped by the boundary of the refreshed region in that reference picture. For the fractional interpolation filter for samples in that reference picture, such motion vectors are clipped by the boundary of the refreshed region in that reference picture. It will be done.
[0156] A detailed description of the embodiments of the present disclosure is provided. The description is related to the base text and is based on the The text is JVET contribution JVET-M1001-v5. That is, only the differences are listed below. Any text in the base text that is not stated applies as is. Text that is modified compared to the previous text is in italics.
[0157] A definition is given.
[0158] 3.1 Clean Random Access (CRA) pictures: nal_unit_type equal to CRA_NUT for each VCL An IRAP picture contained in a NAL unit.
[0159] NOTE - A CRA picture cannot use any other picture than itself for inter prediction in its decoding process. It does not refer to any picture in the bitstream and is the first picture in decoding order. A CRA picture may appear before or after the associated NoIncorrectPicOutputF equal to 1. When a CRA picture has a lag, the RASL picture is a picture that does not exist in the bitstream. Since the associated RASL picture may contain a reference to a The character is not output by the decoder.
[0160] 3.2 Coded Video Sequence (CVS): NoIncorrectPicOutputFlag equals 1 An IRAP access unit with no IncorrectPicOutputFlag or a GDR access unit with NoIncorrectPicOutputFlag equal to 1 All subsequent (but not including) any subsequent access units up to the next access unit that is an IRAP access unit including the access units where NoIncorrectPicOutputFlag is equal to 1 or GDR access units where NoIncorrectPicOutputFlag is equal to 1, and all subsequent 0 or more access units that are not IRAP access units where NoIncorrectPicOutputFlag is equal to 1 or GDR access units where NoIncorre ctPicOutputFlag is equal to 1, provided in decoding order, in a sequence of access units.
[0161] Note 1 - An IRAP access unit may be an IDR access unit or a CRA access unit. For each IDR access unit, the value of NoIncorrectPicOutputFlag is equal to 1 and each CRA access unit that is the first access unit in the bitstream in decoding order either follows the end-of-sequence NAL unit in decoding order or is the first access unit with a HandleCraAsCvsStartFlag equal to 1.
[0162] Note 2 - For each GDR access unit that is the first access unit in the bitstream in decoding order, the fact that the value of NoIncorrectPicOutputFlag is equal to 1 means that it either follows the end-of-sequence NAL unit in decoding order or is the first access unit with a HandleGdrAsCvsStartFla g equal to 1.
[0163] 3.3 Progressive Decoding Refresh (GDR) Access Unit: When the coded picture is G An access unit that is a DR picture.
[0164] 3.4 Gradual Decoding Refresh (GDR) picture: A picture that each VCL N AL unit has with a nal_unit_type equal to GDR_NUT.
[0165] 3.5 Random Access Skipped Reading (RASL) picture: A coded picture that each VCL NAL unit has with a nal_ unit_type equal to RASL_NUT.
[0166] Note - All RASL pictures are reading pictures of the related CRA picture. When the related CRA picture has a NoIncorrectPicOutputFlag equal to 1, the RASL picture may contain a reference to a picture that does not exist in the bitstream, so the RASL picture is not output and may not be correctly decodable. RASL pictures are not used as reference pictures for the decoding process of non - RASL pictures. When present, all RASL pictures precede, in decoding order, all trailing pictures of the same related CRA picture. Since the RASL picture may contain a reference to a picture that does not exist in the bitstream, so the RASL picture is not output and may not be correctly decodable. RASL pictures are not used as reference pictures for the decoding process of non - RASL pictures. When present, all RASL pictures precede, in decoding order, all trailing pictures of the same related CRA picture. Since the RASL picture may contain a reference to a picture that does not exist in the bitstream, so the RASL picture is not output and may not be correctly decodable. RASL pictures are not used as reference pictures for the decoding process of non - RASL pictures. When present, all RASL pictures precede, in decoding order, all trailing pictures of the same related CRA picture. Since the RASL picture may contain a reference to a picture that does not exist in the bitstream, so the RASL picture is not output and may not be correctly decodable. RASL pictures are not used as reference pictures for the decoding process of non - RASL pictures. When present, all RASL pictures precede, in decoding order, all trailing pictures of the same related CRA picture. Since the RASL picture may contain a reference to a picture that does not exist in the bitstream, so the RASL picture is not output and may not be correctly decodable. RASL pictures are not used as reference pictures for the decoding process of non - RASL pictures. When present, all RASL pictures precede, in decoding order, all trailing pictures of the same related CRA picture. precede, in decoding order, all trailing pictures of the same related CRA picture.
[0167] The syntax and semantics of the Sequence Parameter Set Low - byte Sequence Payload (RBSP: raw byte sequence payload). [Table 1]
[0168] A gdr_enabled_flag equal to 1 indicates that there are GDR pictures in the coded video sequence. Specify that there may be a "ャ". A gdr_enabled_flag equal to 0 specifies that there is no GDR picture in the coded video sequence.
[0169] The syntax and semantics of the picture parameter set RBSP.
Table 2
[0170] A rect_tile_group_info_in_pps_flag equal to 1 specifies that the rectangular tile group information is signaled in the PPS. A rect_tile_group_info_in_pps_flag equal to 0 specifies that the rectangular tile group information is not signaled in the PPS.
[0171] When the value of gdr_enabled_flag in the active SPS is equal to 0, it is a bitstream conformity requirement that the value of rect_tile_group_info _in_pps_flag must be equal to 0.
[0172] A loop_filter_across_refreshed_region_enabled_flag equal to 1 specifies that loop filter processing operations may be performed across the boundaries of tile groups where refreshed_region_flag is equal to 1 in the pictures that reference the PPS. A loop_filter_across_ refreshed_region_enabled_flag equal to 0 specifies that loop filter processing operations are not performed across the boundaries of tile groups where refreshed_region_f lag is equal to 1 in the pictures that reference the PPS. shall not be performed. Specify this. The in-loop filter processing operations include deblocking filter, sample adaptive offset filter, and adaptive loop filter operations. When not present, the value of loop_filter_across_refreshed_region_enabled_flag is assumed to be equal to 0. When not present, the value of loop_filter_across_refreshed_region_enabled_flag is assumed to be equal to 0. When not present, the value of loop_filter_across_refreshed_region_enabled_flag is assumed to be equal to 0.
[0173] A signalled_tile_group_id_flag equal to 1 specifies that a tile group ID for each tile group is signalled. A signalled_tile_group_index_flag equal to 0 specifies that a tile group ID is not signalled. When not present, the value of signalled_tile_group_index_flag is assumed to be equal to 0. A signalled_tile_group_id_flag equal to 1 specifies that a tile group ID for each tile group is signalled. A signalled_tile_group_index_flag equal to 0 specifies that a tile group ID is not signalled. When not present, the value of signalled_tile_group_index_flag is assumed to be equal to 0. A signalled_tile_group_id_flag equal to 1 specifies that a tile group ID for each tile group is signalled. A signalled_tile_group_index_flag equal to 0 specifies that a tile group ID is not signalled. When not present, the value of signalled_tile_group_index_flag is assumed to be equal to 0. A signalled_tile_group_id_flag equal to 1 specifies that a tile group ID for each tile group is signalled. A signalled_tile_group_index_flag equal to 0 specifies that a tile group ID is not signalled. When not present, the value of signalled_tile_group_index_flag is assumed to be equal to 0.
[0174] signalled_tile_group_id_length_minus1 + 1, when present, specifies the number of bits used to represent the syntax element tile_group_id[ i ] and the syntax element tile_group_address in the tile group header. The value of signalled_tile_group_index_length_minus1 shall be in the range of 0 to 15, inclusive. When not present, the value of signalled_tile_group_index_length_minus1 is assumed as follows. signalled_tile_group_id_length_minus1 + 1, when present, specifies the number of bits used to represent the syntax element tile_group_id[ i ] and the syntax element tile_group_address in the tile group header. The value of signalled_tile_group_index_length_minus1 shall be in the range of 0 to 15, inclusive. When not present, the value of signalled_tile_group_index_length_minus1 is assumed as follows. signalled_tile_group_id_length_minus1 + 1, when present, specifies the number of bits used to represent the syntax element tile_group_id[ i ] and the syntax element tile_group_address in the tile group header. The value of signalled_tile_group_index_length_minus1 shall be in the range of 0 to 15, inclusive. When not present, the value of signalled_tile_group_index_length_minus1 is assumed as follows. signalled_tile_group_id_length_minus1 + 1, when present, specifies the number of bits used to represent the syntax element tile_group_id[ i ] and the syntax element tile_group_address in the tile group header. The value of signalled_tile_group_index_length_minus1 shall be in the range of 0 to 15, inclusive. When not present, the value of signalled_tile_group_index_length_minus1 is assumed as follows. signalled_tile_group_id_length_minus1 + 1, when present, specifies the number of bits used to represent the syntax element tile_group_id[ i ] and the syntax element tile_group_address in the tile group header. The value of signalled_tile_group_index_length_minus1 shall be in the range of 0 to 15, inclusive. When not present, the value of signalled_tile_group_index_length_minus1 is assumed as follows.
[0175] If rect_tile_group_info_in_pps_flag is equal to 1, Ceil( Log2( num_tile_groups_in_pic_minus1 + 1 ) ) - 1. If rect_tile_group_info_in_pps_flag is equal to 1, Ceil( Log2( num_tile_groups_in_pic_minus1 + 1 ) ) - 1.
[0176] Otherwise, Ceil(Log2(NumTilesInPic)) - 1.
[0177] The syntax and semantics of the general tile group header. [Table 3]
[0178] tile_group_address specifies the tile address of the first tile in the tile group. When it does not exist, the value of tile_group_address is assumed to be equal to 0.
[0179] When rect_tile_group_flag is equal to 0, the following applies. tile_group_address is the tile ID specified by Equation 6-7. The length of tile_group_address is Ceil(Log2(NumTilesInPic)) bits. The value of tile_group_address must be within the range of 0 to NumTilesInPic - 1, inclusive.
[0180] Otherwise, when rect_tile_group_flag is equal to 1 and rect_tile_group_info_in_pps is equal to 0, the following applies. tile_group_address is the tile index of the tile located at the upper left corner of the i-th tile group. The length of tile_group_address is signalled_tile_group_index_length_minus1 + 1 bits. When the signalled_tile_group_id_flag is equal to 0, the value of tile_group_address must be within the range of 0 to NumTilesInPic - 1, including both end values. Otherwise, the value of tile_group_address must be within the range of 0 to 2, including both end values. If not, the value of tile_group_address must be within the range of 0 to 2, including both end values. If not, the value of tile_group_address must be within the range of 0 to 2, including both end values. ( signalled_tile_group_index_length_minu s1 + 1 ) - 1, including both end values.
[0181] In other cases (when rect_tile_group_flag is equal to 1 and rect_tile_group_info_in_pps is equal to 1), the following applies. In other cases (when rect_tile_group_flag is equal to 1 and rect_tile_group_info_in_pps is equal to 1), the following applies. tile_group_address is the tile group ID of the tile group. The length of tile_group_address is signalled_tile_group_index_length_minus1 + 1 bits. The length of tile_group_address is signalled_tile_group_index_length_minus1 + 1 bits. When the signalled_tile_group_id_flag is equal to 0, the value of tile_group_address must be within the range of 0 to num_tile_groups_in_pic_minus1, including both end values. Otherwise, the value of tile_group_address must be within the range of 0 to 2, including both end values. If not, the value of tile_group_address must be within the range of 0 to 2, including both end values. If not, the value of tile_group_address must be within the range of 0 to 2, including both end values. ( signalled_tile_group_index _length_minus1 + 1 ) - 1, including both end values.
[0182] bottom_right_tile_id specifies the tile index of the tile located at the bottom right corner of the tile group. When single_tile_per_tile_group_flag is equal to 1, bottom_right_t bottom_right_tile_id specifies the tile index of the tile located at the bottom right corner of the tile group. When single_tile_per_tile_group_flag is equal to 1, bottom_right_t The ile_id is inferred to be equal to the tile_group_address. The length of a pixel element is Ceil(Log2(NumTilesInPic)) bits.
[0183] A variable, NumTilesInCurrTileGroup, that specifies the number of tiles in the current tile group. TopLeftTileIdx, which specifies the tile index of the top-left tile in the tile group; BottomRightTileIdx, which specifies the tile index of the bottom right tile of the TgTileIdx[ i ], which specifies the tile index of the i-th tile in the tile group, is It is derived as follows. if( rect_tile_group_flag ) { if ( tile_group_info_in_pps ) { tileGroupIdx = 0 while( tile_group_address != rect_tile_group_id[ tileGroupIdx ] ) tileGroupIdx++ tileIdx = top_left_tile_idx[ tileGroupIdx ] BottomRightTileIdx = bottom_right_tile_idx[ tileGroupIdx ] } else { tileIdx = tile_group_address BottomRightTileIdx = bottom_right_tile_id } TopLeftTileIdx = tileIdx deltaTileIdx = BottomRightTileIdx - TopLeftTileIdx NumTileRowsInTileGroupMinus1 = deltaTileIdx / ( num_tile_columns_minus1 + 1 ) (7-35) NumTileColumnsInTileGroupMinus1 = deltaTileIdx % ( num_tile_columns_minus1 + 1 ) NumTilesInCurrTileGroup = ( NumTileRowsInTileGroupMinus1 + 1 ) * ( NumTile ColumnsInTileGroupMinus1 + 1 ) for( j = 0, tIdx = 0; j < NumTileRowsInTileGroupMinus1 + 1; j++, tileIdx + = num_tile_columns_minus1 + 1 ) for( i = 0, currTileIdx = tileIdx; i < NumTileColumnsInTileGroupMinus1 + 1 ; i++, currTileIdx++, tIdx++ ) TgTileIdx[ tIdx ] = currTileIdx } else { NumTilesInCurrTileGroup = num_tiles_in_tile_group_minus1 + 1 TgTileIdx
[0000] = tile_group_address for( i = 1; i < NumTilesInCurrTileGroup; i++ ) TgTileIdx[ i ] = TgTileIdx[ i - 1 ] + 1 }
[0184] The recovery_poc_cnt specifies the recovery point of the decoded picture in the output order. CVS among the pictures in decoding order that follow the current picture (i.e., the GDR picture), and whose PicOrderCntVal is equal to the value of PicOrderCntVal of the current picture + recovery_poc_cnt if there is a picture picA, picA is called a recovery point picture. Otherwise, the first picture in output order that has a PicOrderCntVal greater than the value of PicOrderCntVal of the current picture + recovery_poc_cnt is called a recovery point picture. The recovery point picture must not precede the current picture in decoding order. All decoded pictures in output order are shown to be exactly or approximately exact within the content starting at the output order position of the recovery point picture. The value of recovery_poc_cnt must be within the range of -MaxPicOrderCntLsb / 2 to MaxPicOrderCntLsb / 2 - 1, inclusive. If so, picture picA is called a recovery point picture. Otherwise, if there is a picture picA, picA is called a recovery point picture. Otherwise, the first picture in output order that has a PicOrderCntVal greater than the value of PicOrderCntVal of the current picture + recovery_poc_cnt is called a recovery point picture. The recovery point picture must not precede the current picture in decoding order. All decoded pictures in output order are shown to be exactly or approximately exact within the content starting at the output order position of the recovery point picture. The value of recovery_poc_cnt must be within the range of -MaxPicOrderCntLsb / 2 to MaxPicOrderCntLsb / 2 - 1, inclusive. The first picture in output order that has a PicOrderCntVal greater than the value of PicOrderCntVal of the current picture + recovery_poc_cnt is called a recovery point picture. The recovery point picture must not precede the current picture in decoding order. All decoded pictures in output order are shown to be exactly or approximately exact within the content starting at the output order position of the recovery point picture. The value of recovery_poc_cnt must be within the range of -MaxPicOrderCntLsb / 2 to MaxPicOrderCntLsb / 2 - 1, inclusive. The recovery point picture must not precede the current picture in decoding order. All decoded pictures in output order are shown to be exactly or approximately exact within the content starting at the output order position of the recovery point picture. The value of recovery_poc_cnt must be within the range of -MaxPicOrderCntLsb / 2 to MaxPicOrderCntLsb / 2 - 1, inclusive. All decoded pictures in output order are shown to be exactly or approximately exact within the content starting at the output order position of the recovery point picture. The value of recovery_poc_cnt must be within the range of -MaxPicOrderCntLsb / 2 to MaxPicOrderCntLsb / 2 - 1, inclusive. The value of recovery_poc_cnt must be within the range of -MaxPicOrderCntLsb / 2 to MaxPicOrderCntLsb / 2 - 1, inclusive. The value of recovery_poc_cnt must be within the range of -MaxPicOrderCntLsb / 2 to MaxPicOrderCntLsb / 2 - 1, inclusive.
[0185] The value of RecoveryPointPocVal is derived as follows.
[0186] RecoveryPointPocVal = PicOrderCntVal + recovery_poc_cnt
[0187] A refreshed_region_flag equal to 1 specifies that the decoding of the tile group generates exact reconstructed sample values regardless of the value of NoIncorrectPicOutputFlag of the related GDR. A refreshed_region_flag equal to 0 specifies that the decoding of the tile group is A refreshed_region_flag equal to 1 specifies that the decoding of the tile group generates exact reconstructed sample values regardless of the value of NoIncorrectPicOutputFlag of the related GDR. A refreshed_region_flag equal to 0 specifies that the decoding of the tile group is A refreshed_region_flag equal to 1 specifies that the decoding of the tile group generates exact reconstructed sample values regardless of the value of NoIncorrectPicOutputFlag of the related GDR. A refreshed_region_flag equal to 0 specifies that the decoding of the tile group is When starting from the associated GDR where the Flag is equal to 1, it is specified that even if an inaccurate reconstructed sample value is generated. When it does not exist, the value of the refreshed_region_flag is assumed to be equal to 1. It is inferred.
[0188] Note x - The current picture itself is a GDR picture where the NoIncorrectPicOutputFlag is equal to 1. Obtained.
[0189] The boundaries at which the tile group is refreshed are derived as follows. tileColIdx = TopLeftTileIdx % ( num_tile_columns_minus1 + 1 ) tileRowIdx = TopLeftTileIdx / ( num_tile_columns_minus1 + 1 ) TGRefreshedLeftBoundary = ColBd[ tileColIdx ] << CtbLog2SizeY TGRefreshedTopBoundary = RowBd[ tileRowIdx ] << CtbLog2SizeY tileColIdx = BottomRightTileIdx % ( num_tile_columns_minus1 + 1 ) tileRowIdx = BottomRightTileIdx / ( num_tile_columns_minus1 + 1 ) TGRefreshedRightBoundary = ( ( ColBd[ tileColIdx ] + ColWidth[ tileColIdx ] ) << CtbLog2SizeY ) - 1 TGRefreshedRightBoundary = TGRefreshedRightBoundary > pic_width_in_luma_samp les ? pic_width_in_luma_samples : TGRefreshedRightBoundary TGRefreshedBotBoundary = ( ( RowBd[ tileRowIdx ] + RowHeight[ tileRowIdx ] ) << CtbLog2SizeY ) - 1 TGRefreshedBotBoundary = TGRefreshedBotBoundary > pic_height_in_luma_samples ? pic_height_in_luma_samples : TGRefreshedBotBoundary
[0190] Semantics of the NAL unit header.
Table 4
[0191] ...
[0192] When nal_unit_type is equal to GDR_NUT and the coded tile group belongs to a GDR picture the TemporalId shall be equal to 0.
[0193] The order of access units and their association with CVS are described.
[0194] A bitstream conforming to this specification (i.e., JVET contribution JVET-M1001-v5) contains one or more CVSs.
[0195] A CVS contains one or more access units. The order of NAL units and coded pictures and their association with access units are described in Section 7.4.2.4.4. specified.
[0196] The first access unit of the CVS is one of the following.
[0197] - The IRAP access unit where NoBrokenPictureOutputFlag is equal to 1.
[0198] - The GDR access unit where NoIncorrectPicOutputFlag is equal to 1.
[0199] When present, the next access unit after the access unit containing the end-of-sequence NAL unit or end-of-bitstream N AL unit shall be one of the following, which is a requirement for bitstream compliance.
[0200] - An IRAP access unit that may be an IDR access unit or a CRA access unit.
[0201] - A GDR access unit.
[0202] 8.1.1 The decoding process for the coded picture is described.
[0203] ...
[0204] When the current picture is an IRAP picture, the following applies.
[0205] - When the current picture is an IDR picture, the first picture in the bitstream in decoding order or the first picture following the end-of-sequence NAL unit in decoding order, the variable NoIncorrectPicOutputFlag is set equal to 1.
[0206] - Otherwise, by some external means (e.g., user input) not specified in this specification is available to set the variable HandleCraAsCvsStartFlag to the value for the current picture When that is the case, the variable HandleCraAsCvsStartFlag is equal to the value provided by external means is set, and the variable NoIncorrectPicOutputFlag is set equal to HandleCraAsCvsStartFlag is set.
[0207] - Otherwise, the variable HandleCraAsCvsStartFlag is set equal to 0, and the variable NoInco rrectPicOutputFlag is set equal to 0.
[0208] When the current picture is a GDR picture, the following applies.
[0209] - When the current picture is the first picture in the bitstream in decoding order for a GDR picture or the first picture following an end-of-sequence NAL unit in decoding order, the variable NoIncorrectPicOutputFlag is set equal to 1. is set.
[0210] - Otherwise, when some external means not specified in this specification is available to set the variable HandleGdrAsCvs StartFlag to the value for the current picture, the variable HandleG drAsCvsStartFlag is set equal to the value provided by external means, and the variable NoIncorrec tPicOutputFlag is set equal to HandleGdrAsCvsStartFlag.
[0211] - Otherwise, the variable HandleGdrAsCvsStartFlag is set equal to 0, and the variable NoInco The rrectPicOutputFlag is set equal to 0.
[0212] ...
[0213] The decoding process for the current picture CurrPic operates as follows.
[0214] 1. The decoding of NAL units is specified in Section 8.2.
[0215] 2. The process in Section 8.3 specifies the following decoding process that uses the syntax elements in the tile group header layer, and the above.
[0216] - Variables and functions related to the picture order count are derived as specified in Section 8.3.1. This needs to be called only for the first tile group of a picture.
[0217] - At the start of the decoding process for each tile group of a non-IDR picture, for the derivation of reference picture list 0 (RefPicList
[0000] ) and reference picture list 1 (RefPicList
[0001] ), the decoding process for the reference picture list construction specified in Section 8.3.2 is called.
[0218] - The decoding process for reference picture marking in Section 8.3.3 is called, and the reference pictures may be marked as "unused for reference" or "used for long-term reference". This is called only for the first tile group of a picture.
[0219] - PicOutputFlag is set as follows.
[0220] When one of the following conditions is true, PictureOutputFlag is set equal to 0 .
[0221] - The current picture is a RASL picture and the NoIncorrectPicOutput Flag of the related IRAP picture is equal to 1.
[0222] - gdr_enabled_flag is equal to 1 and the current picture is a GDR picture for which NoIncorrectPicOutputFlag is equal to 1.
[0223] - gdr_enabled_flag is equal to 1 and the current picture contains one or more tile groups for which refreshed_region_flag is equal to 0 and the NoBrokenPictureOutputFlag of the related GDR picture is equal to 1.
[0224] - Otherwise, PicOutputFlag is set equal to 1.
[0225] 3. The decoding process is called for coding tree units, scaling, conversion, in-loop filtering, etc.
[0226] 4. After all tile groups of the current picture have been decoded, the current decoded picture is marked as "used for short-term reference".
[0227] The decoding process for picture order count is described.
[0228] The output of this process is PicOrderCntVal, i.e., the picture order count of the current picture.
[0229] Each coded picture is associated with a picture order count variable shown as PicOrderCntVal. It is related to the variable.
[0230] When the current picture is not an IRAP picture with NoIncorrectPicOutputFlag equal to 1 or a GDR picture with NoIncorrectPicOutputFlag equal to 1, the variables prevPicOrderCntLsb and p revPicOrderCntMsb are derived as follows.
[0231] - The previous picture in decoding order that has a TemporalId equal to 0 and is not a RASL or RADL picture is set as prevTid0Pic. Let the previous picture be prevTid0Pic.
[0232] - The variable prevPicOrderCntLsb is set equal to the tile_group_pic_order_cnt_lsb of prevTid0Pic. It is set.
[0233] - The variable prevPicOrderCntMsb is set equal to the PicOrderCntMsb of prevTid0Pic.
[0234] The variable PicOrderCntMsb of the current picture is derived as follows.
[0235] - When the current picture is an IRAP picture with NoIncorrectPicOutputFlag equal to 1 or a GDR picture with NoIncorrectPicOutputFlag equal to 1, PicOrderCntMsb is set equal to 0. It is set. It is set.
[0236] - Otherwise, PicOrderCntMsb is derived as follows. if( ( tile_group_pic_order_cnt_lsb < prevPicOrderCntLsb ) && ( ( prevPicOrderCntLsb - tile_group_pic_order_cnt_lsb ) >= ( MaxPicOrderCntLsb / 2 ) ) ) PicOrderCntMsb = prevPicOrderCntMsb + MaxPicOrderCntLsb (8 - 1) else if( (tile_group_pic_order_cnt_lsb > prevPicOrderCntLsb ) && ( ( tile_gr oup_pic_order_cnt_lsb - prevPicOrderCntLsb ) > ( MaxPicOrderCntLsb / 2 ) ) ) PicOrderCntMsb = prevPicOrderCntMsb - MaxPicOrderCntLsb else PicOrderCntMsb = prevPicOrderCntMsb PicOrderCntVal is derived as follows. PicOrderCntVal = PicOrderCntMsb + tile_group_pic_order_cnt_lsb (8 - 2)
[0237] Note 1 - For an IRAP picture where NoIncorrectPicOutputFlag is equal to 1, PicOrderCntMsb is set to 0 so that all IRAP pictures where NoIncorrectPicOutputFlag is equal to 1 have a PicOrderCntVal equal to tile_group_pic_order_cnt_lsb.
[0238] Note 1 - For GDR pictures where NoIncorrectPicOutputFlag is equal to 1, PicOrderCntMsb is set to 0 and thus, all GDR pictures where NoIncorrectPicOutputFlag is equal to 1 have a PicOrderCntVal equal to tile_group_pic_order_cnt_lsb.
[0239] The value of PicOrderCntVal must be within the range of -2 31 ~2 31 - 1, inclusive. i.e.
[0240] When the current picture is a GDR picture, the value of LastGDRPocVal is set to be equal to PicOrderCntVal such that.
[0241] A decoding process for the boundary positions where the picture is refreshed is described.
[0242] This process is only called when gdr_enabled_flag is equal to 1.
[0243] This process is called after the tile group header syntax parsing is completed.
[0244] The output of this process is PicRefreshedLeftBoundaryPos, PicRefreshedRightBoundaryPo s, PicRefreshedTopBoundaryPos, and PicRefreshedBotBoundaryPos, i.e., the boundary positions of the refreshed area of the current picture.
[0245] Each coded picture has a refreshed area indicated as PicOrderCntVal It is related to the set of domain boundary position variables.
[0246] PicRefreshedLeftBoundaryPos, PicRefreshedRightBoundaryPos, PicRefreshedTopBoun daryPos, and PicRefreshedBotBoundaryPos are derived as follows.
[0247] When the tile group is the first received tile group of the current picture where refreshed_region_flag is equal to 1, the following applies. The following applies.
[0248] PicRefreshedLeftBoundaryPos = TGRefreshedLeftBoundary
[0249] PicRefreshedRightBoundaryPos = TGRefreshedRightBoundary
[0250] PicRefreshedTopBoundaryPos = TGRefreshedTopBoundary
[0251] PicRefreshedBotBoundaryPos = TileGroupBotBoundary
[0252] Otherwise, when refreshed_region_flag is equal to 1, the following applies.
[0253] PicRefreshedLeftBoundaryPos = TGRefreshedLeftBoundary < PicRefreshedLeftBounda ryPos?
[0254] TGRefreshedLeftBoundary : PicRefreshedLeftBoundaryPos
[0255] PicRefreshedRightBoundaryPos = TGRefreshedRightBoundary > PicRefreshedRightBou ndaryPos?
[0256] TGRefreshedRightBoundary : PicRefreshedRightBoundaryPos
[0257] PicRefreshedTopBoundaryPos = TGRefreshedTopBoundary < PicRefreshedTopBoundaryP os?
[0258] TGRefreshedTopBoundary : RefreshedRegionTopBoundaryPos
[0259] PicRefreshedBotBoundaryPos = TileGroupBotBoundary > PicRefreshedBotBoundaryPos ?
[0260] TileGroupBotBoundary : PicRefreshedBotBoundaryPos
[0261] The decoding process for the reference picture list configuration is described.
[0262] ...
[0263] For each current picture that is not an IRAP picture for which NoIncorrectPicOutputFlag is equal to 1 or a GDR picture for which NoIncorrectPicOutputFlag is equal to 1, maxPicOrderCnt - minPicOrd The value of erCnt must be less than MaxPicOrderCntLsb / 2, which is a requirement for bitstream - compliance.
[0264] ...
[0265] Decoding process for reference picture marking.
[0266] ...
[0267] If the current picture is an IRAP picture where NoIncorrectPicOutputFlag is equal to 1 or a GDR picture where NoIncorre ctPicOutputFlag is equal to 1, then (if any) all reference pictures in the DPB are marked as "unused for reference".
[0268] ...
[0269] The derivation process for time Luma motion vector prediction is described.
[0270] ...
[0271] The variable currCb specifies the current Luma coding block at the Luma location (xCb, yCb).
[0272] The variables mvLXCol and availableFlagLXCol are derived as follows.
[0273] - If tile_group_temporal_mvp_enabled_flag is equal to 0, both components of mvLXCol are set equal to 0, and availableFlagLXCol is set equal to 0.
[0274] - If not (tile_group_temporal_mvp_enabled_flag is equal to 1), the following order The following steps are applied.
[0275] 1. The bottom-right collocated motion vector is derived as follows.
[0276] xColBr = xCb + cbWidth (8-414)
[0277] yColBr = yCb + cbHeight (8-415)
[0278] leftBoundaryPos = gdr_enabled_flag?
[0279] PicRefreshedLeftBoundar of the picture referenced by RefPicList[ X ][ refIdxLX ] yPos :
[0280] 0 (8-415)
[0281] topBoundaryPos = gdr_enabled_flag?
[0282] PicRefreshedTopBoundary of the picture referenced by RefPicList[ X ][ refIdxLX ] Pos :
[0283] 0 (8-415)
[0284] rightBoundaryPos = gdr_enabled_flag?
[0285] PicRefreshedRightBounda of the picture referenced by RefPicList[ X ][ refIdxLX ] ryPos :
[0286] pic_width_in_luma_samples (8 - 415)
[0287] botBoundaryPos = gdr_enabled_flag?
[0288] PicRefreshedBotBoundary of the picture referenced by RefPicList[ X ][ refIdxLX ] Pos :
[0289] pic_height_in_luma_samples (8 - 415)
[0290] - yCb >> CtbLog2SizeY is equal to yColBr >> CtbLog2SizeY, and yColBr is t including both end values within the range from opBoundaryPos to botBoundaryPos, and xColBr is l including both end values If it is within the range from eftBoundaryPos to rightBoundaryPos, the following applies holds.
[0291] - The variable colCb specifies the modified location given by ((xColBr >> 3) << 3, (yColBr >> 3) << 3) inside the collocate picture specified by ColPic to capture the luma coding block.
[0292] - The luma location (xColCb, yColCb) is set equal to the top - left sample of the collocate luma coding block specified by colCb with respect to the top - left luma sample of the collocate picture specified by ColPic picture. is set equal to the top - left sample of the collocate luma coding block specified by colCb with respect to the top - left luma sample of the collocate picture specified by ColPic
[0293] - The derivation process for collocated motion vectors as specified in Section 8.5.2.12 is called with currCb, colCb, (xColCb, yColCb), refIdxLX, and sbFlag set equal to 0 as input, and the output is assigned to mvLXCol and availableFlagLXCol .
[0294] - Otherwise, both components of mvLXCol are set equal to 0 and availableFlagLXCol is set equal to 0.
[0295] 2. ...
[0296] The luma sample bilinear interpolation process is described.
[0297] The input to this process is as follows.
[0298] - The luma location in full sample units (xInt L , yInt L ),
[0299] - The luma location in fractional sample units (xFrac L , yFrac L ),
[0300] - The luma reference sample array refPicLX L ,
[0301] - The refreshed region boundaries of the reference picture PicRefreshedLeftBoundaryPos, PicRefres hedTopBoundaryPos, PicRefreshedRightBoundaryPos, and PicRefreshedBotBoundaryPo s.
[0302] ...
[0303] Full sample unit based luma location (xInt i , yInt i ) is derived as follows for i = 0..1 as follows
[0304] - When gdr_enabled_flag is equal to 1, the following applies
[0305] xInt i = Clip3( PicRefreshedLeftBoundaryPos, PicRefreshedRightBoundaryPos, xInt L + i ) (8-458)
[0306] yInt i = Clip3( PicRefreshedTopBoundaryPos, PicRefreshedBotBoundaryPos, yInt L + i ) (8-458)
[0307] - Otherwise (gdr_enabled_flag is equal to 0), the following applies
[0308] xInt i = sps_ref_wraparound_enabled_flag?
[0309] ClipH( ( sps_ref_wraparound_offset_minus1 + 1 ) * MinCbSizeY, picW, ( xInt L + i ) ) : (8-459)
[0310] Clip3( 0, picW - 1, xInt L + i )
[0311] yInti = Clip3( 0, picH - 1, yInt L + i ) (8-460)
[0312] ...
[0313] The process of 8-tap interpolation filter processing for luma samples is described.
[0314] The input to this process is as follows.
[0315] - Luma location in full sample units (xInt L , yInt L ),
[0316] - Luma location in fractional sample units (xFrac L , yFrac L ),
[0317] - Luma reference sample array refPicLX L ,
[0318] - A list padVal dir ] with dir = 0,1 specifying the direction and amount of reference sample padding.
[0319] - Refreshed region boundaries of the reference picture PicRefreshedLeftBoundaryPos, PicRefres hedTopBoundaryPos, PicRefreshedRightBoundaryPos, and PicRefreshedBotBoundaryPo s.
[0320] ...
[0321] The luma location in full sample units (xInt i , yInt i ) is derived as follows for i = 0..7 as follows.
[0322] - When the gdr_enabled_flag is equal to 1, the following applies.
[0323] xInt i = Clip3( PicRefreshedLeftBoundaryPos, PicRefreshedRightBoundaryPos, xInt L + i - 3 ) (8-830)
[0324] yInti = Clip3( PicRefreshedTopBoundaryPos, PicRefreshedBotBoundaryPos, yInt L + i - 3 ) (8-830)
[0325] - Otherwise (when gdr_enabled_flag is equal to 0), the following applies.
[0326] xInt i = sps_ref_wraparound_enabled_flag?
[0327] ClipH( ( sps_ref_wraparound_offset_minus1 + 1 ) * MinCbSizeY, picW, xInt L + i - 3 ) : (8-831)
[0328] Clip3( 0, picW - 1, xInt L + i - 3 )
[0329] yInt i = Clip3( 0, picH - 1, yInt L + i - 3 ) (8-832)
[0330] The chroma sample interpolation process is described.
[0331] The inputs to this process are as follows:
[0332] - Chroma location in full sample units (xInt C , yInt C ),
[0333] - Chroma location in 1 / 32 fractional sample units (xFrac C , yFrac C ),
[0334] - Chroma reference sample array refPicLX C .
[0335] - Refreshed region boundaries of the reference picture PicRefreshedLeftBoundaryPos, PicRefres hedTopBoundaryPos, PicRefreshedRightBoundaryPos, and PicRefreshedBotBoundaryPo s.
[0336] ...
[0337] The variable xOffset is set equal to (sps_ref_wraparound_offset_minus1 + 1) * MinCbSizeY) / SubWi dthC.
[0338] The chroma location in full sample units (xInt i , yInt i ) is derived as follows for i = 0..3 as follows:
[0339] - When gdr_enabled_flag is equal to 1, the following applies:
[0340] xInt i= Clip3( PicRefreshedLeftBoundaryPos / SubWidthC,
[0341] PicRefreshedRightBoundaryPos / SubWidthC, xInt L + i ) (8-844)
[0342] yInti = Clip3( PicRefreshedTopBoundaryPos / SubHeightC,
[0343] PicRefreshedBotBoundaryPos / SubHeightC, yInt L + i ) (8-844)
[0344] - Otherwise (if gdr_enabled_flag is equal to 0), the following applies.
[0345] xInt i = sps_ref_wraparound_enabled_flag? ClipH( xOffset, picW C , xInt C + i - 1 ) : (8-845)
[0346] Clip3( 0, picW C - 1, xInt C + i - 1 )
[0347] yInt i = Clip3( 0, picH C - 1, yInt C + i - 1 ) (8-846)
[0348] The deblocking filter process is described.
[0349] General process.
[0350] ...
[0351] The deblocking filter process is applied to all coding sub-block edges and transform block edges of the picture, except for the following types of edges.
[0352] - Edges on the boundary of the picture.
[0353] - Edges that coincide with the upper boundary of tile group tgA when all of the following are satisfied.
[0354] - gdr_enabled_flag is equal to 1.
[0355] - loop_filter_across_refreshed_region_enabled_flag is equal to 0.
[0356] - The edge coincides with the lower boundary of tile group tgB and the value of the refreshed_region_flag of tgB is different from the value of the refreshed_region_flag of tgA.
[0357] - Edges that coincide with the left boundary of tile group tgA when all of the following are satisfied.
[0358] - gdr_enabled_flag is equal to 1.
[0359] - loop_filter_across_refreshed_region_enabled_flag is equal to 0.
[0360] - The edge coincides with the right boundary of tile group tgB and the value of the refreshed_region_flag of tgB is different from the value of the refreshed_region_flag of tgA.
[0361] - When loop_filter_across_tiles_enabled_flag is equal to 0, the edges that coincide with the tile boundaries. Edges.
[0362] - For tile groups where tile_group_loop_filter_across_tile_groups_enabled_flag is equal to 0 or tile_group_deblocking_filter_disabled_flag is equal to 1, the edges that coincide with the upper or left boundaries. Edges. - Edges within a tile group where tile_group_deblocking_filter_disabled_flag is equal to 1.
[0363] - Edges that do not correspond to the 8×8 sample grid boundaries of the components to be considered. Edges.
[0364] - Edges within a chroma component where both sides of the edge use inter prediction.
[0365] - Edges of chroma transform blocks that are not edges of the related transform unit.
[0366] - Edges that cross the loop filter blocks of coding units where the IntraSubPartitionsSplit value is not equal to ISP_NO_SPLIT.
[0367] - The deblocking filter process for one direction is described. Edges that cross the loop filter blocks of coding units where the IntraSubPartitionsSplit value is not equal to ISP_NO_SPLIT.
[0368] The deblocking filter process for one direction is described.
[0369] ...
[0370] A coding unit having a coding block width log2CbW, a coding block height log2CbH, and the location (xCb, yCb) of the top - left sample of the coding block. A coding unit having a coding block width log2CbW, a coding block height log2CbH, and the location (xCb, yCb) of the top - left sample of the coding block. For each block, when edgeType is equal to EDGE_VER and xCb % 8 is equal to 0, or when edgeType is equal to EDGE_HOR and yCb % 8 is equal to 0, the edge is filtered by the following ordered steps. The edge is filtered.
[0371] 1. The coding block width nCbW is set equal to 1 << log2CbW, and the coding block height nCbH is set equal to 1 << log2CbH.
[0372] 2. The variable filterEdgeFlag is derived as follows.
[0373] - When edgeType is equal to EDGE_VER and one or more of the following conditions are true, filt erEdgeFlag is set equal to 0.
[0374] - The left boundary of the current coding block is the left boundary of the picture.
[0375] - The left boundary of the current coding block is the left boundary of the tile and loop_filter_acr oss_tiles_enabled_flag is equal to 0.
[0376] - The left boundary of the current coding block is the left boundary of the tile group and tile_gr oup_loop_filter_across_tile_groups_enabled_flag is equal to 0.
[0377] - The left boundary of the current coding block is the left boundary of the current tile group and all of the following conditions are satisfied.
[0378] - gdr_enabled_flag is equal to 1.
[0379] - The loop_filter_across_refreshed_region_enabled_flag is equal to 0.
[0380] - There is a tile group sharing a boundary with the left boundary of the current tile group, and the value of its ref reshed_region_flag is different from the value of the refreshed_region_flag of the current tile group. .
[0381] - Otherwise, if the edgeType is equal to EDGE_HOR and one or more of the following conditions are true , the variable filterEdgeFlag is set to be equal to 0.
[0382] - The upper boundary of the current luma coding block is the upper boundary of the picture.
[0383] - The upper boundary of the current coding block is the upper boundary of the tile, and loop_filter_acr oss_tiles_enabled_flag is equal to 0.
[0384] - The upper boundary of the current coding block is the upper boundary of the tile group, and tile_gr oup_loop_filter_across_tile_groups_enabled_flag is equal to 0.
[0385] - The upper boundary of the current coding block is the upper boundary of the current tile group, and the following all conditions are satisfied.
[0386] - The gdr_enabled_flag is equal to 1.
[0387] - The loop_filter_across_refreshed_region_enabled_flag is equal to 0.
[0388] - There exists a tile group sharing a boundary with the upper boundary of the current tile group, and the value of its ref reshed_region_flag is different from the value of the refreshed_region_flag of the current tile group .
[0389] - Otherwise, the filterEdgeFlag is set equal to 1.
[0390] When the tiles are integrated, the syntax is conformed.
[0391] 3. All elements of the two-dimensional (nCbW)×(nCbH) array edgeFlags are initialized to be equal to 0. initialized.
[0392] The CTB modification process for SAO is described.
[0393] ...
[0394] For all sample locations (xS i , yS j ) and (xY i , yY j ) with i = 0..nCtbSw - 1 and j = 0..nCtbSh - 1, depending on the values of pcm_loop_filter_disabled_flag, pcm_flag[x Y [yY i j , and cu_transquant_bypass_flag of the coding unit containing the coding block covering recPicture[xSi][ySj], the following applies .
[0395] - ....
[0396] Modify the highlighted section pending in the future decision conversion / quantization bypass.
[0397] - Instead, when SaoTypeIdx[ cIdx ][ rx ][ ry ] is equal to 2, the following ordered steps are applied.
[0398] 1. The values of hPos[ k ] and vPos[ k ] for k = 0..1 are specified in Table 8-18 based on SaoEoClass[ cIdx ][ rx ][ ry ].
[0399] 2. The variable edgeIdx is derived as follows.
[0400] - The modified sample locations (xS ik' , yS jk' ) and (xY ik' , yY ik' ) are derived as follows .
[0401] (xS ik' , yS jk' ) = (xS i + hPos[ k ], yS j + vPos[ k ] ) (8-1128)
[0402] (xY ik ', yY jk' ) = (cIdx == 0)? (xS ik' , yS jk' ) : (xS ik' * SubWidthC, yS jk' * SubHeightC ) (8-1129)
[0403] - All sample locations (xS with k = 0..1ik' , yS jk' ) and (xY ik' , yY jk' ) If one or more of the following conditions are true for, edgeIdx is set equal to 0 .
[0404] - Location (xS ik' , yS jk' ) The sample at is outside the picture boundary.
[0405] - gdr_enabled_flag is equal to 1, loop_filter_across_refreshed_region_enabled_fla g is equal to 0, the refreshed_region_flag of the current tile group is equal to 1, and the location (xS ik' , yS jk' ) The refreshed_region_flag of the tile group containing the sample at is equal to 0. .
[0406] - Location (xS ik' , yS jk' ) The sample at belongs to a different tile group and one of the following two conditions is true.
[0407] - MinTbAddrZs[xY ik' >> MinTbLog2SizeY ][yY jk' >> MinTbLog2SizeY ] is smaller than MinTbAdd rZs[xY i >> MinTbLog2SizeY ][yY j >> MinTbLog2SizeY ] and the sample recPi cture[xS i [yS jbelongs to the tile group in tile_group_loop_filter_across_til The e_groups_enabled_flag is equal to 0.
[0408] - MinTbAddrZs[xY i >> MinTbLog2SizeY][yY j >> MinTbLog2SizeY] is MinTbAddrZs xY ik' >> MinTbLog2SizeY][yY jk' >> MinTbLog2SizeY] is smaller than the sample recPic ture[xS ik' [yS jk' belongs to the tile group in tile_group_loop_filter_across_ The tile_groups_enabled_flag is equal to 0.
[0409] - The loop_filter_across_tiles_enabled_flag is equal to 0, and the sample at location (xS ik' , yS jk ' ) belongs to different tiles.
[0410] When a tile without a tile group is incorporated, the emphasized section is corrected .
[0411] - Otherwise, edgeIdx is derived as follows.
[0412] - The following applies.
[0413] edgeIdx = 2 + Sign(recPicture[xS i [yS j - recPicture[xS i+ hPos
[0000] ] yS j + vPos
[0000] ] ) +
[0414] Sign( recPicture[ xS i [ yS j - recPicture[ xS i + hPos
[0001] ][ yS j + vPos
[0001] ] ) (8-1130)
[0415] - When edgeIdx is equal to 0, 1, or 2, edgeIdx is modified as follows.
[0416] edgeIdx = ( edgeIdx == 2 )? 0 : ( edgeIdx + 1 ) (8-1131)
[0417] 3. The corrected picture sample array saoPicture[ xS i [ yS j is derived as follows is derived.
[0418] saoPicture[ xS i [ yS j = Clip3( 0, ( 1 << bitDepth ) - 1, recPicture[ xS i yS j +
[0419] SaoOffsetVal[ cIdx ][ rx ][ ry ][ edgeIdx ] ) (8-1132)
[0420] Coding tree block filtering processes for luma samples for ALF are described. are described.
[0421] ...
[0422] For the derivation of the filtered reconstructed luma sample alfPictureL[x][y], each reconstructed luma sample recPictureL[x [y] inside the current luma coding tree block is filtered as follows with x, y = 0..CtbSizeY - 1.
[0423] -...
[0424] - For each corresponding luma sample (x, y) inside a given array recPicture of luma samples, the location (h , v x ) is derived as follows. y
[0425] - If gdr_enabled_flag is equal to 1, loop_filter_across_refreshed_region_enabled_fl ag is equal to 0, and the refreshed_region_flag of the tile group tgA containing the luma sample at location (x, y ) is equal to 1, the following applies.
[0426] - If the location (h x , v y ) is located within another tile group tgB and the refresh ed_region_flag of tgB is equal to 0, the variables leftBoundary, rightBoundary, topBoundary, and botBoundary are set equal to TGRefreshedLeftBoundary, TGRefreshedRightBoundary, TGRefreshedTopBoundary, and TGRefreshedBotBoundary, respectively.
[0427] - Otherwise, the variables leftBoundary, rightBoundary, topBoundary, and botBoun dary are set equal to PicRefreshedLeftBoundaryPos, PicRefreshedRightBoundaryPos, Pic RefreshedTopBoundaryPos, and PicRefreshedBotBoundaryPos, respectively.
[0428] h x = Clip3( leftBoundary, rightBoundary, xCtb + x ) (8-1140)
[0429] v y = Clip3( topBoundary, botBoundary, yCtb + y ) (8-1141)
[0430] - Otherwise, the following applies.
[0431] h x = Clip3( 0, pic_width_in_luma_samples - 1, xCtb + x ) (8-1140)
[0432] v y = Clip3( 0, pic_height_in_luma_samples - 1, yCtb + y ) (8-1141)
[0433] -...
[0434] The derivation process for the ALF transposition and filter index for luma samples is described.
[0435] ...
[0436] The corresponding luma sample (x, y) inside the given array recPicture of luma samples The location (h x , v y ) for each is derived as follows.
[0437] - When gdr_enabled_flag is equal to 1 and loop_filter_across_refreshed_region_enabled_fl ag is equal to 0, and the refreshed_region_flag of tile group tgA including the luma sample at location (x, y) is equal to 1, the following applies.
[0438] - When the location (h x , v y ) is located within another tile group tgB and the refresh ed_region_flag of tgB is equal to 0, the variables leftBoundary, rightBoundary, topBoundary, and botBoundary are set equal to TGRefreshedLeftBoundary, TGRefreshedRightBoundary, TGRefreshedTopBoundary, and TGRefreshedBotBoundary, respectively.
[0439] - Otherwise, the variables leftBoundary, rightBoundary, topBoundary, and botBoun dary are set equal to PicRefreshedLeftBoundaryPos, PicRefreshedRightBoundaryPos, Pic RefreshedTopBoundaryPos, and PicRefreshedBotBoundaryPos, respectively.
[0440] h x= Clip3( leftBoundary, rightBoundary, x ) (8-1140)
[0441] v y = Clip3( topBoundary, botBoundary, y ) (8-1141)
[0442] - Otherwise, the following applies.
[0443] h x = Clip3( 0, pic_width_in_luma_samples - 1, x ) (8-1145)
[0444] v y = Clip3( 0, pic_height_in_luma_samples - 1, y ) (8-1146)
[0445] The coding tree block filtering process for chroma samples is described herein.
[0446] ...
[0447] For the derivation of the filtered reconstructed chroma sample alfPicture[ x ][ y ], each reconstructed chroma sample recPicture[ x ] y ] inside the current chroma coding tree block is filtered as follows with x, y = 0..ctbSizeC - 1.
[0448] - For each corresponding chroma sample ( x, y ) inside a given array recPicture of chroma samples, the location ( h x , v y ) is derived as follows.
[0449] - The gdr_enabled_flag is equal to 1, and the loop_filter_across_refreshed_region_enabled_fl ag is equal to 0, and when the refreshed_region_flag of the tile group tgA containing the luma sample at location (x, y) is equal to 1, the following applies.
[0450] - When the location (h x , v y ) is located within another tile group tgB and the refresh ed_region_flag of tgB is equal to 0, the variables leftBoundary, rightBoundary, topBoundary, and botBoundary are set equal to TGRefreshedLeftBoundary, TGRefreshedRightBoundary, TGRefreshedTopBoundary, and TGRefreshedBotBoundary, respectively.
[0451] - Otherwise, the variables leftBoundary, rightBoundary, topBoundary, and botBoun dary are set equal to PicRefreshedLeftBoundaryPos, PicRefreshedRightBoundaryPos, Pic RefreshedTopBoundaryPos, and PicRefreshedBotBoundaryPos, respectively.
[0452] h x = Clip3( leftBoundary / SubWidthC, rightBoundary / SubWidthC, xCtbC + x ) ( 8 - 1140)
[0453] v y= Clip3( topBoundary / SubWidthC, botBoundary / SubWidthC, yCtbC + y ) (8-1 141)
[0454] - If not, the following applies:
[0455] h x = Clip3( 0, pic_width_in_luma_samples / SubWidthC - 1, xCtbC + x ) (8-1177)
[0456] v y = Clip3( 0, pic_height_in_luma_samples / SubHeightC - 1, yCtbC + y ) (8-117 8)
[0457] FIG. 7 illustrates a gradual decoding refresh (GDR) technique 700 according to one embodiment of the present disclosure. The GDR technique 700 is similar to the GDR technique 500 of FIG. As used herein, a video bitstream 750 may be a coding Also called encoded video bitstream, bitstream, or any of their variants. As shown in FIG. 7, the bitstream 750 includes a sequence parameter set (SPS: seq picture parameter set (PPS) 754 , slice header 756 , and image data 758 .
[0458] SPS752 is a function that stores all pictures in a sequence of pictures (SOP). includes data common to it. In contrast, PPS754 includes data common to the entire picture. Sla ice header 756 includes information about the current slice, such as which of the reference pictures is used . SPS752 and PPS754 may be collectively referred to as parameter sets. SPS752, PPS754, and slice header 756 are of the type of network abstraction layer (NAL) units. An NAL unit is a syntax structure that includes a representation of the type of data to follow (e.g., coded video data). NAL units are classified into video coding layer (VCL) NAL units and non-VCL NAL units. A VCL NAL unit contains data representing the values of samples in a video picture, and a non-VCL NAL unit contains any additional relevant information, such as parameter sets (important header data that can be applied to multiple VCL NAL units), and supplementary enhancement ment information (timing information and other additional data that can improve the usefulness of the decoded video signal but is not necessary to decode the values of samples in the video picture). It will be understood by those skilled in the art that the bit stream 750 may include other parameters and information in actual applications.
[0459] The picture data 758 in FIG. 7 includes data related to the picture or video during encoding or decoding. The picture data 758 may simply be referred to as the payload or the data being carried in the bitstream 750. In one embodiment, the picture data 758 includes CVS708, which includes GDR picture 702, one or more trailing pictures 704, and recovery point pictures 706. Or, it includes a CLVS. In one embodiment, the trailing picture 704 is the picture that leads the recovery point picture 706 during the GDR period, and thus may be regarded as the format of the GDR picture. Since it precedes the recovery point picture 706, it may be regarded as the format of the GDR picture.
[0460] In one embodiment, the GDR picture 702, the trailing picture 704, and the recovery point picture 706 may define the GDR period within the CVS 708. In one embodiment, the decoding order starts with the GDR picture 702, follows the trailing picture 704, and then proceeds to the recovery picture 706. In one embodiment, the GDR picture 702, the trailing picture 704, and the recovery point picture 706 may define the GDR period within the CVS 708. In one embodiment, the decoding order starts with the GDR picture 702, follows the trailing picture 704, and then proceeds to the recovery picture 706. Since it precedes the recovery point picture 706, it may be regarded as the format of the GDR picture. Proceed to the recovery picture 706.
[0461] When a value (for example, 1) is received by the video decoder 30 via the user interface 84, the first flag is set equal to the value provided by the user interface (for example, an external input), and the second flag is set equal to the first flag to prevent the GDR picture 702 and any trailing picture 704 between the GDR picture 702 and the recovery point picture 706 from being output in the output order (for example, the presentation order 410). Instead, when no value is received by the video coder 30 via the user interface 84, the first flag and the second flag are set equal to different values (for example, 0). In one embodiment, when the first flag is set equal to the value provided by the user interface, the output of only the GDR picture 702 is prevented. When a value (for example, 1) is received by the video decoder 30 via the user interface 84, the GDR picture 702 and any trailing picture 704 between the GDR picture 702 and the recovery point picture 706 are output in the output order (for example, the presentation order 410). To prevent this, the first flag is set equal to the value provided by the user interface (for example, an external input), and the second flag is set equal to the first flag. Instead, when no value is received by the video coder 30 via the user interface 84, the first flag and the second flag are set equal to different values (for example, 0). In one embodiment, when the first flag is set equal to the value provided by the user interface, the output of only the GDR picture 702 is prevented. And any trailing picture 704 between the GDR picture 702 and the recovery point picture 706 are output in the output order (for example, the presentation order 410). To prevent this, the first flag is set equal to the value provided by the user interface (for example, an external input), and the second flag is set equal to the first flag. Instead, when no value is received by the video coder 30 via the user interface 84, the first flag and the second flag are set equal to different values (for example, 0). In one embodiment, when the first flag is set equal to the value provided by the user interface, the output of only the GDR picture 702 is prevented. To prevent the GDR picture 702 and any trailing picture 704 between the GDR picture 702 and the recovery point picture 706 from being output in the output order (for example, the presentation order 410), the first flag is set equal to the value provided by the user interface (for example, an external input), and the second flag is set equal to the first flag. Instead, when no value is received by the video coder 30 via the user interface 84, the first flag and the second flag are set equal to different values (for example, 0). In one embodiment, when the first flag is set equal to the value provided by the user interface, the output of only the GDR picture 702 is prevented. To prevent the GDR picture 702 and any trailing picture 704 between the GDR picture 702 and the recovery point picture 706 from being output in the output order (for example, the presentation order 410), the first flag is set equal to the value provided by the user interface (for example, an external input), and the second flag is set equal to the first flag. Instead, when no value is received by the video coder 30 via the user interface 84, the first flag and the second flag are set equal to different values (for example, 0). In one embodiment, when the first flag is set equal to the value provided by the user interface, the output of only the GDR picture 702 is prevented. Instead, when no value is received by the video coder 30 via the user interface 84, the first flag and the second flag are set equal to different values (for example, 0). In one embodiment, when the first flag is set equal to the value provided by the user interface, the output of only the GDR picture 702 is prevented. Instead, when no value is received by the video coder 30 via the user interface 84, the first flag and the second flag are set equal to different values (for example, 0). In one embodiment, when the first flag is set equal to the value provided by the user interface, the output of only the GDR picture 702 is prevented. In one embodiment, when the first flag is set equal to the value provided by the user interface, the output of only the GDR picture 702 is prevented. In one embodiment, when the first flag is set equal to the value provided by the user interface, the output of only the GDR picture 702 is prevented.
[0462] The CVS 708 is the coded video sequence for all the coded layer video sequences (CLVS) in the video bitstream 750. In particular, the video Is the coded video sequence for all the coded layer video sequences (CLVS) in the video bitstream 750. In particular, the video When the bitstream 750 contains a single layer, CVS and CLVS are the same. CVS differs from CLVS only when the stream 750 contains multiple layers.
[0463] As shown in FIG. 7, the GDR technique 700 or principle starts with a GDR picture 702 and The GDR technique 700 works over a series of pictures ending with the first picture 706. 702, trailing picture 704, and recovery point picture 706 correspond to GDR technique 5 of FIG. 00, a GDR picture 502, a trailing picture 504, and a recovery point picture 506. Thus, for simplicity, the manner in which the GDR technique 700 is implemented will be described with reference to FIG. Do not repeat.
[0464] As shown in FIG. 7, a GDR picture 702, a trailing picture 704, and a leading picture 705 in a CVS 708 are The coverage point pictures 706 are each contained within their own VCL NAL unit 730. The set of VCL NAL units 730 in 708 may be referred to as an access unit.
[0465] The NAL unit 730 containing the GDR picture 702 in CVS 708 is of the GDR NAL unit type (GDR_NU That is, in one embodiment, the NAL unit containing the GDR picture 702 in the CVS 708 The trailing picture 704 and the recovery point picture 706 are automatically In one embodiment, GDR_NUT is a unique NAL unit type for the bitstream. Instead of 750 having to start with an IRAP picture, bitstream 750 must start with a GDR picture 7 Enable starting with 02. Designating the VCL NAL unit 730 of the GDR picture 702 as GDR_NUT may indicate to the decoder, for example, that the initial VCL NAL unit 730 in the CVS708 contains the GDR picture 702. For example, it may be indicated to the decoder.
[0466] In one embodiment, the GDR picture 702 is the initial picture in the CVS708. In one embodiment, the GDR picture 702 is the initial picture during the GDR period. In one embodiment, the GDR picture 702 has a time identifier (ID) equal to 0. The time ID is a value or number that identifies the position or order of the picture relative to other pictures. In one embodiment, the GDR picture 702 is the initial picture during the GDR period. In one embodiment, the GDR picture 702 has a time identifier (ID) equal to 0. The time ID is a value or number that identifies the position or order of the picture relative to other pictures. For example, it may be indicated to the decoder. In one embodiment, an access unit containing the VCL NAL unit 730 with GDR_NUT is designated as a GDR access unit. In one embodiment, an access unit containing the VCL NAL unit 730 with GDR_NUT is designated as a GDR access unit. In one embodiment, the GDR picture 702 is the coded slice of another (e.g., larger) GDR picture. That is, the GDR picture 702 may be a part of a larger GDR picture. In one embodiment, the GDR picture 702 is the coded slice of another (e.g., larger) GDR picture. That is, the GDR picture 702 may be a part of a larger GDR picture.
[0467] FIG. 8 is one embodiment of a method 800 for decoding a coded video bitstream implemented by a video decoder (e.g., video decoder 30). The method 800 may be executed after the decoded bitstream is received directly or indirectly from a video encoder (e.g., video encoder 20). The method 800 improves the decoding process because sequential intra refresh enables random access without the need to use IRAP pictures. To prevent the GDR picture from being output, a filter FIG. 8 is one embodiment of a method 800 for decoding a coded video bitstream implemented by a video decoder (e.g., video decoder 30). The method 800 may be executed after the decoded bitstream is received directly or indirectly from a video encoder (e.g., video encoder 20). The method 800 improves the decoding process because sequential intra refresh enables random access without the need to use IRAP pictures. To prevent the GDR picture from being output, a filter FIG. 8 is one embodiment of a method 800 for decoding a coded video bitstream implemented by a video decoder (e.g., video decoder 30). The method 800 may be executed after the decoded bitstream is received directly or indirectly from a video encoder (e.g., video encoder 20). The method 800 improves the decoding process because sequential intra refresh enables random access without the need to use IRAP pictures. To prevent the GDR picture from being output, a filter FIG. 8 is one embodiment of a method 800 for decoding a coded video bitstream implemented by a video decoder (e.g., video decoder 30). The method 800 may be executed after the decoded bitstream is received directly or indirectly from a video encoder (e.g., video encoder 20). The method 800 improves the decoding process because sequential intra refresh enables random access without the need to use IRAP pictures. To prevent the GDR picture from being output, a filter FIG. 8 is one embodiment of a method 800 for decoding a coded video bitstream implemented by a video decoder (e.g., video decoder 30). The method 800 may be executed after the decoded bitstream is received directly or indirectly from a video encoder (e.g., video encoder 20). The method 800 improves the decoding process because sequential intra refresh enables random access without the need to use IRAP pictures. To prevent the GDR picture from being output, a filter FIG. 8 is one embodiment of a method 800 for decoding a coded video bitstream implemented by a video decoder (e.g., video decoder 30). The method 800 may be executed after the decoded bitstream is received directly or indirectly from a video encoder (e.g., video encoder 20). The method 800 improves the decoding process because sequential intra refresh enables random access without the need to use IRAP pictures. To prevent the GDR picture from being output, a filter The value of the flag is set by the user of the video decoder via a user interface (e.g., user interface 84 in FIG. 3, or some other external means). In one embodiment, to prevent any trailing pictures between the GDR picture and the recovery point picture in the GDR picture and output order from being output, the value of the flag is set by the user of the video decoder via a user interface (e.g., user interface 84 in FIG. 3, or some other external means). In this way, setting the flag prevents potentially dirty data from being output to the display and enables the video decoder to operate according to the user's preferences. Thus, in practice, the codec performance is improved, which leads to a better user experience. In block 802, the video decoder determines whether the value for the first flag is provided by an external input (e.g., user interface 84 in FIG. 3, or some other external means). In one embodiment, the external input is a graphical user interface (GUI) of the video decoder. In one embodiment, the user of the video decoder provides the value of the first flag using the external input. In one embodiment, the first flag is designated as Handle GdrAsCvsStartFlag. In block 804, when the value for the first flag is provided by the external input, the video decoder outputs a progressive decode refresh (GDR) picture (e.g., GDR picture 702) and prevents potentially dirty data from being output to the display and enables the video decoder to operate according to the user's preferences. Thus, in practice, the codec performance is improved, which leads to a better user experience. In block 802, the video decoder determines whether the value for the first flag is provided by an external input (e.g., user interface 84 in FIG. 3, or some other external means). In one embodiment, the external input is a graphical user interface (GUI) of the video decoder. In one embodiment, the user of the video decoder provides the value of the first flag using the external input. In one embodiment, the first flag is designated as Handle GdrAsCvsStartFlag.
[0468] In block 802, the video decoder determines whether the value for the first flag is provided by an external input (e.g., user interface 84 in FIG. 3, or some other external means). In one embodiment, the external input is a graphical user interface (GUI) of the video decoder. In one embodiment, the user of the video decoder provides the value of the first flag using the external input. In one embodiment, the first flag is designated as Handle GdrAsCvsStartFlag. In one embodiment, the external input is a graphical user interface (GUI) of the video decoder. In one embodiment, the user of the video decoder provides the value of the first flag using the external input. In one embodiment, the first flag is designated as Handle GdrAsCvsStartFlag. In block 804, when the value for the first flag is provided by the external input, the video decoder outputs a progressive decode refresh (GDR) picture (e.g., GDR picture 702) and prevents any trailing pictures between the GDR picture and the recovery point picture in the GDR picture and output order from being output.
[0469] In block 804, when the value for the first flag is provided by the external input, the video decoder outputs a progressive decode refresh (GDR) picture (e.g., GDR picture 702) and prevents any trailing pictures between the GDR picture and the recovery point picture in the GDR picture and output order from being output. Any trailing picture 704 between the GDR picture 702 and the recovery point picture 706 in the call output order is prevented from being output by setting a first flag equal to a value provided by an external input and setting a second flag equal to the first flag. In one embodiment, the value of the first flag is set to 1 to prevent the GDR picture and any trailing picture between the GDR picture and the recovery point picture in the call output order from being output. In one embodiment, when the value for the first flag is not provided by an external input, the value of the first flag is set to 0. To prevent any trailing picture 704 between the GDR picture 702 and the recovery point picture 706 in the call output order from being output, a first flag is set equal to a value provided by an external input and a second flag is set equal to the first flag. In one embodiment, the value of the first flag is set to 1. In one embodiment, when the value for the first flag is not provided by an external input, the value of the first flag is set to 0. In one embodiment, the GDR picture is the initial picture in the CVS of the coded video bitstream. In one embodiment, the GDR picture is the initial picture in the layer of the coded video bitstream. In one embodiment, the layer is the CLVS of the CVS of the coded video bitstream. In one embodiment, the GDR picture is the initial picture in the CVS of the coded video bitstream. In one embodiment, the GDR picture is the initial picture in the layer of the coded video bitstream. In one embodiment, the layer is the CLVS of the CVS of the coded video bitstream. In block 806, the video decoder decodes the GDR picture. The trailing picture and the recovery point picture are then decoded in order. In block 808, the video decoder stores the GDR picture in the decoded picture buffer (DPB). In one embodiment, when the output of the GDR picture is not restricted by the setting of the first and second flags, an image generated based on the GDR picture may be displayed for the user of an electronic device (e.g., a smartphone, tablet, laptop, personal computer, etc.).
[0470] In one embodiment, the GDR picture is the initial picture in the CVS of the coded video bitstream. In one embodiment, the GDR picture is the initial picture in the layer of the coded video bitstream. In one embodiment, the layer is the CLVS of the CVS of the coded video bitstream. In one embodiment, the GDR picture is the initial picture in the CVS of the coded video bitstream. In one embodiment, the GDR picture is the initial picture in the layer of the coded video bitstream. In one embodiment, the layer is the CLVS of the CVS of the coded video bitstream. In one embodiment, the GDR picture is the initial picture in the CVS of the coded video bitstream. In one embodiment, the GDR picture is the initial picture in the layer of the coded video bitstream. In one embodiment, the layer is the CLVS of the CVS of the coded video bitstream. In one embodiment, the GDR picture is the initial picture in the CVS of the coded video bitstream. In one embodiment, the GDR picture is the initial picture in the layer of the coded video bitstream. In one embodiment, the layer is the CLVS of the CVS of the coded video bitstream.
[0471] In block 806, the video decoder decodes the GDR picture. The trailing picture and the recovery point picture are then decoded in order. In block 808, the video decoder stores the GDR picture in the decoded picture buffer (DPB). In one embodiment, when the output of the GDR picture is not restricted by the setting of the first and second flags, an image generated based on the GDR picture may be displayed for the user of an electronic device (e.g., a smartphone, tablet, laptop, personal computer, etc.). In block 806, the video decoder decodes the GDR picture. The trailing picture and the recovery point picture are then decoded in order. In block 808, the video decoder stores the GDR picture in the decoded picture buffer (DPB). In one embodiment, when the output of the GDR picture is not restricted by the setting of the first and second flags, an image generated based on the GDR picture may be displayed for the user of an electronic device (e.g., a smartphone, tablet, laptop, personal computer, etc.). In block 806, the video decoder decodes the GDR picture. The trailing picture and the recovery point picture are then decoded in order. In block 808, the video decoder stores the GDR picture in the decoded picture buffer (DPB). In one embodiment, when the output of the GDR picture is not restricted by the setting of the first and second flags, an image generated based on the GDR picture may be displayed for the user of an electronic device (e.g., a smartphone, tablet, laptop, personal computer, etc.). In block 806, the video decoder decodes the GDR picture. The trailing picture and the recovery point picture are then decoded in order. In block 808, the video decoder stores the GDR picture in the decoded picture buffer (DPB). In one embodiment, when the output of the GDR picture is not restricted by the setting of the first and second flags, an image generated based on the GDR picture may be displayed for the user of an electronic device (e.g., a smartphone, tablet, laptop, personal computer, etc.). In one embodiment, when the output of the GDR picture is not restricted by the setting of the first and second flags, an image generated based on the GDR picture may be displayed for the user of an electronic device (e.g., a smartphone, tablet, laptop, personal computer, etc.). In one embodiment, when the output of the GDR picture is not restricted by the setting of the first and second flags, an image generated based on the GDR picture may be displayed for the user of an electronic device (e.g., a smartphone, tablet, laptop, personal computer, etc.). In one embodiment, when the output of the GDR picture is not restricted by the setting of the first and second flags, an image generated based on the GDR picture may be displayed for the user of an electronic device (e.g., a smartphone, tablet, laptop, personal computer, etc.).
[0472] FIG. 9 is a schematic diagram of a video coding device 900 (e.g., a video encoder 20 or a video decoder 30) according to an embodiment of the present disclosure. The video coding device 900 is suitable for implementing the disclosed embodiments as described herein. The video co ding device 900 includes an input port 910 and a receiver unit (Rx) 920 for receiving data, a processor, a logic unit, or a central processing unit ( CPU) 930 for processing data, a transmitter unit (Tx) 940 and an output port 950 for transmitting data, and a memory 960 for storing data. The video coding device 900 may also include optical-to-electrical (OE) and electrical-to-optical (EO) components coupled to the input port 910, the receiver unit 920, the transmitter unit 940, and the output port 950 for the exit or entry of optical or electrical signals.
[0473] The processor 930 is implemented by hardware and software. The processor 9 30 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), and a digital signal processor (DSP). The processor 930 communicates with the input port 910, the receiver uni t 920, the transmitter unit 940, the output port 950, and the memory 960. The pro cessor 930 includes a coding module 970. The coding module 970 implements the disclosed embodiments described above. For example, the coding module 970 may perform various operations as described above. Implement, process, prepare, or provide various codec functions. Thus, including the coding module 970 can bring about a significant improvement to the functions of the video coding device 900 and affect the conversion of the video coding device 900 to different states. Alternatively, the coding module 970 is implemented as instructions stored in the memory 960 and executed by the processor 930.
[0474] The video coding device 900 may also include an input and / or output (I / O) device 980 for communicating data with the user. The I / O device 980 may include output devices such as a display for displaying video data, speakers for outputting audio data, etc. The I / O device 980 may also include input devices such as a keyboard, mouse, trackball, etc., and / or a corresponding interface for interacting with such output devices. In one embodiment, the I / O device 980 is an external means used by the user of the video coding device 900 to input the value of a first flag.
[0475] The memory 960 includes one or more disks, tape drives, and solid state drives for storing such programs when a program is selected for execution and for storing the instructions and data read during program execution, and may be used as an auxiliary flow data storage device. The memory 960 may be volatile and / or non-volatile, such as read-only memory (ROM), random access memory (RAM), It may be a ternary content-addressable memory (TCAM) and / or a static random access memory (SRAM).
[0476] FIG. 10 is a schematic diagram of an embodiment of means 1000 for coding. In one embodiment, the means 1000 for coding is implemented in a video coding device 1002 (e.g., a video encoder 20 or a video decoder 30). The video coding device 1002 includes receiving means 1001. The receiving means 1001 is configured to receive a picture to be coded or a bitstream to be decoded. The video coding device 1002 includes transmitting means 1007 coupled to the receiving means 1001. The transmitting means 1007 is configured to transmit a bitstream to a decoder or to transmit a decoded image to a display means (e.g., one of the I / O devices 1080).
[0477] The video coding device 1002 includes storage means 1003. The storage means 1003 is coupled to at least one of the receiving means 1001 or the transmitting means 1007. The storage means 1003 is configured to store instructions. The video coding device 1002 also includes processing means 1005 . The processing means 1005 is coupled to the storage means 1003. The processing means 1005 is configured to execute instructions stored in the storage means 1003 to execute the methods disclosed herein .
[0478] The steps of the exemplary methods described herein are not necessarily executed in the order described It should also be understood that the order of steps in such methods is not necessarily required. It should be understood that such methods are merely exemplary. Similarly, additional steps may be included in such methods. Some steps may be included or omitted in methods consistent with various embodiments of the present disclosure. They may be omitted or combined.
[0479] Although several embodiments are provided in this disclosure, the disclosed system and method include: This disclosure may be embodied in many other specific forms without departing from its spirit or scope. It should be understood that the examples are to be regarded as illustrative rather than limiting, and that the present invention is not intended to be limiting. The present invention is not limited to the details provided herein. For example, various elements or components may be The elements may be combined or integrated in another system, or any number of Some features may be omitted or not implemented.
[0480] In addition, the techniques and systems described and illustrated in various embodiments as being separate or distinct may be used in various ways. The systems, subsystems, and methods may be adapted for use with other systems without departing from the scope of the present disclosure. The present invention may be combined or integrated with any other system, module, technique, or method. may not be shown or described as being joined or directly coupled or in communication with each other. Other items described herein may include, but are not limited to, any of the following: Indirectly coupled through some interface, device, or intermediate component Other examples of modifications, substitutions, and alterations may be identified by those of ordinary skill in the art. is permissible and may be made without departing from the spirit and scope of the disclosure herein. do. [Explanation of symbols]
[0481] 10 Coding system 12 Source device, video device 14 Destination device 16 Computer-readable medium 18 Video source 20 Video encoder 22 Output interface 28 Input interface 30 Video decoder 32 Display device 40 Mode selection unit 42 Motion estimation unit 44 Motion compensation unit 46 Intra prediction unit 48 Partitioning unit 50 Adder 52 Transformation processing unit 54 Quantization unit 56 Entropy coding unit, entropy encoding unit 58 Inverse quantization unit 60 Inverse transformation unit 62 Adder 64 Reference frame memory 70 Entropy decoding unit 72 Motion compensation unit 74 Intra prediction unit 76 Inverse quantization unit 78 Inverse transformation unit 80 Adder 82 Reference frame memory 84 User interface (UI) 402 IRAP picture 404 Leading picture 406 Trailing picture 408 Decoding order 410 Presentation order 502 GDR picture 504 Trailing picture 506 Recovery Point Picture 508 Coded Video Sequence 510 Refreshed / Clean Region 512 Non-Refreshed / Dirty Region 602 Current Picture 604 Reference Picture, Refreshed Region 606 Refreshed Region 608 Non-Refreshed Region, Refreshed Region 610 Motion Vector 612 Reference Block 614 Current Block 702 GDR Picture 704 Trailing Picture 706 Recovery Point Picture 708 CVS 730 NAL Unit 750 Video Bitstream 752 Sequence Parameter Set (SPS) 754 Picture Parameter Set (PPS) 756 Slice Header 758 Image Data 900 Video Coding Device 910 Input Port 920 Receiver Unit (Rx) 930 Processor, Logic Unit, Central Processing Unit (CPU) 940 Transmitter Unit (Tx) 950 Output Port 960 Memory 970 Coding Module 980 Input and / or Output (I / O) Device 1000 Means for Coding 1001 Receiving Means 1002 Video Coding Device 1003 Storage Means 1005 Processing Means 1007 Transmitting Means
Claims
1. A method performed by a video encoder, comprising: determining whether a value for a first flag is provided by an external means; when the value for the first flag is provided by the external means, setting the value of the first flag to be equal to the value provided by the external means and setting the value of a second flag to be equal to the value of the first flag, in order to prevent a coded progressive decode refresh (GDR) picture from being output; encoding the coded GDR picture; and a method comprising the above.
2. The method according to claim 1, further preventing any trailing picture between the coded GDR picture and the recovery point picture in the output order from being output when the value for the first flag is provided by the external means, by setting the value of the second flag to be equal to the value of the first flag.
3. The method according to claim 1, wherein the external means is a graphic user interface (GUI) of the video encoder, and the value of the first flag is provided by a user of the video encoder using the external means.
4. The method according to claim 1, wherein the first flag is designated as HandleGdrAsCvsStartFlag.
5. The method according to claim 2, wherein the value of the second flag is set to 1 in order to prevent the coded GDR picture and any trailing picture between the coded GDR picture and the recovery point picture in the output order from being output.
6. The method according to claim 1, wherein when the value for the first flag is not provided by the external means, the value of the first flag is set to 0.
7. An encoding device, comprising: a receiver configured to receive a picture to be encoded; a memory coupled to the receiver for storing instructions; and a processor coupled to the memory, the processor causing the encoding device to: determine whether a value for a first flag is provided by an external means; When the value for the first flag is provided by the external means, in order to prevent a coded Gradual Decoding Refresh (GDR) picture from being output, setting the value of the first flag equal to the value provided by the external means and setting the value of the second flag equal to the value of the first flag; encoding the coded GDR picture; configured to execute the instructions to cause; an encoding device.
8. The encoding device according to claim 7, wherein when the value for the first flag is provided by the external means, in order to further prevent a coded Gradual Decoding Refresh (GDR) picture and any trailing pictures between the coded GDR picture and the recovery point picture in output order from being output, the value of the second flag is set equal to the value of the first flag.
9. The encoding device according to claim 7, wherein the external means is a Graphic User Interface (GUI) of a video encoder, and the value of the first flag is provided by a user of the video encoder using the external means.
10. The encoding device according to claim 7, wherein the first flag is designated as HandleGdrAsCvsStartFlag.
11. The encoding device according to claim 8, wherein in order to prevent the coded GDR picture and any trailing pictures between the coded GDR picture and the recovery point picture in output order from being output, the value of the second flag is set to 1.
12. The encoding device according to claim 7, wherein when the value for the first flag is not provided by the external means, the value of the first flag is set to 0.
13. a receiver configured to receive a picture to be encoded; a transmitter coupled to the receiver and configured to transmit a bitstream to a decoder; a memory coupled to at least one of the receiver or the transmitter and configured to store instructions; A processor coupled to the memory, configured to execute the instructions stored in the memory in order to execute the method according to any one of claims 1 to 6, and a processor A coding device comprising **Claim 14** A decoder, and An encoder communicating with the decoder, the encoder including an encoding device or a coding device according to any one of claims 7 to 13, and an encoder A system comprising **Claim 15** Receiving means configured to receive a picture to be coded, Transmitting means coupled to the receiving means, configured to transmit a bitstream to decoding means, and transmitting means Storage means coupled to at least one of the receiving means or the transmitting means, configured to store instructions, and storage means Processing means coupled to the storage means, configured to execute the instructions stored in the storage means in order to execute the method according to any one of claims 1 to 6, and processing means Means for coding comprising **Claim 16** Determining whether a value for a first flag is provided by an external means, and when the value for the first flag is provided by the external means, setting the value of the first flag to be equal to the value provided by the external means and setting the value of a second flag to be equal to the value of the first flag in order to prevent a coded progressive refresh (GDR) picture from being output, and a determining unit configured to perform the above, An encoding unit configured to encode the coded GDR picture, and An encoder comprising **Claim 17** A non-transitory computer-readable medium including a computer program for use by a video coding device, the computer program being executable instructions stored in the non-transitory computer-readable medium, including instructions that, when executed by a processor, cause the video coding device to perform the method according to any one of claims 1 to 6, and a non-transitory computer-readable medium An apparatus for storing and transmitting an encoded bitstream for a video signal including a plurality of syntax elements, comprising: a receiver configured to receive a picture to be encoded; a memory storing instructions; a transmitter configured to transmit the encoded bitstream for the video signal; a processor coupled to the receiver, the memory, and the transmitter; wherein when executed by the processor, the instructions cause the processor to determine whether a value for a first flag is provided by external means; when the value for the first flag is provided by external means, set the value of the first flag equal to the value provided by the external means and set the value of a second flag equal to the value of the first flag to prevent a coded progressive refresh (GDR) picture from being output; code a GDR picture; store in the memory an encoded bitstream for the video signal, including the coded GDR and the plurality of syntax elements including the first flag and the second flag; An apparatus that causes the above to be executed.
Citation Information
Patent Citations
JPP7302000B