Processing of Multiple Image Sizes and Conformance Windows for Reference Image Resampling in Video Coding
By ensuring that the image parameter set with the same image size has the same compliance window, the complexity of multi-image size and window processing in video decoding is solved, and the resource utilization rate is reduced and video quality is improved.
Patent Information
- Application Number
- CN202310336555.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-07-08
- Filing Date
- 2020-07-07
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-07-07
AI Technical Summary
In video decoding, the prior art is difficult to effectively process multiple image sizes and compliance windows, resulting in high resource utilization of encoder and decoder side, affecting video quality and transmission efficiency.
Avoid overly complex processing when enabling reference image resampling by ensuring that image parameter sets with the same image size have the same compliance window. The specific implementation includes a video decoder and an encoder receiving an image parameter set and determining the size of the compliance window based on the image width and height.
The utilization rate of processors, memory and network resources on the encoder and decoder side is reduced, the video decoding process is improved, and the user experience and higher transmission efficiency are provided.
Smart Images

Figure CN116347101B_ABST
Abstract
Description
[0001] This application is a divisional application. The application number of the original application is 202080045733.0, the filing date of the original application is July 7, 2020, and the entire content of the original application is incorporated herein by reference.
[0002] Cross - reference to related applications
[0003] This patent application claims the benefit of U.S. Provisional Patent Application No. 62 / 871,493, filed on July 8, 2019, by Jianle Chen et al., titled "Handling of Multiple Picture Size and Conformance Windows for Reference Picture Resampling in Video Coding", the entire content of which is incorporated herein by reference. Technical field
[0004] The present invention generally describes techniques for supporting multiple picture sizes and conformance windows in video coding. More specifically, the present invention ensures that picture parameter sets with the same picture size also have the same conformance window. Background art
[0005] Even in the case of short videos, a large amount of video data is required for description, which may cause difficulties when the data is to be streamed or otherwise transmitted in a communication network with limited bandwidth capacity. Therefore, video data is usually compressed first and then sent in modern telecommunication networks. Since memory resources may be limited, the size of the video may also be a problem when storing the video on a storage device. Video compression devices typically use software and / or hardware on the source side to encode video data and then transmit or store it, thereby reducing the amount of data required to represent digital video images. Then, a video decompression device that decodes the video data receives the compressed data on the destination side. In the context of limited network resources and the growing demand for higher video quality, improved compression and decompression techniques are needed, which can increase the compression ratio with little impact on image quality. Summary of the invention
[0006] A first aspect relates to a method for decoding a decoded video bitstream implemented by a video decoder. The method includes: the video decoder receives a first picture parameter set and a second picture parameter set that refer to the same sequence parameter set, wherein when the picture width and picture height of the first picture parameter set and the second picture parameter set have the same values, the compliance window of the first picture parameter set and the second picture parameter set has the same value; the video decoder applies the compliance window to a current picture corresponding to the first picture parameter set or the second picture parameter set.
[0007] The method provides a technique for constraining picture parameter sets with the same picture size to also have the same compliance window size (e.g., cropping window size). By keeping the compliance window sizes of picture parameter sets with the same picture size the same, overly complex processing can be avoided when reference picture resampling (RPR) is enabled. Therefore, the utilization of processor, memory, and / or network resources on the encoder and decoder sides can be reduced. Thus, the encoder / decoder (also known as "codec") in video coding is improved compared to the current codec. In fact, the improved video coding process provides a better user experience when sending, receiving, and / or viewing videos.
[0008] Optionally, according to any of the above aspects, in another implementation of the aspect, the compliance window includes a compliance window left offset, a compliance window right offset, a compliance window top offset, and a compliance window bottom offset.
[0009] Optionally, according to any of the above aspects, in another implementation of the aspect, the instruction further causes the decoding device to, after applying the compliance window, perform decoding on the current picture corresponding to the first picture parameter set or the second picture parameter set using inter prediction, wherein the inter prediction is based on a resampled reference picture.
[0010] Optionally, according to any of the above aspects, in another implementation of the aspect, the method further includes resampling a reference picture associated with the current picture corresponding to the first picture parameter set or the second picture parameter set using reference picture resampling (RPS).
[0011] Optionally, according to any of the above aspects, in another implementation of the aspect, the resampling of the reference picture changes the resolution of the reference picture, and the reference picture is used for inter prediction of the current picture corresponding to the first picture parameter set or the second picture parameter set.
[0012] Optionally, according to any of the above aspects, in another implementation of the aspect, the image width and the image height are measured in luminance samples.
[0013] Optionally, according to any of the above aspects, in another implementation of the aspect, the method further includes determining whether to enable bi-direction optical flow (BDOF) to decode the image according to the image width, the image height, and the compliance window of the current image and a reference image of the current image.
[0014] Optionally, according to any of the above aspects, in another implementation of the aspect, the method further includes determining whether to enable decoder-side motion vector refinement (DMVR) to decode the image according to the image width, the image height, and the compliance window of the current image and a reference image of the current image.
[0015] Optionally, according to any of the above aspects, in another implementation of the aspect, the method further includes displaying an image generated using the current block on a display of an electronic device.
[0016] A second aspect relates to a method for encoding a video bitstream implemented by a video encoder. The method includes: the video encoder generating a first image parameter set and a second image parameter set that reference the same sequence parameter set, wherein when the image widths and image heights of the first image parameter set and the second image parameter set have the same values, the compliance windows of the first image parameter set and the second image parameter set have the same values; the video encoder encoding the first image parameter set and the second image parameter set into the video bitstream; the video encoder storing the video bitstream, wherein the video bitstream is for sending to a video decoder.
[0017] The method provides a technique for constraining image parameter sets having the same image size to also have the same compliance window size (e.g., crop window size). By keeping the sizes of the compliance windows of image parameter sets having the same image size the same, overly complex processing can be avoided when enabling reference picture resampling (RPR). Therefore, the utilization of processors, memories, and / or network resources on the encoder and decoder sides can be reduced. Thus, the encoder / decoder (also known as the "codec") in video coding is improved compared to the current codec. In fact, the improved video coding process provides a better user experience when sending, receiving, and / or viewing videos.
[0018] Optionally, according to any of the above aspects, in another implementation of the aspect, the compliance window includes a left offset of the compliance window, a right offset of the compliance window, an upper offset of the compliance window, and a lower offset of the compliance window.
[0019] Optionally, according to any of the above aspects, in another implementation of the aspect, the image width and the image height are measured in units of luminance samples.
[0020] Optionally, according to any of the above aspects, in another implementation of the aspect, the method further includes sending the video bitstream including the first picture parameter set and the second picture parameter set to the video decoder.
[0021] A third aspect relates to a decoding device. The decoding device includes: a receiver for receiving a decoded video bitstream; a memory coupled to the receiver and storing instructions; and a processor coupled to the memory and configured to execute the instructions to cause the decoding device to: receive a first picture parameter set and a second picture parameter set that refer to the same sequence parameter set, wherein when the picture widths and picture heights of the first picture parameter set and the second picture parameter set have the same values, the compliance windows of the first picture parameter set and the second picture parameter set have the same values; and apply the compliance window to a current picture corresponding to the first picture parameter set or the second picture parameter set.
[0022] The decoding device provides a technique for constraining picture parameter sets having the same picture size to also have the same compliance window size (e.g., cropping window size). By keeping the sizes of the compliance windows of picture parameter sets having the same picture size the same, overly complex processing can be avoided when enabling reference picture resampling (RPR). Accordingly, the utilization of processor, memory, and / or network resources on the encoder and decoder sides can be reduced. Thus, the encoder / decoder (also known as a "codec") in video coding is improved relative to current codecs. In fact, the improved video coding process provides a better user experience when sending, receiving, and / or viewing video.
[0023] Optionally, according to any of the above aspects, in another implementation of the aspect, the compliance window includes a left offset of the compliance window, a right offset of the compliance window, an upper offset of the compliance window, and a lower offset of the compliance window.
[0024] Optionally, according to any of the above aspects, in another implementation of the aspect, the instruction further causes the decoding device to use inter prediction to decode the current image corresponding to the first image parameter set or the second image parameter set after applying the compliance window, wherein the inter prediction is based on a resampled reference image.
[0025] Optionally, according to any of the above aspects, in another implementation of the aspect, the decoding device further includes a display, wherein the display is configured to display an image generated based on the current image.
[0026] A fourth aspect relates to an encoding device. The encoding device includes: a memory including instructions; a processor coupled to the memory and configured to implement the instructions to cause the encoding device to: generate a first image parameter set and a second image parameter set that reference the same sequence parameter set, wherein when the image width and image height of the first image parameter set and the second image parameter set have the same values, the compliance windows of the first image parameter set and the second image parameter set have the same values; encode the first image parameter set and the second image parameter set into a video bitstream; a transmitter coupled to the processor and configured to send the video bitstream including the first image parameter set and the second image parameter set to a video decoder.
[0027] The encoding device provides a technique for constraining image parameter sets having the same image size to also have the same compliance window size (e.g., cropping window size). By keeping the compliance window sizes of image parameter sets having the same image size the same, overly complex processing can be avoided when enabling reference picture resampling (RPR). Accordingly, the utilization of processor, memory, and / or network resources on the encoder and decoder sides can be reduced. Thus, the encoder / decoder (aka "codec") in video coding is improved relative to current codecs. In fact, the improved video coding process provides a better user experience when sending, receiving, and / or viewing video.
[0028] Optionally, according to any of the above aspects, in another implementation of the aspect, the compliance window includes a compliance window left offset, a compliance window right offset, a compliance window top offset, and a compliance window bottom offset.
[0029] Optionally, according to any of the above aspects, in another implementation of the aspect, the image width and the image height are measured in terms of luminance samples.
[0030] A fifth aspect relates to a decoding apparatus. The decoding apparatus includes: a receiver for receiving an image for encoding or receiving a bitstream for decoding; a transmitter coupled to the receiver and for transmitting the bitstream to a decoder or transmitting the decoded image to a display; a memory coupled to at least one of the receiver or the transmitter and for storing instructions; and a processor coupled to the memory and for executing the instructions stored in the memory to perform any of the methods disclosed herein.
[0031] The decoding apparatus provides a technique for constraining image parameter sets having the same image size to also have the same compliance window size (e.g., cropping window size). By keeping the size of the compliance window of image parameter sets having the same image size the same, overly complex processing can be avoided when enabling reference picture resampling (RPR). Accordingly, the utilization of processor, memory, and / or network resources on the encoder and decoder sides can be reduced. Thus, the encoder / decoder (also known as a "codec") in video decoding is improved relative to current codecs. In fact, the improved video decoding process provides a better user experience when sending, receiving, and / or viewing video.
[0032] Optionally, according to any of the above aspects, in another implementation of the aspect, the decoding apparatus further includes a display, wherein the display is for displaying an image.
[0033] A sixth aspect relates to a system. The system includes: an encoder; and a decoder in communication with the encoder, wherein the encoder or the decoder includes the decoding device, encoding device, or decoding apparatus disclosed herein.
[0034] The system provides a technique for constraining image parameter sets having the same image size to also have the same compliance window size (e.g., cropping window size). By keeping the size of the compliance window of image parameter sets having the same image size the same, overly complex processing can be avoided when enabling reference picture resampling (RPR). Accordingly, the utilization of processor, memory, and / or network resources on the encoder and decoder sides can be reduced. Thus, the encoder / decoder (also known as a "codec") in video decoding is improved relative to current codecs. In fact, the improved video decoding process provides a better user experience when sending, receiving, and / or viewing video.
[0035] The seventh aspect relates to a decoding module. The decoding module includes: a receiving module for receiving an image for encoding or receiving a bitstream for decoding; a transmitting module coupled to the receiving module and for transmitting the bitstream to a decoding module or transmitting a decoded image to a display module; a storage module coupled to at least one of the receiving module or the transmitting module and for storing instructions; and a processing module coupled to the storage module and for executing the instructions stored in the storage module to perform any one of the methods disclosed herein.
[0036] The decoding module provides a technique for constraining a set of picture parameters having the same picture size to also have the same compliance window size (e.g., cropping window size). By keeping the size of the compliance window of a set of picture parameters having the same picture size the same, overly complex processing can be avoided when enabling reference picture resampling (RPR). Accordingly, the utilization of processor, memory, and / or network resources on the encoder and decoder sides can be reduced. Thus, the encoder / decoder (aka "codec") in video decoding is improved relative to current codecs. In fact, the improved video decoding process provides a better user experience when sending, receiving, and / or viewing video.
[0037] For clarity of description, any one of the above embodiments may be combined with any one or more of the other above embodiments to create new embodiments within the scope of the present invention.
[0038] These and other features will be more clearly understood from the following detailed description in conjunction with the drawings and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] To understand the present invention more thoroughly, reference is now made to the following brief description, taken in conjunction with the accompanying drawings and detailed description, in which like reference numerals represent like parts.
[0040] Figure 1 A flowchart of an exemplary method for decoding a video signal.
[0041] Figure 2 A schematic diagram of an exemplary encoding and decoding (codec) system for video decoding.
[0042] Figure 3 A schematic diagram of an exemplary video encoder.
[0043] Figure 4 A schematic diagram of an exemplary video decoder.
[0044] Figure 5A decoded video sequence representing the relationship of an intra random access point (IRAP) image with respect to the preceding and succeeding images in the decoding order and the presentation order.
[0045] Figure 6 An example of multi-layer decoding for spatial scalability is shown.
[0046] Figure 7 A schematic diagram of an example of unidirectional inter prediction.
[0047] Figure 8 A schematic diagram of an example of bidirectional inter prediction.
[0048] Figure 9 A video bitstream is shown.
[0049] Figure 10 An image segmentation technique is shown.
[0050] Figure 11 An embodiment of a method for decoding a decoded video bitstream.
[0051] Figure 12 An embodiment of a method for encoding a decoded video bitstream.
[0052] Figure 13 A schematic diagram of a video decoding device.
[0053] Figure 14 A schematic diagram of an embodiment of a decoding module. Detailed Description
[0054] First, it should be understood that although illustrative implementations of one or more embodiments are provided below, the systems and / or methods disclosed by the present invention can be implemented using any number of techniques, whether currently known or existing. The present invention should in no way be limited to the illustrative implementations, drawings, and techniques described below, including the exemplary designs and implementations illustrated and described herein, but can be modified within the full scope of the appended claims and their equivalents.
[0055] The definitions of the following terms are as described below, unless used in the opposite context herein. Specifically, the following definitions are intended to more clearly describe the present invention. However, the terms may have different descriptions in different contexts. Therefore, the following definitions should be regarded as supplementary information and should not be regarded as limiting any other definitions provided for these terms herein.
[0056] A bitstream is a series of bits that includes video data, which is compressed for transmission between an encoder and a decoder. An encoder is a device that compresses video data into a bitstream using an encoding process. A decoder is a device that reconstructs video data from a bitstream for display using a decoding process. An image is an array composed of luminance samples and / or chrominance samples that create a frame or its fields. For the sake of clear discussion, the image being encoded or decoded can be referred to as the current image.
[0057] A reference picture is an image that includes reference samples that can be used when decoding other images by reference according to inter - frame prediction and / or inter - layer prediction. A reference picture list is a list of reference pictures used for inter - frame prediction and / or inter - layer prediction. Some video decoding systems use two reference picture lists, which can be denoted as reference picture list 1 and reference picture list 0. A reference picture list structure is an addressable syntax structure that includes multiple reference picture lists. Inter - frame prediction is a mechanism for decoding samples of a current image by referring to indicated samples in a reference picture that is different from the current image in the same layer as the reference picture and the current image. A reference picture list structure entry is an addressable location in the reference picture list structure that represents a reference picture related to a reference picture list.
[0058] A slice header is a part of a decoded slice that includes data elements related to all the video data within a block represented in the slice. A picture parameter set (PPS) is a parameter set that includes data related to an entire image. More specifically, a PPS is a syntax structure that includes syntax elements applicable to zero or more complete decoded images, determined by syntax elements in each picture header. A sequence parameter set (SPS) is a parameter set that includes data related to an image sequence. An access unit (AU) is a collection of one or more decoded images associated with the same display time (e.g., the same picture order number), for output from a decoded picture buffer (DPB) (e.g., for display to a user). A decoded video sequence is a series of images that have been reconstructed by a decoder in preparation for display to a user.
[0059] The compliance cropping window (or simply the compliance window) refers to a window composed of samples of an image in a decoded video sequence output from the decoding process. The bitstream can provide compliance window cropping parameters to indicate the output area of the decoded image. The image width is the width of the image measured in luminance samples. The image height is the height of the image measured in luminance samples. According to the rectangular area specified in the image coordinates for output, the compliance window offsets (e.g., conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset) represent the samples of the image of the reference PPS output from the decoding process.
[0060] Decoder-Side Motion Vector Refinement (DMVR) is a process, algorithm, or decoding tool for refining the motion or motion vectors of a predicted block. DMVR can find a motion vector based on two motion vectors found for bidirectional prediction using a bilateral template matching process. In DMVR, a weighted combination of the predicted coding units generated using each of the two motion vectors can be found, and the two motion vectors can be refined by replacing them with a new motion vector that best points to the combined predicted coding unit. Bi-directional optical flow (BDOF) is a process, algorithm, or decoding tool for refining the motion or motion vectors of a predicted block. BDOF can find a motion vector for a sub-coding unit based on the gradient of the difference between two reference images.
[0061] The reference picture resampling (RPR) feature is the ability to change the spatial resolution of a decoded image in the middle of the bitstream without the need for intra-coding the image at the resolution change location. The resolution used herein describes the number of pixels in a video file. That is, the resolution is the width and height of the projected image measured in pixels. For example, the resolution of a video may be 1280 (horizontal pixels) × 720 (vertical pixels). This is usually abbreviated as 1280×720, or simply 720p.
[0062] Decoder-Side Motion Vector Refinement (DMVR) is a process, algorithm, or decoding tool for refining the motion or motion vectors of predicted blocks. Bi-directional optical flow (BDOF), also known as bi-directional optical flow (BIO), is a process, algorithm, or decoding tool for refining the motion or motion vectors of predicted blocks. The reference picture resampling (RPR) feature is the ability to change the spatial resolution of a decoded image in the middle of a bitstream without intra-decoding the image at the resolution change positions.
[0063] The following abbreviations are used in this document: Coding Tree Block (CTB), Coding Tree Unit (CTU), Coding Unit (CU), Coded Video Sequence (CVS), Joint Video Experts Team (JVET), Motion-Constrained Tile Set (MCTS), Maximum Transfer Unit (MTU), Network Abstraction Layer (NAL), Picture Order Count (POC), Raw Byte Sequence Payload (RBSP), Sequence Parameter Set (SPS), Versatile Video Coding (VVC), and Working Draft (WD).
[0064] Figure 1 Flowchart of an exemplary operating method 100 for decoding a video signal. Specifically, the video signal is encoded on the encoder side. The encoding process compresses the video signal by using various mechanisms, thereby reducing the video file. The smaller file size helps to transmit the compressed video file to the user while reducing the associated bandwidth overhead. Then, the decoder decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process typically works in the same way as the encoding process to help the decoder reconstruct the video signal in the same manner.
[0065] In step 101, a video signal is input into an encoder. For example, the video signal can be an uncompressed video file stored in a memory. Alternatively, the video file can be captured by a video capture device (e.g., a camera) and encoded to support real-time streaming of the video. The video file can include both an audio component and a video component. The video component includes a series of image frames that, when viewed in sequence, produce the visual effect of motion. These frames include pixels represented by light (referred to herein as the luminance component (or luminance samples)) and color (referred to as the chrominance component (or color samples)). In some examples, the frames can also include depth values to support three-dimensional viewing.
[0066] In step 103, the video is segmented into blocks. The segmentation includes subdividing the pixels in each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC) (also known as H.265 and MPEG-H Part 2), the frame can first be divided into coding tree units (CTUs), which are blocks of a predefined size (e.g., 64 pixels × 64 pixels). A CTU includes luminance samples and chrominance samples. The CTU can be divided into blocks using a coding tree, and these blocks can then be recursively subdivided until a configuration structure that supports further encoding is obtained. For example, the luminance component of the frame can be subdivided until the individual blocks include relatively uniform lighting values. Additionally, the chrominance component of the frame can be subdivided until the individual blocks include relatively uniform color values. Thus, the segmentation mechanism varies depending on the content of the video frame.
[0067] In step 105, various compression mechanisms are used to compress the image blocks segmented in step 103. For example, inter-frame prediction and / or intra-frame prediction can be used. Inter-frame prediction aims to take advantage of the fact that objects in a common scene tend to appear in consecutive frames. Thus, there is no need to repeatedly describe the blocks depicting an object in adjacent frames. An object (e.g., a table) can remain in a constant position across multiple frames. Therefore, the table is described only once, and adjacent frames can refer back to the reference frame. A pattern matching mechanism can be used to match objects across multiple frames. Additionally, due to reasons such as object movement or camera movement, a moving object can be represented across multiple frames. In a specific example, a video can show a car moving across the screen over multiple frames. Motion vectors can be used to describe such movement. A motion vector is a two-dimensional vector that provides the offset between the coordinates of an object in one frame and the coordinates of that object in a reference frame. Thus, inter-frame prediction can encode the image blocks in the current frame as a set of motion vectors representing the offset between the image blocks in the current frame and the corresponding blocks in the reference frame.
[0068] Intra prediction encodes blocks in a common frame. Intra prediction exploits the fact that luminance and chrominance components tend to cluster in a frame. For example, a patch of green in a part of a tree tends to be adjacent to several similar patches of green. Intra prediction uses a variety of directional prediction modes (e.g., 33 modes in HEVC), a planar mode, and a direct current (DC) mode. The directional modes indicate that the samples of the current block are similar / same as the samples of the adjacent blocks in the corresponding directions. The planar mode indicates that a series of blocks on a row / column (e.g., a plane) can be interpolated based on the adjacent blocks of the row edge. In fact, the planar mode represents a smooth transition of luminance / color between rows / columns by using a relatively constant slope in the changing values. The DC mode is used for boundary smoothing and indicates that the average value of the samples of all adjacent blocks related to the block and the angular direction of the directional prediction mode is similar / same. Thus, an intra prediction block can represent an image block as various relationship prediction mode values instead of actual values. Additionally, an inter prediction block can represent an image block as motion vector values instead of actual values. In both cases, the prediction block may not fully represent the image block in some cases. Any differences are stored in the residual block. The residual block can be transformed to further compress the file.
[0069] In step 107, various filtering techniques can be applied. In HEVC, filters are applied according to an in-loop filtering scheme. The block-based prediction discussed above can create a blocky image in the decoder. Additionally, the block-based prediction scheme can encode blocks and then reconstruct the encoded blocks for later use as reference blocks. The in-loop filtering scheme iteratively applies a noise suppression filter, a deblocking filter, an adaptive loop filter, and a sample adaptive offset (SAO) filter to the blocks / frames. These filters reduce these block artifacts so that the encoded file can be accurately reconstructed. Additionally, these filters reduce the reconstructed reference block artifacts, making it less likely that other artifacts will be generated in subsequent blocks encoded based on the reconstructed reference blocks.
[0070] In step 109, once the video signal is segmented, compressed, and filtered, the resulting data is encoded into a bitstream. The bitstream includes the above data and any indication data desired to support proper video signal reconstruction in the decoder. For example, this data can include segmentation data, prediction data, residual blocks, and various flags that provide decoding instructions to the decoder. The bitstream can be stored in a memory and is used to be sent to the decoder upon request. The bitstream can also be broadcast and / or multicast to multiple decoders. The creation of the bitstream is an iterative process. Thus, steps 101, 103, 105, 107, and 109 can occur continuously and / or simultaneously on multiple frames and blocks. Figure 1The order shown is presented for clarity and ease of discussion and is not intended to limit the video decoding process to a particular order.
[0071] In step 111, the decoder receives the bitstream and begins the decoding process. Specifically, the decoder uses an entropy decoding scheme to convert the bitstream into corresponding syntax and video data. In step 111, the decoder uses the syntax data in the bitstream to determine the partitioning of the frame. The partitioning should match the result of the block partitioning in step 103. Now describe the entropy encoding / decoding used in step 111. The encoder makes many choices during the compression process, such as choosing a block partitioning scheme from multiple possible options based on the spatial location of the values in the input image. Indicating the exact option can use a large number of binary bits. The binary bits used herein are binary values regarded as variables (e.g., bit values that can vary according to the context). Entropy encoding helps the encoder discard any options that are clearly not suitable for a particular situation, leaving a set of available options. Then, a codeword is assigned to each available option. The length of the codeword depends on the number of available options (e.g., one binary bit corresponds to two options, and two binary bits correspond to three to four options). Then, the encoder encodes the codewords of the selected options. This scheme reduces the size of the codewords because the size of the codewords is as large as desired to uniquely indicate one option in a small subset of available options, rather than uniquely indicating an option in a possibly large set of all possible options. Then, the decoder decodes the options by determining the set of available options in a manner similar to the encoder. By determining the set of available options, the decoder can read the codewords and determine the choices made by the encoder.
[0072] In step 113, the decoder performs block decoding. Specifically, the decoder performs an inverse transform to generate residual blocks. Then, the decoder uses the residual blocks and the corresponding prediction blocks to reconstruct the image blocks according to the partitioning. The prediction blocks may include intra-prediction blocks and inter-prediction blocks generated by the encoder in step 105. Then, the reconstructed image blocks are placed in the frame of the reconstructed video signal according to the partitioning data determined in step 111. The syntax of step 113 can also be indicated in the bitstream by the entropy encoding discussed above.
[0073] In step 115, the frame of the reconstructed video signal is filtered in a manner similar to that of the encoder in step 107. For example, a noise suppression filter, a deblocking filter, an adaptive loop filter, and an SAO filter can be used on the frame to eliminate block artifacts. Once the frame is filtered, the video signal can be output to a display for viewing by the end user in step 117.
[0074] Figure 2FIG. 0 is a schematic diagram of an exemplary encoding and decoding (codec) system 200 for video decoding. Specifically, the codec system 200 is capable of implementing the operation method 100. Broadly, the codec system 200 is used to describe the components used in an encoder and a decoder. As discussed with respect to steps 101 and 103 in the operation method 100, the codec system 200 receives a video signal and segments the video signal to generate a segmented video signal 201. Then, when acting as an encoder, the codec system 200 compresses the segmented video signal 201 into an encoded bitstream, as discussed with respect to steps 105, 107, and 109 in the method 100. When acting as a decoder, the codec system 200 generates an output video signal from the bitstream, as described in connection with steps 111, 113, 115, and 117 in the operation method 100. The codec system 200 includes a general decoder control component 211, a transform scaling and quantization component 213, an intra prediction component 215, an intra prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded image buffer component 223, a header format and context adaptive binary arithmetic coding (CABAC) component 231. These components are coupled as shown. In Figure 2 FIG. 1, the black lines represent the motion of the data to be encoded / decoded, and the dashed lines represent the motion of the control data that controls the operation of other components. All components in the codec system 200 can be present in the encoder. The decoder may include a subset of the components in the codec system 200. For example, the decoder may include an intra prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded image buffer component 223. These components will now be described.
[0075] The segmented video signal 201 is a captured video sequence that has been segmented into pixel blocks by an encoding tree. The encoding tree uses various partitioning modes to subdivide the pixel blocks into smaller pixel blocks. These blocks can then be further subdivided into even smaller blocks. The blocks can be referred to as nodes on the encoding tree. A larger parent node is divided into smaller child nodes. The number of times a node is subdivided is referred to as the depth of the node / encoding tree. In some cases, the partitioned blocks can be included in a coding unit (CU). For example, a CU can be a subpart of a CTU, including a luminance block, a chrominance red difference (Cr) block, a chrominance blue difference (Cb) block, and the corresponding syntax instructions for the CU. The partitioning modes can include a binary tree (BT), a triple tree (TT), and a quad tree (QT), which are used to divide a node into two, three, or four child nodes with different shapes, respectively, depending on the partitioning mode used. The segmented video signal 201 is forwarded to a general decoder control component 211, a transform scaling and quantization component 213, an intra prediction component 215, a filter control analysis component 227, and a motion estimation component 221 for compression.
[0076] The general decoder control component 211 is used to make decisions related to encoding the images of a video sequence into a bitstream according to application constraints. For example, the general decoder control component 211 manages the optimization of the bitrate / bitstream size with respect to the reconstructed quality. These decisions can be made based on the storage space / bandwidth availability and the image resolution request. The general decoder control component 211 also manages the utilization of the buffer according to the transmission speed to alleviate the problems of buffer underflow and overflow. To manage these problems, the general decoder control component 211 manages the segmentation, prediction, and filtering performed by other components. For example, the general decoder control component 211 can dynamically increase the compression complexity to increase the resolution and bandwidth utilization, or reduce the compression complexity to reduce the resolution and bandwidth utilization. Therefore, the general decoder control component 211 controls other components of the codec system 200 to balance the video signal reconstruction quality and the bitrate issue. The general decoder control component 211 creates control data to control the operations of other components. The control data is also forwarded to a header format and CABAC component 231 for encoding in the bitstream to indicate the parameters decoded in the decoder.
[0077] The segmented video signal 201 is also sent to a motion estimation component 221 and a motion compensation component 219 for inter prediction. The frames or stripes of the segmented video signal 201 can be divided into a plurality of video blocks. The motion estimation component 221 and the motion compensation component 219 perform inter prediction decoding on the received video blocks based on one or more blocks in one or more reference frames to provide temporal prediction. The codec system 200 can perform multiple decoding processes to select an appropriate decoding mode for each video data block, and so on.
[0078] The motion estimation component 221 and the motion compensation component 219 can be highly integrated but are described separately for conceptual purposes. The motion estimation performed by the motion estimation component 221 is a process of generating motion vectors, which are used to estimate the motion of video blocks. For example, a motion vector can indicate the displacement of an encoded object relative to a prediction block. A prediction block is a block that is found to closely match the block to be encoded in terms of pixel differences. A prediction block can also be referred to as a reference block. Such pixel differences can be determined by sum of absolute difference (SAD), sum of square difference (SSD), or other difference metrics. HEVC uses several encoded objects, including CTUs, coding tree blocks (CTBs), and CUs. For example, a CTU can be divided into multiple CTBs, and then the CTBs can be divided into multiple CUs including CUs. A CU can be encoded as a prediction unit (PU) including prediction data and / or a transform unit (TU) including transform residual data of the CU. The motion estimation component 221 uses rate-distortion analysis as part of a rate-distortion optimization process to generate motion vectors, PUs, and TUs. For example, the motion estimation component 221 can determine multiple reference blocks, multiple motion vectors, etc. of the current block / frame, and can select the reference blocks, motion vectors, etc. with the best rate-distortion characteristics. The best rate-distortion characteristics balance the quality of video reconstruction (e.g., the amount of data loss caused by compression) and decoding efficiency (e.g., the final encoded size).
[0079] In some examples, the codec system 200 can calculate the values of sub-integer pixel positions of the reference images stored in the decoded image buffer component 223. For example, the video codec system 200 can interpolate the values of quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference images. Therefore, the motion estimation component 221 can perform motion searches regarding integer pixel positions and fractional pixel positions and output motion vectors with fractional pixel accuracy. The motion estimation component 221 calculates the motion vectors of the PUs in the inter-frame encoded strips of video blocks by comparing the positions of the PUs with the positions of the prediction blocks of the reference images. The motion estimation component 221 outputs the calculated motion vectors as motion data to the header format and CABAC component 231 for encoding and outputs the motion to the motion compensation component 219.
[0080] The motion compensation performed by the motion compensation component 219 may involve obtaining or generating a prediction block according to the motion vector determined by the motion estimation component 221. Similarly, in some examples, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated. After receiving the motion vector of the PU of the current video block, the motion compensation component 219 may locate the prediction block pointed to by the motion vector. Then, by subtracting the pixel values of the prediction block from the pixel values of the currently encoded current video block, a pixel difference is generated, thereby forming a residual video block. Generally, the motion estimation component 221 performs motion estimation on the luminance component, and the motion compensation component 219 uses the motion vector calculated according to the luminance component for both the chrominance component and the luminance component. The prediction block and the residual block are forwarded to the transform scaling and quantization component 213.
[0081] The split video signal 201 is also sent to the intra estimation component 215 and the intra prediction component 217. Similar to the motion estimation component 221 and the motion compensation component 219, the intra estimation component 215 and the intra prediction component 217 may be highly integrated, but are separately described for conceptual purposes. The intra estimation component 215 and the intra prediction component 217 perform intra prediction on the current block according to the blocks in the current frame to replace the inter prediction performed between frames by the motion estimation component 221 and the motion compensation component 219 as described above. Specifically, the intra estimation component 215 determines the intra prediction mode for encoding the current block. In some examples, the intra estimation component 215 selects an appropriate intra prediction mode from a plurality of tested intra prediction modes to encode the current block. Then, the selected intra prediction mode is forwarded to the header format and CABAC component 231 for encoding.
[0082] For example, the intra estimation component 215 uses rate-distortion analysis of various tested intra prediction modes to calculate rate-distortion values and selects the intra prediction mode with the best rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block encoded to generate the encoded block and the code rate (e.g., the number of bits) used to generate the encoded block. The intra estimation component 215 calculates a ratio according to the distortion and rate of various encoded blocks and determines which intra prediction mode gives the best rate-distortion value for the block. In addition, the intra estimation component 215 can be used to encode depth blocks of the depth map using the depth modeling mode (DMM) according to rate-distortion optimization (RDO).
[0083] When implemented on the encoder, the intra prediction component 217 may generate a residual block from a prediction block according to the selected intra prediction mode determined by the intra estimation component 215, or when implemented on the decoder, read the residual block from the bitstream. The residual block includes the value difference between the prediction block and the original block, represented as a matrix. Then, the residual block is forwarded to the transform scaling and quantization component 213. The intra estimation component 215 and the intra prediction component 217 may operate on the luminance component and the chrominance component.
[0084] The transform scaling and quantization component 213 is used to further compress the residual block. The transform scaling and quantization component 213 applies a transform such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform to the residual block, generating a video block including residual transform coefficient values. Wavelet transforms, integer transforms, subband transforms, or other types of transforms may also be used. The transform may transform the residual information from the pixel value domain to the transform domain, such as the frequency domain. The transform scaling and quantization component 213 is also used to scale the transform residual information according to frequency, etc. Such scaling involves applying a scaling factor to the residual information in order to quantify different frequency information at different granularities, which can affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 is also used to quantize the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting the quantization parameter. In some examples, the transform scaling and quantization component 213 may then scan the matrix including the quantized transform coefficients. The quantized transform coefficients are forwarded to the header format and CABAC component 231 for encoding into the bitstream.
[0085] The scaling and inverse transform component 229 performs the inverse operations of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 performs inverse scaling, inverse transform, and / or inverse quantization to reconstruct the residual block in the pixel domain, e.g., for subsequent use as a reference block, which may become the prediction block for another current block. The motion estimation component 221 and / or the motion compensation component 219 may calculate the reference block by adding the residual block to the corresponding prediction block for use in motion estimation of subsequent blocks / frames. A filter is applied to the reconstructed reference block to reduce the artifacts generated during the scaling, quantization, and transform processes. These artifacts may produce inaccurate predictions (and generate other artifacts) when predicting subsequent blocks.
[0086] The filter control analysis component 227 and the in-loop filter component 225 apply filters to the residual blocks and / or the reconstructed image blocks. For example, the transformed residual blocks in the scaling and inverse transform component 229 can be combined with the corresponding prediction blocks in the intra prediction component 217 and / or the motion compensation component 219 to reconstruct the original image blocks. Then, filters can be applied to the reconstructed image blocks. In some examples, filters can be applied to the residual blocks. Similar to Figure 2 the other components in Figure 2 , the filter control analysis component 227 and the in-loop filter component 225 are highly integrated and can be implemented together, but are described separately for conceptual purposes. The filters applied to the reconstructed reference blocks are applied to specific spatial regions and include multiple parameters to adjust how these filters are applied. The filter control analysis component 227 analyzes the reconstructed reference blocks to determine where these filters should be applied and sets the corresponding parameters. This data is forwarded as filter control data to the header format and CABAC component 231 for encoding. The in-loop filter component 225 applies these filters according to the filter control data. The filters can include deblocking filters, noise suppression filters, SAO filters, and adaptive loop filters. These filters can be applied, according to examples, in the spatial / pixel domain (e.g., on the reconstructed pixel blocks) or in the frequency domain.
[0087] When operating as an encoder, the filtered reconstructed image blocks, residual blocks, and / or prediction blocks are stored in the decoded image buffer component 223 for later motion estimation as described above. When operating as a decoder, the decoded image buffer component 223 stores the reconstructed blocks and the filtered blocks and forwards the reconstructed blocks and the filtered blocks to the display as part of the output video signal. The decoded image buffer component 223 can be any memory device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks.
[0088] The header format and CABAC component 231 receive data from various components of the codec system 200 and encode this data into an encoded bitstream for transmission to a decoder. Specifically, the header format and CABAC component 231 generate various headers to encode control data such as overall control data and filter control data. In addition, prediction data including intra prediction and motion data, as well as residual data in the form of quantized transform coefficient data, are encoded into the bitstream. The final bitstream includes all the information that the decoder needs to reconstruct the original segmented video signal 201. This information may also include an intra prediction mode index table (also referred to as a codeword mapping table), definitions of the coding contexts of various blocks, indications of the most likely intra prediction modes, indications of segmentation information, etc. This data can be encoded by entropy coding techniques. For example, the information can be encoded by using context adaptive variable length coding (CAVLC), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) decoding, or other entropy decoding techniques. After entropy decoding, the decoded bitstream can be sent to another device (e.g., a video decoder) or archived for later transmission or retrieval.
[0089] Figure 3 FIG. Exemplary block diagram of a video encoder 300. The video encoder 300 can be used to implement the encoding function of the codec system 200 and / or implement steps 101, 103, 105, 107, and / or 109 of the operation method 100. The encoder 300 segments the input video signal to produce a segmented video signal 301 that is substantially similar to the segmented video signal 201. Then, the segmented video signal 301 is compressed and encoded into a bitstream by the components of the encoder 300.
[0090] Specifically, the segmented video signal 301 is forwarded to the intra prediction component 317 for intra prediction. The intra prediction component 317 can be substantially similar to the intra estimation component 215 and the intra prediction component 217. The segmented video signal 301 is also forwarded to the motion compensation component 321 for inter prediction based on the reference blocks in the decoded picture buffer 323. The motion compensation component 321 can be substantially similar to the motion estimation component 221 and the motion compensation component 219. The prediction blocks and residual blocks in the intra prediction component 317 and the motion compensation component 321 are forwarded to the transform and quantization component 313 to perform transform and quantization on the residual blocks. The transform and quantization component 313 can be substantially similar to the transform scaling and quantization component 213. The transformed and quantized residual blocks and the corresponding prediction blocks (and related control data) are forwarded to the entropy coding component 331 to be encoded into the bitstream. The entropy coding component 331 can be substantially similar to the header format and CABAC component 231.
[0091] The transformed and quantized residual blocks and / or the corresponding prediction blocks are also forwarded from the transform and quantization component 313 to the inverse transform and quantization component 329 to be reconstructed as reference blocks for use by the motion compensation component 321. The inverse transform and quantization component 329 can be substantially similar to the scaling and inverse transform component 229. According to an example, the in-loop filter in the in-loop filter component 325 is also applied to the residual blocks and / or the reconstructed reference blocks. The in-loop filter component 325 can be substantially similar to the filter control analysis component 227 and the in-loop filter component 225. As discussed with respect to the in-loop filter component 225, the in-loop filter component 325 can include multiple filters. Then, the filtered blocks are stored in the decoded picture buffer component 323 for use as reference blocks by the motion compensation component 321. The decoded picture buffer component 323 can be substantially similar to the decoded picture buffer component 223.
[0092] Figure 4 It is a block diagram of an exemplary video decoder 400. The video decoder 400 can be used to implement the decoding function of the codec system 200 and / or implement steps 111, 113, 115, and / or 117 of the operation method 100. For example, the decoder 400 receives a bitstream from the encoder 300 and generates a reconstructed output video signal according to the bitstream for display to the end user.
[0093] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 is used to implement an entropy decoding scheme, such as CAVLC, CABAC, SBAC, PIPE decoding, or other entropy decoding techniques. For example, the entropy decoding component 433 can use header information to provide context for interpreting other data encoded as codewords in the bitstream. The decoded information includes any information required to decode the video signal, such as overall control data, filter control data, segmentation information, motion data, prediction data, and quantization transform coefficients of residual blocks. The quantization transform coefficients are forwarded to the inverse transform and quantization component 429 for reconstruction into residual blocks. The inverse transform and quantization component 429 can be substantially similar to the inverse transform and quantization component 329.
[0094] The reconstructed residual block and / or prediction block is forwarded to the intra prediction component 417 for reconstruction into an image block according to the intra prediction operation. The intra prediction component 417 can be similar to the intra estimation component 215 and the intra prediction component 217. Specifically, the intra prediction component 417 uses a prediction mode to locate a reference block in the frame and applies the residual block to the result to reconstruct the intra prediction image block. The reconstructed intra prediction image block and / or residual block, along with the corresponding inter prediction data, are forwarded to the in-loop filter component 425 and then to the decoded image buffer component 423. The decoded image buffer component 423 and the in-loop filter component 425 can be substantially similar to the decoded image buffer component 223 and the in-loop filter component 225, respectively. The in-loop filter component 425 filters the reconstructed image block, residual block, and / or prediction block, and this information is stored in the decoded image buffer component 423. The reconstructed image block in the decoded image buffer component 423 is forwarded to the motion compensation component 421 for inter prediction. The motion compensation component 421 can be substantially similar to the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 uses the motion vector in the reference block to generate a prediction block and applies the residual block to the result to reconstruct the image block. The resulting reconstructed block can also be forwarded to the decoded image buffer component 423 through the in-loop filter component 425. The decoded image buffer component 423 continues to store other reconstructed image blocks, which can be reconstructed into frames through segmentation information. These frames can also be arranged in sequence. This sequence is output as the reconstructed output video signal to the display screen.
[0095] Remember that video compression techniques perform spatial (intra-frame) prediction and / or temporal (inter-frame) prediction to reduce or remove the redundancy inherent in a video sequence. For block-based video coding, a video slice (e.g., a video picture or a portion of a video picture) may be partitioned into video blocks, which may also be referred to as tree blocks, coding tree blocks (CTBs), coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of an image are coded using spatial prediction with reference samples in adjacent blocks in the same image. Video blocks in an inter-coded (P or B) slice of an image may use spatial prediction with reference samples in adjacent blocks in the same image or temporal prediction with reference samples in other reference images. An image may be referred to as a frame, and a reference image may be referred to as a reference frame.
[0096] Spatial or temporal prediction produces a prediction block for the block to be coded. Residual data represents the pixel difference between the original block to be coded and the prediction block. An inter-coded block is coded according to a motion vector pointing to a block of reference samples that form the prediction block and residual data indicating the difference between the coded block and the prediction block. An intra-coded block is coded according to an intra-coding mode and residual data. For further compression, the residual data may be transformed from the pixel domain to a transform domain, thereby producing residual transform coefficients, which may then be quantized. The quantized transform coefficients, initially arranged in a two-dimensional array, may be scanned to produce a one-dimensional transform coefficient vector, and entropy coding may be applied to achieve greater compression.
[0097] Image and video compression has developed rapidly, and coding standards have diversified. These video coding standards include ITU-T H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) MPEG-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, Advanced Video Coding (AVC) (also known as ITU-T H.264 or ISO / IEC MPEG-4 Part 10), and High Efficiency Video Coding (HEVC) (also known as ITU-T H.265 or MPEG-H Part 2). AVC includes extended versions such as Scalable Video Coding (SVC), Multiview Video Coding (MVC), Multiview Video Coding plus Depth (MVC+D), and 3D AVC (3D-AVC). HEVC includes extended versions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC).
[0098] There is also a new video coding standard called Versatile Video Coding (VVC) being developed by the joint video experts team (JVET) of ITU-T and ISO / IEC. Although there are several working drafts of the VVC standard, one working draft (WD) of VVC, namely JVET-N1001-v3 "Versatile Video Coding (Draft 5)" proposed by B. Bross, J. Chen, and S. Liu at the 13th JVET meeting on March 27, 2019 (VVC Draft 5), is cited in this article. Each reference in this paragraph and the previous paragraph is incorporated by full reference.
[0099] The description of the technology disclosed herein is based on the Versatile Video Coding (VVC), a video coding standard being developed by the joint video experts team (JVET) of ITU-T and ISO / IEC. However, these technologies are also applicable to other video coding and decoding specifications.
[0100] Figure 5 Reference numeral 500 is a representation of the relationship of an intra random access point (IRAP) picture 502 with respect to a previous picture 504 and a subsequent picture 506 in decoding order 508 and presentation order 510. In one embodiment, the IRAP picture 502 is referred to as a clean random access (CRA) picture or an instantaneous decoder refresh (IDR) picture accompanying a random access decodable (RADL) picture. In HEVC, IDR pictures, CRA pictures, and Broken Link Access (BLA) pictures are all considered to be IRAP pictures 502. For VVC, it was agreed at the 12th JVET meeting in October 2018 that both IDR pictures and CRA pictures are to be considered IRAP pictures. In one embodiment, Broken Link Access (BLA) pictures and Gradual Decoder Refresh (GDR) pictures may also be considered IRAP pictures. The decoding process of a decoded video sequence always starts from an IRAP picture.
[0101] As Figure 5 shown, the previous pictures 504 (e.g., pictures 2 and 3) are after the IRAP picture 502 in decoding order 508, but before the IRAP picture 502 in presentation order 510. The subsequent picture 506 is after the IRAP picture 502 in both decoding order 508 and presentation order 510. Although Figure 5 two previous pictures 504 and one subsequent picture 506 are shown, those skilled in the art will understand that in a practical application, there may be more or fewer previous pictures 504 and / or subsequent pictures 506 in decoding order 508 and presentation order 510.
[0102] Figure 5The leading picture 504 therein is divided into two types, namely, the random access skipped leading (RASL) picture and the RADL picture. When decoding starts from the IRAP picture 502 (e.g., picture 1), the RADL picture (e.g., picture 3) can be correctly decoded; however, the RASL picture (e.g., picture 2) cannot be correctly decoded. Therefore, the RASL picture is discarded. Given the difference between the RADL picture and the RASL picture, the type of the leading picture 504 associated with the IRAP picture 502 can be identified as RADL or RASL to achieve efficient and correct decoding. In HEVC, when there are RASL pictures and RADL pictures, the constraints are as follows: for the RASL picture and the RADL picture associated with the same IRAP picture 502, the RASL picture can be before the RADL picture in the presentation order 510.
[0103] The IRAP picture 502 provides the following two important functions / advantages. First, the presence of the IRAP picture 502 indicates that the decoding process can start from this picture. This function has a random access feature, where the decoding process starts from a position in the bitstream, not necessarily from the start of the bitstream, provided that the IRAP picture 502 exists at this position. Second, the presence of the IRAP picture 502 refreshes the decoding process, such that the decoded pictures starting from the IRAP picture 502 (excluding the RASL pictures) are decoded without referring to the pictures before the IRAP picture. Therefore, the presence of the IRAP picture 502 in the bitstream can stop any error propagation that may occur during the decoding of the pictures before the IRAP picture 502 to the IRAP picture 502 and those pictures after the IRAP picture 502 in the decoding order 508.
[0104] Although the IRAP picture 502 provides important functions, these functions will reduce the compression efficiency. The presence of the IRAP picture 502 will cause a significant increase in the bit rate. This reduction in compression efficiency is due to two reasons. First, since the IRAP picture 502 is an intra-predicted picture, when compared with other pictures that are inter-predicted pictures (e.g., the leading picture 504, the trailing picture 506), the picture itself will require relatively more bits to represent. Second, since the presence of the IRAP picture 502 interrupts the temporal prediction (this is because the decoder refreshes the decoding process, and one of the actions of the decoding process is to remove the previous reference pictures in the decoded picture buffer (DPB)), the IRAP picture 502 results in lower decoding efficiency for the pictures after the IRAP picture 502 in the decoding order 508 (i.e., requires more bits to represent), because these pictures have fewer reference pictures for their inter-predicted decoding.
[0105] Among the image types considered to be the IRAP image 502, the IDR image in HEVC has different indication and derivation methods compared to other image types. Some of these differences are as follows.
[0106] For the indication and derivation of the picture order count (POC) value of the IDR image, the most significant bit (MSB) part of the POC is not derived based on the previous key image, but is simply set to 0.
[0107] For the indication information required for reference image management, the slice header of the IDR image does not include the information that needs to be indicated to assist in reference image management. For other image types (i.e., CRA, post, temporal sub-layer access (TSA), etc.), the reference image identification process (i.e., the process of determining the status (for reference or not for reference) of the reference images in the decoded picture buffer (DPB)) requires information such as the reference picture set (RPS) described below or other forms of similar information (e.g., reference image lists). However, for the IDR image, since the presence of the IDR image indicates that the decoding process can mark all the reference images in the DPB as not for reference, there is no need to signal this information.
[0108] In addition to the IRAP image concept, there are leading images. If a leading image exists, it is associated with the IRAP image. A leading image is an image whose decoding order is after its associated IRAP image but whose output order is before the IRAP image. According to the decoding configuration and the picture reference structure, leading images are further divided into two types. The first type is a leading image that may not be correctly decoded when the decoding process starts from the associated IRAP image. This occurs because these leading images are decoded with reference to pictures whose decoding order is before the IRAP image. These leading images are called random access skipped leading (RASL) images. The second type is a leading image that can be correctly decoded even when the decoding process starts from the associated IRAP image. This may be because these leading images are decoded without directly or indirectly referring to pictures whose decoding order is before the IRAP image. These leading images are called random access decodable leading (RADL) images. In HEVC, when both RASL images and RADL images exist, the constraints are as follows: for RASL images and RADL images associated with the same IRAP image, the RASL images can be output before the RADL images in output order.
[0109] In HEVC and VVC, both the IRAP picture 502 and the leading picture 504 can be included in a single network abstraction layer (NAL) unit. A set of NAL units can be called an access unit. The IRAP picture 502 and the leading picture 504 are given different NAL unit types so that these pictures can be easily recognized by system-level applications. For example, a video splicer needs to understand the decoded picture type without needing to understand too many details of the syntax elements in the decoded bitstream. In particular, it needs to identify the IRAP picture 502 from non-IRAP pictures and identify the leading picture 504 from the trailing pictures 506, including determining RASL pictures and RADL pictures. The trailing pictures 506 are those pictures that are associated with the IRAP picture 502 and are after the IRAP picture 502 in presentation order 510. The pictures can be after a particular IRAP picture 502 in decoding order 508 and before any other IRAP picture 502 in decoding order 508. To this end, giving the IRAP picture 502 and the leading picture 504 their own NAL unit types is helpful for such applications.
[0110] For HEVC, the NAL unit types of IRAP pictures include:
[0111] BLA picture (BLA_W_LP) accompanied by a leading picture: The NAL unit of a Broken Link Access (BLA) picture that can be before one or more leading pictures in decoding order.
[0112] BLA picture (BLA_W_RADL) accompanied by a RADL picture: The NAL unit of a BLA picture that can be before one or more RADL pictures in decoding order but is not accompanied by a RASL picture.
[0113] BLA picture (BLA_N_LP) not accompanied by a leading picture: The NAL unit of a BLA picture that is not before a leading picture in decoding order.
[0114] IDR picture (IDR_W_RADL) accompanied by a RADL picture: The NAL unit of an IDR picture that can be before one or more RADL pictures in decoding order but is not accompanied by a RASL picture.
[0115] IDR picture (IDR_N_LP) not accompanied by a leading picture: The NAL unit of an IDR picture that is not before a leading picture in decoding order.
[0116] CRA: The NAL unit of a Clean Random Access (CRA) picture that can be before a leading picture (i.e., a RASL picture or a RADL picture or both).
[0117] RADL: The NAL unit of a RADL picture.
[0118] RASL: The NAL unit of a RASL picture.
[0119] For VVC, the NAL unit types of IRAP picture 502 and leading picture 504 are as follows:
[0120] IDR picture (IDR_W_RADL) accompanied by a RADL picture: The NAL unit of an IDR picture that can be before one or more RADL pictures in decoding order but is not accompanied by a RASL picture.
[0121] IDR picture (IDR_N_LP) not accompanied by a leading picture: The NAL unit of an IDR picture that is not before a leading picture in decoding order.
[0122] CRA: The NAL unit of a Clean Random Access (CRA) picture that can be before a leading picture (i.e., a RASL picture or a RADL picture or both).
[0123] RADL: The NAL unit of a RADL picture.
[0124] RASL: NAL unit of RASL picture.
[0125] The reference picture resampling (RPR) feature is the ability to change the spatial resolution of a decoded picture in the middle of a bitstream without intra - decoding the picture at the resolution change positions. To implement this feature, the picture needs to be able to reference one or more reference pictures with a spatial resolution different from that of the current picture for inter - prediction. Therefore, it is necessary to resample such a reference picture or a part of it for encoding and decoding the current picture. Hence, it is called RPR. This feature can also be called adaptive resolution change (ARC) or other names. The RPR feature can be beneficial for some use cases or application scenarios as follows.
[0126] Rate adaptation in video telephony and conferencing. This is to adapt the decoded video to changing network conditions. When the network conditions deteriorate, resulting in a reduced available bandwidth, the encoder can adapt by encoding pictures at a lower resolution.
[0127] Active speaker change in multi - party video conferencing. For multi - party video conferencing, the video size of the active speaker is usually larger than that of other participants. When the active speaker changes, the image resolution of each participant may also need to be adjusted. When the active speaker changes frequently, the ARC feature is more needed.
[0128] Fast start in streaming. For streaming applications, the application usually buffers a certain length of decoded pictures before starting to display the pictures. Starting the bitstream at a lower resolution enables the application to have enough pictures in the buffer to start displaying faster.
[0129] Adaptive stream switching in streaming. The HTTP Dynamic Adaptive Streaming over HTTP (DASH) specification includes a feature called @mediaStreamStructureId. This feature enables switching between different representations at the random access points of an open group of pictures (GOP) with undecodable pre-images (e.g., CRA images accompanied by associated RASL images in HEVC). When the bitrates of two different representations of the same video are different but the spatial resolutions are the same, and their @mediaStreamStructureId values are the same, switching between the two representations can be performed at the CRA image accompanied by the associated RASL image, and the RASL image associated with the switch at the CRA image can be decoded with acceptable quality, thus enabling seamless switching. Using ARC, the @mediaStreamStructureId feature can also be used for switching between DASH representations with different spatial resolutions.
[0130] Various methods are beneficial for supporting the basic technologies of RPR / ARC, such as the indication of the image resolution list, some constraints on the resampling of reference images in the DPB, etc.
[0131] One component of the technology required to support RPR is a method for indicating the image resolutions that can exist in the bitstream. In some examples, this can be addressed by changing the current indication of the image resolution using the image resolution list in the SPS, as shown below.
[0132]
[0133] num_pic_size_in_luma_samples_minus1 + 1 represents the number of image sizes (width and height) in luma samples that can exist in the decoded video sequence.
[0134] pic_width_in_luma_samples[i] represents the i-th width in luma samples of the decoded image that can exist in the decoded video sequence. pic_width_in_luma_samples[i] cannot be equal to 0 and can be an integer multiple of MinCbSizeY.
[0135] pic_height_in_luma_samples[i] represents the i-th height in luma samples of a decoded picture that can exist in the decoded video sequence. pic_height_in_luma_samples[i] shall not be equal to 0 and shall be an integer multiple of MinCbSizeY.
[0136] Another technique for signaling the picture size and compliance window for RPR was discussed at the 15th JVET meeting. The signaling is as follows.
[0137] – Signal the maximum picture size (i.e., picture width and picture height) in the SPS
[0138] – Signal the picture size in the picture parameter set (PPS)
[0139] – Move the current signaling of the compliance window from the SPS to the PPS. The compliance window information is used to crop the reconstructed / decoded picture during the process of preparing the picture output. The cropped picture size is the picture size after cropping the picture using the compliance window associated with the picture.
[0140] The signaling of the picture size and compliance window is as follows.
[0141]
[0142] max_width_in_luma_samples represents the requirement for bitstream compliance that pic_width_in_luma_samples of any picture for which this SPS is active shall be less than or equal to max_width_in_luma_samples.
[0143] max_height_in_luma_samples represents the requirement for bitstream compliance that pic_height_in_luma_samples of any picture for which this SPS is active shall be less than or equal to max_height_in_luma_samples.
[0144]
[0145]
[0146] pic_width_in_luma_samples represents the width in luma samples of each decoded picture with reference to the PPS. pic_width_in_luma_samples shall not be equal to 0 and shall be an integer multiple of MinCbSizeY.
[0147] pic_height_in_luma_samples represents the height in luma samples of each decoded picture of the reference PPS. pic_height_in_luma_samples shall not be equal to 0 and shall be an integer multiple of MinCbSizeY.
[0148] The requirements for bitstream conformance are as follows: Each active reference picture with width and height of reference_pic_width_in_luma_samples and reference_pic_height_in_luma_samples shall satisfy all of the following conditions:
[0149] –2*pic_width_in_luma_samples >= reference_pic_width_in_luma_samples
[0150] –2*pic_height_in_luma_samples >= reference_pic_height_in_luma_samples
[0151] –pic_width_in_luma_samples <= 8*reference_pic_width_in_luma_samples
[0152] –pic_height_in_luma_samples <= 8*reference_pic_height_in_luma_samples
[0153] The variables PicWidthInCtbsY, PicHeightInCtbsY, PicSizeInCtbsY, PicWidthInMinCbsY, PicHeightInMinCbsY, PicSizeInMinCbsY, PicSizeInSamplesY, PicWidthInSamplesC, and PicHeightInSamplesC are derived as follows:
[0154] PicWidthInCtbsY = Ceil( pic_width_in_luma_samples ÷ CtbSizeY ) (1)
[0155] PicHeightInCtbsY = Ceil( pic_height_in_luma_samples ÷ CtbSizeY )(2)
[0156] PicSizeInCtbsY = PicWidthInCtbsY * PicHeightInCtbsY (3)
[0157] PicWidthInMinCbsY = pic_width_in_luma_samples / MinCbSizeY (4)
[0158] PicHeightInMinCbsY = pic_height_in_luma_samples / MinCbSizeY (5)
[0159] PicSizeInMinCbsY = PicWidthInMinCbsY * PicHeightInMinCbsY (6)
[0160] PicSizeInSamplesY = pic_width_in_luma_samples * pic_height_in_luma_samples (7)
[0161] PicWidthInSamplesC = pic_width_in_luma_samples / SubWidthC (8)
[0162] PicHeightInSamplesC = pic_height_in_luma_samples / SubHeightC (9)
[0163] Conformance_window_flag = 1 indicates that the conformance cropping window offset parameter is after the next parameter in the PPS. conformance_window_flag = 0 indicates that the conformance cropping window offset parameter does not exist.
[0164] According to the rectangular area specified for output in the image coordinates, conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset represent samples of the image of the reference PPS output from the decoding process. When conformance_window_flag = 0, the values of conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset are inferred to be 0.
[0165] The conformance cropping window includes luma samples with horizontal and vertical image coordinates, where the range of the horizontal image coordinates is from SubWidthC * conf_win_left_offset to pic_width_in_luma_samples – (SubWidthC * conf_win_right_offset + 1) (including the end values), and the range of the vertical image coordinates is from SubHeightC * conf_win_top_offset to pic_height_in_luma_samples – (SubHeightC * conf_win_bottom_offset + 1) (including the end values).
[0166] The value of SubWidthC * (conf_win_left_offset + conf_win_right_offset) can be less than pic_width_in_luma_samples, and the value of SubHeightC * (conf_win_top_offset + conf_win_bottom_offset) can be less than pic_height_in_luma_samples.
[0167] The variables PicOutputWidthL and PicOutputHeightL are derived as follows:
[0168] PicOutputWidthL = pic_width_in_luma_samples – (10) SubWidthC * (conf_win_right_offset + conf_win_left_offset)
[0169] PicOutputHeightL = pic_height_in_pic_size_units – (11) SubHeightC * (conf_win_bottom_offset + conf_win_top_offset)
[0170] When ChromaArrayType is not equal to 0, the corresponding specified samples of the two chroma arrays are the samples with image coordinates (x / SubWidthC, y / SubHeightC), where (x, y) are the image coordinates of the specified luma sample.
[0171] Note: The compliance cropping window offset parameters are only applied to the output. All internal decoding processes are applied to the uncropped image size.
[0172] The indication of the picture size and the compliance window in the PPS introduces the following problems.
[0173] – Since there can be multiple PPSs in a coded video sequence (CVS), two PPSs can include the same picture size indication but different compliance window indications. This will result in two pictures referring to different PPSs having the same picture size but different cropping sizes.
[0174] – To support RPR, when the current picture and the reference picture of a block have different picture sizes, it is recommended to turn off several decoding tools to decode the block. However, since now the cropping sizes can be different even if two pictures have the same picture size, it is necessary to perform additional checks based on the cropping size.
[0175] This document discloses a technique for constraining picture parameter sets with the same picture size to also have the same compliance window size (e.g., cropping window size). By keeping the sizes of the compliance windows of picture parameter sets with the same picture size the same, overly complex processing can be avoided when reference picture resampling (RPR) is enabled. Therefore, the utilization of processor, memory, and / or network resources on the encoder and decoder sides can be reduced. As a result, the encoder / decoder (also known as "codec") in video coding is improved compared to the current codec. In fact, the improved video coding process provides a better user experience when sending, receiving, and / or viewing videos.
[0176] The adaptability of video coding is typically supported by using multi-layer coding techniques. A multi-layer bitstream includes a base layer (BL) and one or more enhancement layers (EL). Examples of adaptability include spatial adaptability, quality / signal-to-noise (SNR) adaptability, multi-view adaptability, etc. When multi-layer coding techniques are used, an image or a part thereof can be coded in the following cases: (1) without using a reference image, i.e., using intra prediction, (2) using a reference image in the same layer, i.e., using inter prediction, or (3) using a reference image in one or more other layers, i.e., using inter-layer prediction. The reference image used for inter-layer prediction of the current image is called an inter-layer reference picture (ILRP).
[0177] Figure 6 A schematic diagram of an example of layer-based prediction 600. For example, layer-based prediction 600 is performed at block compression step 105, block decoding step 113, motion estimation component 221, motion compensation component 219, motion compensation component 321, and / or motion compensation component 421 to determine the MV. Layer-based prediction 600 coexists with uni-directional inter prediction and / or bi-directional inter prediction, but is also performed between images in different layers.
[0178] Layer-based prediction 600 is applied between images 611, 612, 613, and 614 in different layers and images 615, 616, 617, and 618. In the example shown, images 611, 612, 613, and 614 are part of layer N+1 632, and images 615, 616, 617, and 618 are part of layer N 631. Layers such as layer N 631 and / or layer N+1 632 are a set of images that are all associated with similar eigenvalue characteristics such as similar size, quality, resolution, signal-to-noise ratio, capabilities, etc. In the example shown, layer N+1 632 is associated with a larger image size compared to layer N 631. Thus, in this example, images 611, 612, 613, and 614 in layer N+1 632 are larger than images 615, 616, 617, and 618 in layer N 631 (e.g., greater in height and width and thus having more samples). However, these images can be divided into layer N+1 632 and layer N 631 by other characteristics. Although only two layers are shown: layer N+1 632 and layer N 631, a set of images can be divided into any number of layers according to the associated characteristics. Layer N+1 632 and layer N 631 can also be represented by layer IDs. A layer ID is a data item associated with an image and indicates that the image is part of the indicated layer. Thus, each of images 611 to 618 can be associated with a corresponding layer ID to indicate which layer in layer N+1 632 or layer N 631 includes the corresponding image.
[0179] Images 611 to 618 in different layers 631 and 632 are alternately displayed. Thus, images 611 to 618 in different layers 631 and 632 can share the same time identifier (ID) and can be included in the same AU. An AU as used herein is a set of one or more decoded images associated with the same display time for output from the DPB. For example, if a smaller image is needed, the decoder can decode and display image 615 at the current display time, or if a larger image is needed, the decoder can decode and display image 611 at the current display time. Thus, images 611 to 614 in the higher layer N+1 632 include substantially the same image data as the corresponding images 615 to 618 in the lower layer N 631 (although the image sizes are different). Specifically, image 611 includes substantially the same image data as image 615, image 612 includes substantially the same image data as image 616, and so on.
[0180] Images 611 to 618 can be decoded with reference to other images 611 to 618 in the same layer N 631 or N+1 632. Decoding one image with reference to another image in the same layer is inter-frame prediction 623, and inter-frame prediction 623 includes uni-directional inter-frame prediction and / or bi-directional inter-frame prediction. Inter-frame prediction 623 is represented by solid arrows. For example, image 613 can be decoded by inter-frame prediction 623 with reference to one or two of images 611, 612, and / or 614 in layer N+1 632, where uni-directional inter-frame prediction uses one image as a reference, and / or bi-directional inter-frame prediction uses two images as references. Additionally, image 617 can be decoded by inter-frame prediction 623 with reference to one or two of images 615, 616, and / or 618 in layer N 631, where uni-directional inter-frame prediction uses one image as a reference, and / or bi-directional inter-frame prediction uses two images as references. When performing inter-frame prediction 623, when one image is used as a reference for another image in the same layer, this image can be called a reference image. For example, image 612 can be a reference image for decoding image 613 according to inter-frame prediction 623. Inter-frame prediction 623 can also be referred to as intra-layer prediction in a multi-layer context. Therefore, inter-frame prediction 623 is a mechanism for decoding samples of a current image by referring to indicative samples in a reference image different from the current image, where the reference image and the current image are in the same layer.
[0181] Images 611 to 618 can also be decoded with reference to other images 611 to 618 in different layers. This process is called inter-layer prediction 621 and is represented by dashed arrows. Inter-layer prediction 621 is a mechanism for decoding samples of a current image by referring to indicative samples in a reference image, where the current image and the reference image are in different layers and thus have different layer IDs. For example, an image in the lower layer N 631 can be used as a reference image for decoding the corresponding image in the higher layer N+1 632. In a specific example, image 611 can be decoded according to inter-layer prediction 621 by referring to image 615. In this case, image 615 is used as an inter-layer reference image. An inter-layer reference image is a reference image for inter-layer prediction 621. In most cases, constraints are imposed on inter-layer prediction 621 such that the current image (e.g., image 611) can only use one or more inter-layer reference images contained in the same AU and located in a lower layer, e.g., image 615. When multiple layers (e.g., more than two layers) are available, inter-layer prediction 621 can encode / decode the current image based on multiple inter-layer reference images with a lower level than the current image.
[0182] A video encoder may use layer-based prediction 600 to encode images 611 to 618 through many different combinations and / or permutations of inter-frame prediction 623 and inter-layer prediction 621. For example, image 615 may be decoded according to intra-frame prediction. Then, by using image 615 as a reference image, images 616 to 618 may be decoded according to inter-frame prediction 623. In addition, by using image 615 as an inter-layer reference image, image 611 may be decoded according to inter-layer prediction 621. Then, by using image 611 as a reference image, images 612 to 614 may be decoded according to inter-frame prediction 623. Thus, a reference image may serve as a single-layer reference image and an inter-layer reference image for different decoding mechanisms. By decoding a high-layer N+1 632 image according to a low-layer N 631 image, the high-layer N+1 632 may avoid using intra-frame prediction, which has a much lower decoding efficiency than inter-frame prediction 623 and inter-layer prediction 621. Therefore, intra-frame prediction with low decoding efficiency is limited to the minimum / lowest quality images and thus limited to decoding the smallest amount of video data. Images used as reference images and / or inter-layer reference images may be indicated in entries of one or more reference image lists included in a reference image list structure.
[0183] Previous H.26x video coding series have supported adaptability in one or more individual profiles according to one or more profiles for single-layer decoding. Scalable video coding (SVC) is an extended version of AVC / H.264 that supports spatial scalability, temporal scalability, and quality scalability. For SVC, a flag is indicated in each macroblock (MB) in the EL image to indicate whether the EL MB uses a collocated block in the lower layer for prediction. Prediction based on the collocated block may include texture, motion vectors, and / or decoding modes. The implementation of SVC cannot directly reuse an unmodified H.264 / AVC implementation in its design. The SVC EL macroblock syntax and decoding process are different from the H.264 / AVC syntax and decoding process.
[0184] Scalable High Efficiency Video Coding (SHVC) is an extended version of the HEVC / H.265 standard, supporting spatial scalability and quality scalability; Multiview High Efficiency Video Coding (MV-HEVC) is an extended version of HEVC / H.265, supporting multiview scalability; 3D High Efficiency Video Coding (3D-HEVC) is an extended version of HEVC / H.264, supporting more advanced and efficient three-dimensional (3D) video decoding than MV-HEVC. It should be noted that temporal scalability is a component of a single-layer HEVC codec. The design of the multi-layer extension of HEVC adopts the following concept: the decoded pictures for inter-layer prediction come only from the same access unit (AU), and are regarded as long-term reference pictures (LTRP), and are assigned a reference index in one or more reference picture lists and other temporal reference pictures in the current layer. Inter-layer prediction (ILP) is implemented at the prediction unit (PU) level by setting the value of the reference index to refer to one or more inter-layer reference pictures in one or more reference picture lists.
[0185] It should be noted that both reference picture resampling and spatial scalability features require resampling of the reference picture or a part thereof. Reference picture resampling can be implemented at the picture level or the coded block level. However, when RPR is referred to as a decoding feature, it is a feature of single-layer decoding. Even so, from the perspective of codec design, it is possible or even preferred to use the same resampling filter to implement the RPR feature of single-layer decoding and the spatial scalability feature of multi-layer decoding.
[0186] Figure 7 Schematic diagram of an example of unidirectional inter-frame prediction 700. Unidirectional inter-frame prediction 700 can be used to determine the motion vectors of the coded blocks and / or decoded blocks created when segmenting an image.
[0187] Unidirectional inter-frame prediction 700 uses a reference frame 730 including a reference block 731 to predict a current block 711 in a current frame 710. As shown, the reference frame 730 can be temporally after the current frame 710 (e.g., as a subsequent reference frame), but in some examples, it can also be temporally before the current frame 710 (e.g., as a previous reference frame). The current frame 710 is an exemplary frame / image being encoded / decoded at a specific time. The current frame 710 includes an object in the current block 711 that matches an object in the reference block 731 of the reference frame 730. The reference frame 730 is a frame used as a reference when encoding the current frame 710, and the reference block 731 is a block in the reference frame 730 that includes an object also included in the current block 711 of the current frame 710.
[0188] The current block 711 is any decoding unit being encoded / decoded at a specified point during the decoding process. When the affine inter-frame prediction mode is adopted, the current block 711 can be the entire segmented block or a sub-block. The current frame 710 is separated from the reference frame 730 by a certain temporal distance (TD) 733. The TD 733 represents the amount of time between the current frame 710 and the reference frame 730 in the video sequence, and the unit of measurement can be frames. The prediction information of the current block 711 can refer to the reference frame 730 and / or the reference block 731 through a reference index representing the direction and temporal distance between the frames. During the time period represented by the TD 733, the object in the current block 711 moves from one position in the current frame 710 to another position in the reference frame 730 (e.g., the position of the reference block 731). For example, the object can move along a motion trajectory 713, which represents the direction in which the object moves over time. The motion vector 735 describes the direction and magnitude of the movement of the object along the motion trajectory 713 within the TD 733. Therefore, the encoded motion vector 735, the reference block 731, and the residual including the difference between the current block 711 and the reference block 731 provide sufficient information to reconstruct the current block 711 and locate the current block 711 in the current frame 710.
[0189] Figure 8 A schematic diagram of an example of bidirectional inter-frame prediction 800. Bidirectional inter-frame prediction 800 can be used to determine the motion vectors of the encoded blocks and / or decoded blocks created when segmenting an image.
[0190] Bidirectional inter - frame prediction 800 is similar to unidirectional inter - frame prediction 700, but uses a pair of reference frames to predict the current block 811 in the current frame 810. Thus, the current frame 810 and the current block 811 are respectively substantially similar to the current frame 710 and the current block 711. The current frame 810 is temporally located between a previous reference frame 820 that appears before the current frame 810 in the video sequence and a subsequent reference frame 830 that appears after the current frame 810 in the video sequence. The previous reference frame 820 and the subsequent reference frame 830 are substantially similar to the reference frame 730 in other respects.
[0191] The current block 811 matches a previous reference block 821 in the previous reference frame 820 and a subsequent reference block 831 in the subsequent reference frame 830. This match indicates that during the playback of the video sequence, an object moves along a motion path 813 from the position of the subsequent reference block 821 through the current block 811 to the position of the subsequent reference block 831. The current frame 810 is separated from the previous reference frame 820 by a previous time distance (TD0) 823 and from the subsequent reference frame 830 by a subsequent time distance (TD1) 833. TD0 823 represents the amount of time, in terms of frames, between the previous reference frame 820 and the current frame 810 in the video sequence. TD1 833 represents the amount of time, in terms of frames, between the current frame 810 and the subsequent reference frame 830 in the video sequence. Thus, the object moves along the motion path 813 from the previous reference block 821 to the current block 811 during the time period represented by TD0 823. The object also moves along the motion path 813 from the current block 811 to the subsequent reference block 831 during the time period represented by TD1 833. The prediction information of the current block 811 can refer to the previous reference frame 820 and / or the previous reference block 821 and the subsequent reference frame 830 and / or the subsequent reference block 831 through a pair of reference indices representing the direction and time distance between the frames.
[0192] The previous motion vector (MV0) 825 describes the direction and magnitude of the movement of the object along the motion path 813 within TD0 823 (e.g., between the previous reference frame 820 and the current frame 810). The subsequent motion vector (MV1) 835 describes the direction and magnitude of the movement of the object along the motion path 813 within TD1 833 (e.g., between the current frame 810 and the subsequent reference frame 830). Thus, in bidirectional inter - frame prediction 800, the current block 811 can be decoded and reconstructed by using the previous reference block 821 and / or the subsequent reference block 831, MV0 825, and MV1 835.
[0193] In one embodiment, inter-frame prediction and / or bidirectional inter-frame prediction may be performed on a per-sample (e.g., per-pixel) basis rather than on a per-block basis. That is, a motion vector pointing to each sample in the previous reference block 821 and / or the subsequent reference block 831 may be determined for each sample in the current block 811. In these embodiments, Figure 8 the motion vectors 825 and 835 shown in FIG. represent a plurality of motion vectors corresponding to a plurality of samples in the current block 811, the previous reference block 821, and the subsequent reference block 831.
[0194] In the merge mode and the advanced motion vector prediction (AMVP) mode, the candidate list is generated by adding candidate motion vectors to the candidate list in the order defined by the candidate list determination mode. Such candidate motion vectors may include motion vectors generated according to unidirectional inter-frame prediction 700, bidirectional inter-frame prediction 800, or a combination thereof. Specifically, the motion vectors are generated for such blocks when adjacent blocks are encoded. Such motion vectors are added to the candidate list of the current block, and the motion vector of the current block is selected from the candidate list. Then, the motion vector may be indicated as the index of the selected motion vector in the candidate list. The decoder may construct the candidate list using the same procedure as the encoder and may determine the selected motion vector from the candidate list according to the indicated index. Thus, the candidate motion vectors include motion vectors generated according to unidirectional inter-frame prediction 700 and / or bidirectional inter-frame prediction 800, depending on the method used when encoding such adjacent blocks.
[0195] Figure 9 A video bitstream 900 is shown. The video bitstream 900 used herein may also be referred to as a decoded video bitstream, a bitstream, or a variant thereof. As Figure 9 shown, the bitstream 900 includes a sequence parameter set (SPS) 902, a picture parameter set (PPS) 904, a slice header 906, and picture data 908.
[0196] The SPS 902 includes data common to all the pictures in a sequence of pictures (SOP). In contrast, the PPS 904 includes data common to the entire picture. The slice header 906 includes information about the current slice, such as the slice type, which reference pictures will be used, etc. The SPS 902 and the PPS 904 can be collectively referred to as parameter sets. The SPS 902, the PPS 904, and the slice header 906 are types of Network Abstraction Layer (NAL) units. A NAL unit is a syntax structure that includes an indication of the type of data to be followed (e.g., coded video data). NAL units are divided into video coding layer (VCL) and non-VCL NAL units. VCL NAL units include data representing sample values in a video picture, and non-VCL NAL units include any relevant additional information, such as parameter sets (important header data applicable to a large number of VCL NAL units) and supplementary enhancement information (timing information and other supplementary data that can enhance the usability of the decoded video signal but are not necessary for decoding the sample values in the video picture). Those skilled in the art will understand that the bitstream 900 may include other parameters and information in practical applications.
[0197] Figure 9 The picture data 908 includes data associated with the picture or video being encoded or decoded. The picture data 908 can simply be referred to as the payload or data carried in the bitstream 900. In one embodiment, the picture data 908 includes CVS914 (or CLVS), and the CVS 914 includes a plurality of pictures 910. The CVS 914 is the decoded video sequence of each coded layer video sequence (CLVS) in the video bitstream 900. It should be noted that when the video bitstream 900 includes a single layer, the CVS and the CLVS are the same. The CVS and the CLVS are different only when the video bitstream 900 includes multiple layers.
[0198] As Figure 9 shown, the slices of each picture 910 can be included in their own VCL NAL units 912. A group of VCL NAL units 912 in the CVS 914 can be referred to as an access unit.
[0199] Figure 10 Illustrates a segmentation technique 1000 for a picture 1010. The picture 1010 can be similar to Figure 9 any of the pictures 910. As shown, the picture 1010 can be segmented into a plurality of slices 1012. A slice is a spatially distinct region of a frame (e.g., a picture), and this region is encoded separately from any other region in the same frame. AlthoughFigure 10 Three stripes 1012 are shown, but more or fewer stripes may be used in actual applications. Each stripe 1012 may be divided into a plurality of blocks 1014. Figure 10 The blocks 1014 in may be similar to Figure 8 the current block 811, the previous reference block 821, and the next reference block 831 in. The block 1014 may represent a CU. Although Figure 10 Four blocks 1014 are shown, but more or fewer blocks may be used in actual applications.
[0200] Each block 1014 may be divided into a plurality of samples 1016 (e.g., pixels). In one embodiment, the size of each block 1014 is measured in terms of luminance samples. Although Figure 10 Sixteen samples 1016 are shown, but more or fewer samples may be used in actual applications.
[0201] In one embodiment, a compliance window 1060 is applied to the image 1010. As described above, the compliance window 1060 is used to crop, reduce, or otherwise change the size of the image 1010 (e.g., the reconstructed / decoded image) during the preparation of the image output. For example, the decoder may apply the compliance window 1060 to the image 1010 to crop, trim, shrink, or otherwise change the size of the image 1010 before the image is output for display to the user. The size of the compliance window 1060 is determined by applying the compliance window top offset 1062, the compliance window bottom offset 1064, the compliance window left offset 1066, and the compliance window right offset 1068 to the image 1010 to reduce the size of the image 1010 before output. That is, only a portion of the image 1010 that exists within the compliance window 1060 is output. Thus, the image 1010 is cropped before output. In one embodiment, the first image parameter set and the second image parameter set refer to the same sequence parameter set, and their image width and image height have the same values. Therefore, the compliance windows of the first image parameter set and the second image parameter set have the same values.
[0202] Figure 11An embodiment of a decoding method 1100 implemented by a video decoder (e.g., video decoder 400). Method 1100 may be performed after receiving a decoded bitstream directly or indirectly from a video encoder (e.g., video encoder 300). By keeping the size of the compliance window for an image parameter set with the same image size the same, method 1100 improves the decoding process. Thus, reference picture resampling (RPR) can be kept enabled or turned on for the entire CVS. By keeping the compliance window sizes for image parameter sets with the same image size consistent, the decoding efficiency can be improved. Thus, in effect, the performance of the codec is improved and the user experience is better.
[0203] In step 1102, the video decoder receives a first image parameter set (e.g., ppsA) and a second image parameter set (e.g., ppsB) that reference the same sequence parameter set. When the image width and image height of the first image parameter set and the second image parameter set have the same values, the compliance windows of the first image parameter set and the second image parameter set have the same values. In one embodiment, the image width and image height are measured in terms of luminance samples.
[0204] In one embodiment, the image width is represented as pic_width_in_luma_samples. In one embodiment, the image height is represented as pic_height_in_luma_samples. In one embodiment, pic_width_in_luma_samples represents the width in terms of luminance samples of each decoded image that references the PPS. In one embodiment, pic_height_in_luma_samples represents the height in terms of luminance samples of each decoded image that references the PPS.
[0205] In one embodiment, the compliance window includes a compliance window left offset, a compliance window right offset, a compliance window top offset, and a compliance window bottom offset that jointly represent the compliance window size. In one embodiment, the compliance window left offset is represented as pps_conf_win_left_offset. In one embodiment, the compliance window right offset is represented as pps_conf_win_right_offset. In one embodiment, the compliance window top offset is represented as pps_conf_win_top_offset. In one embodiment, the compliance window bottom offset is represented as pps_conf_win_bottom_offset. In one embodiment, the compliance window size or value is indicated in the PPS.
[0206] In step 1104, the video decoder applies a compliance window to the current picture corresponding to the first picture parameter set or the second picture parameter set. Through the above steps, the video encoder clips the current picture to the size of the compliance window.
[0207] In one embodiment, the method further includes decoding the current picture using inter prediction based on the resampled reference picture. In one embodiment, the method further includes resampling the reference picture corresponding to the current picture using reference picture resampling (RPS). In one embodiment, resampling of the reference picture changes the resolution of the reference picture.
[0208] In one embodiment, the method further includes determining whether to enable bi-direction optical flow (BDOF) to decode the picture according to the picture width, picture height, and compliance window of the current picture and the reference picture of the current picture. In one embodiment, the method further includes determining whether to enable decoder-side motion vector refinement (DMVR) to decode the picture according to the picture width, picture height, and compliance window of the current picture and the reference picture of the current picture.
[0209] In one embodiment, the method further includes displaying, on a display of an electronic device (such as a smart phone, a tablet computer, a laptop computer, a personal computer, etc.), the picture generated using the current block.
[0210] Figure 12 An embodiment of method 1200 for encoding a video bitstream implemented by a video encoder (such as video encoder 300). Method 1200 may be executed when a picture (such as from a video) is to be encoded into a video bitstream and then sent to a video decoder (such as video decoder 400). By keeping the size of the compliance window of the picture parameter set with the same picture size the same, method 1200 improves the encoding process. Thus, reference picture resampling (RPR) can be kept enabled or on for the entire CVS. By keeping the compliance window sizes of picture parameter sets with the same picture size consistent, the decoding efficiency can be improved. Thus, in fact, the performance of the codec is improved and the user experience is better.
[0211] In step 1202, the video encoder generates a first picture parameter set and a second picture parameter set that refer to the same sequence parameter set. When the picture width and picture height of the first picture parameter set and the second picture parameter set have the same values, the compliance windows of the first picture parameter set and the second picture parameter set have the same values. In one embodiment, the picture width and picture height are measured in units of luma samples.
[0212] In one embodiment, the picture width is represented as pic_width_in_luma_samples. In one embodiment, the picture height is represented as pic_height_in_luma_samples. In one embodiment, pic_width_in_luma_samples represents the width in luma samples of each decoded picture that refers to the PPS. In one embodiment, pic_height_in_luma_samples represents the height in luma samples of each decoded picture that refers to the PPS.
[0213] In one embodiment, the compliance window includes a compliance window left offset, a compliance window right offset, a compliance window top offset, and a compliance window bottom offset that jointly represent the compliance window size. In one embodiment, the compliance window left offset is represented as pps_conf_win_left_offset. In one embodiment, the compliance window right offset is represented as pps_conf_win_right_offset. In one embodiment, the compliance window top offset is represented as pps_conf_win_top_offset. In one embodiment, the compliance window bottom offset is represented as pps_conf_win_bottom_offset. In one embodiment, the compliance window size or value is indicated in the PPS.
[0214] In step 1204, the video encoder encodes the first picture parameter set and the second picture parameter set into the video bitstream. In step 1206, the video encoder stores the video bitstream, where the video bitstream is for sending to the video decoder. In one embodiment, the video encoder sends the video bitstream including the first picture parameter set and the second picture parameter set to the video decoder.
[0215] In one embodiment, a method for encoding a video bitstream is provided. The bitstream includes a plurality of parameter sets and a plurality of pictures. Each picture in the plurality of pictures includes a plurality of slices. Each slice in the plurality of slices includes a plurality of coding blocks. The method includes generating a parameter set parameterSetA and writing it into the bitstream including information, where the information includes a picture size picSizeA and a compliance window confWinA. The parameter set may be a picture parameter set (PPS). The method further includes generating another parameter set parameterSetB and writing it into the bitstream including information, where the information includes a picture size picSizeB and a compliance window confWinB. The parameter set may be a picture parameter set (PPS). The method further includes: when the value of picSizeA in parameterSetA is equal to the value of picSizeB in parameterSetB, constraining the value of the compliance window confWinA in parameterSetA to be equal to the value of the compliance window confWinB in parameterSetB; when the value of confWinA in parameterSetA is equal to the value of confWinB in parameterSetB, constraining the value of the picture size picSizeA in parameterSetA to be equal to the value of picSizeB in parameterSetB. The method further includes encoding the bitstream.
[0216] In one embodiment, a method for decoding a video bitstream is provided. The bitstream includes a plurality of parameter sets and a plurality of pictures. Each picture in the plurality of pictures includes a plurality of slices. Each slice in the plurality of slices includes a plurality of coding blocks. The method includes parsing a parameter set to obtain a picture size and a compliance window size associated with a current picture currPic. The information obtained above is used to derive the picture size and the cropped size of the current picture. The method further includes parsing another parameter set to obtain a picture size and a compliance window size associated with a reference picture refPic. The information obtained above is used to derive the picture size and the cropped size of the reference picture. The method further includes: determining refPic as a reference picture to decode a current block curBlock located in the current picture currPic; determining whether to use or enable bi-direction optical flow (BDOF) to decode the current coding block according to the picture sizes and compliance windows of the current picture and the reference picture; decoding the current block.
[0217] In one embodiment, when the image sizes and compliance windows of the current image and the reference image are different, BDOF is not used or disabled to decode the current coding block.
[0218] In one embodiment, a method for decoding a video bitstream is provided. The bitstream includes a plurality of parameter sets and a plurality of images. Each image in the plurality of images includes a plurality of slices. Each slice in the plurality of slices includes a plurality of coding blocks. The method includes parsing the parameter sets to obtain the image size and compliance window size associated with a current image currPic. The information obtained above is used to derive the image size and the cropped size of the current image. The method further includes parsing another parameter set to obtain the image size and compliance window size associated with a reference image refPic. The information obtained above is used to derive the image size and the cropped size of the reference image. The method further includes: determining refPic as the reference image to decode a current block curBlock located in the current image currPic; determining whether to use or enable decoder-side motion vector refinement (DMVR) to decode the current coding block according to the image sizes and compliance windows of the current image and the reference image; and decoding the current block.
[0219] In one embodiment, when the image sizes and compliance windows of the current image and the reference image are different, DMVR is not used or disabled to decode the current coding block.
[0220] In one embodiment, a method for encoding a video bitstream is provided. In one embodiment, the bitstream includes a plurality of parameter sets and a plurality of pictures. Each picture in the plurality of pictures includes a plurality of slices. Each slice in the plurality of slices includes a plurality of coding blocks. The method includes generating a parameter set, where the parameter set includes a picture size and a compliance window size associated with a current picture currPic. The above information is used to derive the picture size and the cropped size of the current picture. The method further includes generating another parameter set, where the another parameter set includes a picture size and a compliance window size associated with a reference picture refPic. The information obtained above is used to derive the picture size and the cropped size of the reference picture. The method further includes: when the picture sizes and the compliance windows of the current picture and the reference picture are different, constraining that the reference picture refPic cannot be used as a collocated reference picture for temporal motion vector prediction (TMVP) of all slices, where all slices belong to the current picture currPic. That is, constraining that if the reference picture refPic is a collocated reference picture for decoding blocks in the current picture currPic for TMVP, the picture sizes and the compliance windows of the current picture and the reference picture can be the same. The method further includes decoding the bitstream.
[0221] In one embodiment, a method for decoding a video bitstream is provided. The bitstream includes a plurality of parameter sets and a plurality of pictures. Each picture in the plurality of pictures includes a plurality of slices. Each slice in the plurality of slices includes a plurality of coded blocks. The method includes parsing the parameter sets to obtain the picture size and the compliance window size associated with the current picture currPic. The information obtained above is used to derive the picture size and the cropped size of the current picture. The method further includes parsing another parameter set to obtain the picture size and the compliance window size associated with the reference picture refPic. The information obtained above is used to derive the picture size and the cropped size of the reference picture. The method further includes: determining refPic as the reference picture for decoding the current block curBlock located in the current picture currPic; parsing the syntax element (slice_DVMR_BDOF_enable_flag) to determine whether to use or enable decoder-side motion vector refinement (DMVR) and / or bi-direction optical flow (BDOF) for decoding the current coded picture and / or slice. The method further includes: when the compliance window confWinA in parameterSetA is different from the compliance window confWinB in parameterSetB or the value of picSizeA in parameterSetA is different from the value of picSizeB in parameterSetB, constraining the value of the syntax element (slice_DVMR_BDOF_enable_flag) to be equal to 0.
[0222] The following describes the base text, i.e., the VVC working draft. That is, only the deltas are described, and the text in the base text not mentioned below applies as is. Deleted text is italicized, and added text is bolded.
[0223] The syntax and semantics of the sequence parameter set are provided.
[0224]
[0225] max_width_in_luma_samples indicates that the requirement for bitstream compliance is that the pic_width_in_luma_samples of any picture for which this SPS is active is less than or equal to max_width_in_luma_samples.
[0226] max_height_in_luma_samples indicates that the requirement for bitstream compliance is that for any picture with this SPS being active, pic_height_in_luma_samples is less than or equal to max_height_in_luma_samples.
[0227] The syntax and semantics of the picture parameter set are provided.
[0228]
[0229]
[0230] pic_width_in_luma_samples represents the width in luma samples of each decoded picture of the reference PPS. pic_width_in_luma_samples shall not be equal to 0 and shall be an integer multiple of MinCbSizeY.
[0231] pic_height_in_luma_samples represents the height in luma samples of each decoded picture of the reference PPS. pic_height_in_luma_samples shall not be equal to 0 and shall be an integer multiple of MinCbSizeY.
[0232] The requirement for bitstream compliance is that for each active reference picture with width and height reference_pic_width_in_luma_samples and reference_pic_height_in_luma_samples, all of the following conditions are satisfied:
[0233] –2*pic_width_in_luma_samples >= reference_pic_width_in_luma_samples
[0234] –2*pic_height_in_luma_samples >= reference_pic_height_in_luma_samples
[0235] –pic_width_in_luma_samples <= 8*reference_pic_width_in_luma_samples
[0236] –pic_height_in_luma_samples <= 8*reference_pic_height_in_luma_samples
[0237] The variables PicWidthInCtbsY, PicHeightInCtbsY, PicSizeInCtbsY, PicWidthInMinCbsY, PicHeightInMinCbsY, PicSizeInMinCbsY, PicSizeInSamplesY, PicWidthInSamplesC, and PicHeightInSamplesC are derived as follows:
[0238] PicWidthInCtbsY = Ceil(pic_width_in_luma_samples ÷ CtbSizeY) (1)
[0239] PicHeightInCtbsY = Ceil(pic_height_in_luma_samples ÷ CtbSizeY) (2)
[0240] PicSizeInCtbsY = PicWidthInCtbsY * PicHeightInCtbsY (3)
[0241] PicWidthInMinCbsY = pic_width_in_luma_samples / MinCbSizeY (4)
[0242] PicHeightInMinCbsY = pic_height_in_luma_samples / MinCbSizeY (5)
[0243] PicSizeInMinCbsY = PicWidthInMinCbsY * PicHeightInMinCbsY (6)
[0244] PicSizeInSamplesY = pic_width_in_luma_samples * pic_height_in_luma_samples (7)
[0245] PicWidthInSamplesC = pic_width_in_luma_samples / SubWidthC (8)
[0246] PicHeightInSamplesC = pic_height_in_luma_samples / SubHeightC (9)
[0247] Conformance_window_flag = 1 indicates that the conformance cropping window offset parameter is after the next parameter in the PPS. conformance_window_flag = 0 indicates that the conformance cropping window offset parameter does not exist.
[0248] According to the rectangular area specified for output in the image coordinates, conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset represent samples of the image of the reference PPS output from the decoding process. When conformance_window_flag = 0, the values of conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset are inferred to be 0.
[0249] The conformance cropping window includes luma samples with horizontal and vertical image coordinates, where the range of the horizontal image coordinates is from SubWidthC * conf_win_left_offset to pic_width_in_luma_samples – (SubWidthC * conf_win_right_offset + 1), and the range of the vertical image coordinates is from SubHeightC * conf_win_top_offset to pic_height_in_luma_samples – (SubHeightC * conf_win_bottom_offset + 1) (including the end values).
[0250] The value of SubWidthC * (conf_win_left_offset + conf_win_right_offset) can be less than pic_width_in_luma_samples, and the value of SubHeightC * (conf_win_top_offset + conf_win_bottom_offset) can be less than pic_height_in_luma_samples.
[0251] The variables PicOutputWidthL and PicOutputHeightL are derived as follows:
[0252] PicOutputWidthL = pic_width_in_luma_samples – (10) SubWidthC * (conf_win_right_offset + conf_win_left_offset)
[0253] PicOutputHeightL = pic_height_in_pic_size_units – (11) SubHeightC * (conf_win_bottom_offset + conf_win_top_offset)
[0254] When ChromaArrayType is not equal to 0, the corresponding specified samples of the two chroma arrays are the samples with image coordinates (x / SubWidthC, y / SubHeightC), where (x, y) are the image coordinates of the specified luma sample.
[0255] Note: The compliance cropping window offset parameters are only applied to the output. All internal decoding processes are applied to the uncropped image size.
[0256] Let PPS_A and PPS_B be picture parameter sets that refer to the same sequence parameter set. The bitstream compliance requirements are as follows: If pic_width_in_luma_samples in PPS_A and PPS_B have the same value and pic_height_in_luma_samples in PPS_A and PPS_B have the same value, then all of the following conditions can be true:
[0257] conf_win_left_offset in PPS_A and PPS_B have the same value,
[0258] conf_win_right_offset in PPS_A and PPS_B have the same value,
[0259] conf_win_top_offset in PPS_A and PPS_B have the same value,
[0260] conf_win_bottom_offset in PPS_A and PPS_B have the same value.
[0261] The following constraints are added to the semantics of collocated_ref_idx:
[0262] collocated_ref_idx represents the reference index of the collocated picture used for temporal motion vector prediction.
[0263] When slice_type is equal to P or slice_type is equal to B and collocated_from_l0_flag is equal to 1, collocated_ref_idx refers to the picture in reference list 0, and the value range of collocated_ref_idx can be from 0 to NumRefIdxActive[0]–1 (including the end values).
[0264] When slice_type = B and collocated_from_l0_flag = 0, collocated_ref_idx refers to the picture in reference list 1, and the value range of collocated_ref_idx can be from 0 to NumRefIdxActive[1]–1 (including the end values).
[0265] When collocated_ref_idx does not exist, the value of collocated_ref_idx is inferred to be 0.
[0266] The requirement for bitstream compliance is that the picture referred to by collocated_ref_idx can be the same for all slices of the decoded picture.
[0267] The requirement for bitstream compliance is that the resolution of the reference picture referred to by collocated_ref_idx can be the same as that of the current picture.
[0268] The requirement for bitstream compliance is that the picture size and compliance window of the reference picture referred to by collocated_ref_idx can be the same as those of the current picture.
[0269] Modify the following conditions for setting dmvrFlag to 1:
[0270] – When all of the following conditions are true, dmvrFlag is set to 1:
[0271] – sps_dmvr_enabled_flag is equal to 1,
[0272] – general_merge_flag[xCb][yCb] is equal to 1,
[0273] – predFlagL0[0][0] and predFlagL1[0][0] are both equal to 1,
[0274] – mmvd_merge_flag[xCb][yCb] is equal to 0,
[0275] –DiffPicOrderCnt(currPic, RefPicList[0][refIdxL0]) is equal to DiffPicOrderCnt(RefPicList[1][refIdxL1], currPic),
[0276] –BcwIdx[xCb][yCb] is equal to 0,
[0277] –Both luma_weight_l0_flag[refIdxL0] and luma_weight_l1_flag[refIdxL1] are equal to 0,
[0278] –cbWidth is greater than or equal to 8,
[0279] –cbHeight is greater than or equal to 8,
[0280] –cbHeight * cbWidth is greater than or equal to 128,
[0281] –When X is 0 or 1, pic_width_in_luma_samples and pic_height_in_luma_samples of the reference picture refPicLX associated with refIdxLX are equal to pic_width_in_luma_samples and pic_height_in_luma_samples of the current picture respectively.
[0282] –When X is 0 or 1, pic_width_in_luma_samples, pic_height_in_luma_samples, conf_win_left_offset, conf_win_right_offset, conf_win_top_offset and conf_win_bottom_offset of the reference picture refPicLX associated with refIdxLX are equal to pic_width_in_luma_samples, pic_height_in_luma_samples, conf_win_left_offset, conf_win_right_offset, conf_win_top_offset and conf_win_bottom_offset of the current picture respectively.
[0283] Modify the following conditions for setting dmvrFlag to 1:
[0284] –If all of the following conditions are true, then bdofFlag is set to true (TRUE):
[0285] – The sps_bdof_enabled_flag is equal to 1,
[0286] – Both predFlagL0[xSbIdx][ySbIdx] and predFlagL1[xSbIdx][ySbIdx] are equal to 1,
[0287] – DiffPicOrderCnt(currPic, RefPicList[0][refIdxL0]) * DiffPicOrderCnt(currPic, RefPicList[1][refIdxL1]) is less than 0,
[0288] – MotionModelIdc[xCb][yCb] is equal to 0,
[0289] – merge_subblock_flag[xCb][yCb] is equal to 0,
[0290] – sym_mvd_flag[xCb][yCb] is equal to 0,
[0291] – BcwIdx[xCb][yCb] is equal to 0,
[0292] – Both luma_weight_l0_flag[refIdxL0] and luma_weight_l1_flag[refIdxL1] are equal to 0,
[0293] – cbHeight is greater than or equal to 8,
[0294] – When X is 0 or 1, the pic_width_in_luma_samples and pic_height_in_luma_samples of the reference picture refPicLX associated with refIdxLX are equal to the pic_width_in_luma_samples and pic_height_in_luma_samples of the current picture, respectively.
[0295] – When X is 0 or 1, pic_width_in_luma_samples, pic_height_in_luma_samples, conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset of the reference picture refPicLX associated with refIdxLX are respectively equal to pic_width_in_luma_samples, pic_height_in_luma_samples, conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset of the current picture.
[0296] – cIdx is equal to 0.
[0297] Figure 13 FIG. is a schematic diagram of a video decoding device 1300 (e.g., video encoder 20 or video encoder 30) provided by an embodiment of the present invention. The video decoding device 1300 is adapted to implement the disclosed embodiments described herein. The video decoding device 1300 includes an input port 1310 for receiving data and a receiver unit (Rx) 1320; a processor, logic unit, or central processing unit (CPU) 1330 for processing data; a transmitter unit (Tx) 1340 and an output port 1350 for transmitting data; and a memory 1360 for storing data. The video decoding device 1300 may further include optical-to-electrical (OE) components and electrical-to-optical (EO) components coupled to the input port 1310, the receiver unit 1320, the transmitter unit 1340, and the output port 1350 for the input or output of optical or electrical signals.
[0298] The processor 1330 is implemented by hardware and software. The processor 1330 can be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), field-programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor 1330 communicates with the input port 1310, the receiver unit 1320, the transmitter unit 1340, the output port 1350, and the memory 1360. The processor 1330 includes a decoding module 1370. The decoding module 1370 implements the embodiments disclosed above. For example, the decoding module 1370 performs, processes, prepares, or provides various encoding and decoding functions. Thus, including the decoding module 1370 provides a substantial improvement to the functions of the video decoding device 1300 and enables the conversion of the video decoding device 1300 to different states. Alternatively, the decoding module 1370 is implemented by instructions stored in the memory 1360 and executed by the processor 1330.
[0299] The video decoding device 1300 may further include an input and / or output (I / O) device 1380 for sending data to and receiving data from a user. The I / O device 1380 may include output devices such as a display for displaying video data, a speaker for outputting audio data, etc. The I / O device 1380 may further include input devices such as a keyboard, a mouse, a trackball, etc. and / or corresponding interfaces for interacting with the above output devices.
[0300] The memory 1360 includes one or more disks, tape drives, and solid state drives, and can be used as an overflow data storage device to store these programs when a selected program is executed, as well as instructions and data read during program execution. The memory 1360 can be volatile and / or non-volatile, and can be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).
[0301] Figure 14Schematic diagram of an embodiment of the decoding module 1400. In one embodiment, the decoding module 1400 is implemented in a video decoding device 1402 (e.g., a video encoder 20 or a video decoder 30). The video decoding device 1402 includes a receiving module 1401. The receiving module 1401 is configured to receive an image for encoding or receive a bitstream for decoding. The video decoding device 1402 includes a transmitting module 1407 coupled to the receiving module 1401. The transmitting module 1407 is configured to send a bitstream to a decoder or send a decoded image to a display module (e.g., one of the I / O devices 1380).
[0302] The video decoding device 1402 includes a storage module 1403. The storage module 1403 is coupled to at least one of the receiving module 1401 or the transmitting module 1407. The storage module 1403 is configured to store instructions. The video decoding device 1402 includes a processing module 1405. The processing module 1405 is coupled to the storage module 1403. The processing module 1405 is configured to execute the instructions stored in the storage module 1403 to perform the methods disclosed herein.
[0303] It should also be understood that the steps of the exemplary methods set forth herein need not be performed in the order described, and the order of these method steps should be understood to be merely exemplary. Similarly, in methods consistent with various embodiments of the present invention, these methods may include other steps, and certain steps may be omitted or combined.
[0304] Although several embodiments have been provided in the present invention, it should be understood that the systems and methods disclosed in the present invention may be embodied in many other specific forms without departing from the spirit or scope of the present invention. The examples of the present invention should be considered illustrative rather than restrictive, and the present invention is not limited to the details given herein. For example, various elements or components may be combined or merged in another system, or certain features may be omitted or not implemented.
[0305] In addition, without departing from the scope of the present invention, the techniques, systems, subsystems, and methods described and illustrated as discrete or separate in various embodiments may be combined or merged with other systems, modules, techniques, or methods. Other items shown or described as being coupled or directly coupled or communicating with each other may be indirectly coupled or communicating via some interface, device, or intermediate component in an electrical, mechanical, or other manner. Examples of other variations, substitutions, and alterations may be determined by those skilled in the art without departing from the spirit and the disclosed scope herein.
Claims
1. A decoding method, characterized in that, Comprising: Receiving a bitstream, the bitstream including a first Picture Parameter Set (PPS) and a second PPS that refer to the same Sequence Parameter Set (SPS), wherein the first PPS includes an image width parameter, an image height parameter, and a conformance window flag, and the second PPS includes an image width parameter, an image height parameter, and a conformance window flag; When the value of the conformance window flag in the first PPS is 1, it indicates that the first PPS further includes conformance window parameters; when the value of the conformance window flag in the second PPS is 1, it indicates that the second PPS further includes conformance window parameters; When the image width parameter and the image height parameter in the first PPS and the image width parameter and the image height parameter in the second PPS have the same values respectively, the conformance window parameters in the first PPS and the conformance window parameters in the second PPS have the same values; Decoding the bitstream to obtain a decoded image.
2. The method according to claim 1, wherein The conformance window parameters include a conformance window left offset, a conformance window right offset, a conformance window top offset, and a conformance window bottom offset.
3. The method according to claim 1 or 2, characterized in that Further comprising applying the conformance window parameters to a current image corresponding to the first PPS or the second PPS, and after applying the conformance window, performing decoding of the current image using inter prediction.
4. The method according to claim 1 or 2, characterized in that, When the value of the conformance window flag in the first PPS is 0, it indicates that the first PPS does not include the conformance window parameters; when the value of the conformance window flag in the second PPS is 0, it indicates that the second PPS does not include the conformance window parameters.
5. The method according to claim 1 or 2, characterized in that, Further comprising resampling a reference image associated with a current image corresponding to the first PPS or the second PPS using a Reference Picture Sampling (RPS).
6. The method according to claim 1 or 2, characterized in that, The image width parameter and the image height parameter are in units of luma samples.
7. The method according to claim 1 or 2, characterized in that, The compliance window includes luminance samples with horizontal image coordinates and vertical image coordinates, where the range of the horizontal image coordinates is from SubWidthC * conf_win_left_offset to pic_width_in_luma_samples – (SubWidthC * conf_win_right_offset + 1), and the range of the vertical image coordinates is from SubHeightC * conf_win_top_offset to pic_height_in_luma_samples – (SubHeightC * conf_win_bottom_offset + 1), where SubWidthC is the sub-width coefficient, SubHeightC is the sub-height coefficient, pic_width_in_luma_samples is the image width in luminance samples, pic_height_in_luma_samples is the image height in luminance samples, conf_win_left_offset is the left offset of the compliance window, conf_win_right_offset is the right offset of the compliance window, conf_win_top_offset is the top offset of the compliance window, and conf_win_bottom_offset is the bottom offset of the compliance window.
8. The method according to claim 1 or 2, characterized in that, It further includes determining whether to enable Bidirectional Optical Flow (BDOF) according to the image width parameter, the image height parameter, and the compliance window parameter of the current image, and the reference image of the current image.
9. The method according to claim 1 or 2, characterized in that It further includes determining whether to enable Decoder-side Motion Vector Refinement (DMVR) according to the image width parameter, the image height parameter, and the compliance window parameter of the current image, and the reference image of the current image.
10. The method according to claim 1 or 2, characterized in that, The bitstream further includes the SPS, and the SPS includes a maximum image width parameter and a maximum image height parameter, and the maximum image width parameter and the maximum image height parameter are in luminance samples.
11. A decoding device, characterized in that, It includes: a receiver and a processor; the receiver is configured to receive a decoded video bitstream; the processor is configured to execute instructions to cause the decoding device to: receive a bitstream, the bitstream including a first Picture Parameter Set (PPS) and a second PPS that refer to the same Sequence Parameter Set (SPS), where the first PPS includes an image width parameter, an image height parameter, and a conformance window flag, and the second PPS includes an image width parameter, an image height parameter, and a conformance window flag; when the value of the conformance window flag in the first PPS is 1, it indicates that the first PPS further includes conformance window parameters; when the value of the conformance window flag in the second PPS is 1, it indicates that the second PPS further includes conformance window parameters; When the image width parameter and the image height parameter in the first PPS and the image width parameter and the image height parameter in the second PPS have the same values respectively, the compliance window parameter in the first PPS and the compliance window parameter in the second PPS have the same value; Decode the bitstream to obtain a decoded image.
12. The decoding device according to claim 11, characterized in that, The compliance window parameter includes a compliance window left offset, a compliance window right offset, a compliance window top offset, and a compliance window bottom offset.
13. The decoding device according to claim 11 or 12, characterized in that, The instruction further causes the decoding device to apply the compliance window parameter to the current image corresponding to the first PPS or the second PPS. After applying the compliance window, perform decoding on the current image using inter-frame prediction.
14. The decoding device according to claim 11 or 12, characterized in that, When the value of the compliance window flag in the first PPS is 0, it indicates that the first PPS does not include the compliance window parameter; when the value of the compliance window flag in the second PPS is 0, it indicates that the second PPS does not include the compliance window parameter.
15. The decoding device according to claim 11 or 12, characterized in that, The instruction further causes the decoding device to resample the reference image associated with the current image corresponding to the first PPS or the second PPS using the reference picture resampling RPS.
16. The decoding device according to claim 11 or 12, characterized in that, The image width parameter and the image height parameter are in units of luma samples.
17. The decoding device according to claim 11 or 12, characterized in that, The compliance window includes luma samples with horizontal image coordinates and vertical image coordinates. Among them, the range of the horizontal image coordinates is from SubWidthC * conf_win_left_offset to pic_width_in_luma_samples – (SubWidthC * conf_win_right_offset + 1), and the range of the vertical image coordinates is from SubHeightC * conf_win_top_offset to pic_height_in_luma_samples – (SubHeightC * conf_win_bottom_offset + 1), where SubWidthC is the sub-width coefficient, SubHeightC is the sub-height coefficient, pic_width_in_luma_samples is the image width in units of luma samples, pic_height_in_luma_samples is the image height in units of luma samples, conf_win_left_offset is the compliance window left offset, conf_win_right_offset is the compliance window right offset, conf_win_top_offset is the compliance window top offset, and conf_win_bottom_offset is the compliance window bottom offset.
18. The decoding device according to claim 11 or 12, characterized in that, The instruction further causes the decoding device to determine whether to enable bidirectional optical flow BDOF according to the image width parameter, the image height parameter, the compliance window parameter of the current image, and the reference image of the current image.
19. The decoding device according to claim 11 or 12, characterized in that, The instruction further causes the decoding device to determine whether to enable decoder-side motion vector refinement (DMVR) according to the image width parameter, the image height parameter, and the compliance window parameter of the current image, and the reference image of the current image.
20. The decoding device according to claim 11 or 12, characterized in that, The bitstream further includes the SPS, and the SPS includes a maximum image width parameter and a maximum image height parameter, and the maximum image width parameter and the maximum image height parameter are in units of luma samples.
21. A codec system, characterized in that, Comprising: An encoder; A decoder, communicating with the encoder, wherein the decoder includes a decoding device according to any one of claims 11 to 20.
22. A decoding module, characterized in that, Comprising: A receiving module, configured to receive a bitstream; A storage module, coupled to the receiving module and configured to store instructions; A processing module, coupled to the storage module and configured to execute the instructions stored in the storage module to perform the decoding method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Conformance window information in multi-layer coding
CN106165429A
360-degree video coding using geometry projection
CN109417632A