Handling multiple picture sizes and conformance windows for reference picture resampling in video coding

By ensuring consistent conformance window sizes for picture parameter sets with the same picture, the solution effectively addresses the inefficiencies in existing technologies by reducing the processor, memory, and/or network resource usage, thereby enhancing the user experience in video coding and decoding processes.

JP7789833B2Active Publication Date: 2025-12-22HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024065997
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-07-08
Filing Date
2024-04-16
Publication Date
2025-12-22
Estimated Expiration
2040-07-07

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently managing multiple picture sizes and conformance windows, leading to excessive processing complexity and resource usage in both encoding and decoding processes, particularly when reference picture resampling is enabled.

Method used

Constraining picture parameter sets with the same picture size to have the same conformance window size, thereby avoiding excessive processing complexity by maintaining consistent window sizes during reference picture resampling.

Benefits of technology

This approach reduces processor, memory, and network resource usage, enhancing the user experience by improving the efficiency of video coding and decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007789833000006
    Figure 0007789833000006
  • Figure 0007789833000007
    Figure 0007789833000007
  • Figure 0007789833000008
    Figure 0007789833000008
Patent Text Reader

Abstract

To provide a method of decoding a coded video bitstream.SOLUTION: A method of decoding includes receiving a first picture parameter set and a second picture parameter set each referring to the same sequence parameter set. The first and second picture parameter sets have the same values of a conformance window 1060 when the first and second picture parameter sets have the same values of picture width and picture height. The method also includes applying the conformance window to a current picture corresponding to the first or second picture parameter set.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This patent application claims the benefit of U.S. Provisional Patent Application No. 62 / 871,493, entitled "Handling of Multiple Picture Size and Conformance Windows for Reference Picture Resampling in Video Coding," filed July 8, 2019, by Jianle Chen et al., which is incorporated herein by reference in its entirety.

[0002] This disclosure generally describes techniques for supporting multiple picture sizes and conformance windows in video coding. More specifically, this disclosure ensures that picture parameter sets with the same picture size also have the same conformance window. [Background technology]

[0003] The amount of video data required to render even a relatively short video can be substantial, which can pose challenges when streaming or otherwise communicating data over communication networks with limited bandwidth capacity. Therefore, video data is typically compressed before being communicated over modern telecommunications networks. The size of the video can also be an issue if the video is stored on a storage device, which may have limited memory resources. Video compression devices often use software and / or hardware at the source to encode video data before transmitting or storing it, thereby reducing the amount of data required to represent a digital video image. The compressed data is then received at the destination by a video decompression device, which decodes the video data. Limited network resources and ever-increasing demands for high video quality dictate improved compression / decompression techniques that increase compression ratios with little or no sacrifice in image quality. Summary of the Invention

[0004] A first aspect relates to a method of decoding an encoded video bitstream performed by a video decoder, the method including receiving, by the video decoder, a first picture parameter set and a second picture parameter set, each of which references the same sequence parameter set, where if the first picture parameter set and the second picture parameter set have the same values ​​for picture width and picture height, the first picture parameter set and the second picture parameter set have the same value for a conformance window, and applying, by the video decoder, the conformance window to a current picture corresponding to the first picture parameter set or the second picture parameter set.

[0005] The method provides a technique for constraining picture parameter sets with the same picture size to also have the same conformance window size (e.g., cropping the window size). By maintaining the same conformance window size for picture parameter sets with the same picture size, excessively complex processing can be avoided when reference picture resampling (RPR) is enabled. Therefore, processor, memory, and / or network resource usage may be reduced in both the encoder and the decoder. This improves the coder / decoder (also known as codec) in video coding compared to current codecs. In practical terms, improvements to the video coding process provide users with a more favorable user experience when transmitting, receiving, and / or watching video.

[0006] Optionally, in any of the aforementioned aspects, another implementation example of this aspect provides that the conformance window has a conformance window left offset, a conformance window right offset, a conformance window top offset, and a conformance window bottom offset.

[0007] Optionally, in any of the aforementioned aspects, another implementation example of the aspect provides decoding the current picture corresponding to the first picture parameter set or the second picture parameter set using inter prediction after applying a conformance window, where the inter prediction is based on a resampled reference picture.

[0008] Optionally, in any of the aforementioned aspects, another implementation example of the present aspect provides a step of resampling a reference picture associated with a current picture corresponding to the first picture parameter set or the second picture parameter set using reference picture resampling (RPS).

[0009] Optionally, in any of the aforementioned aspects, another implementation example of the present aspect specifies that the resampling of the reference picture changes the resolution of the reference picture used to inter-predict the current picture corresponding to the first picture parameter set or the second picture parameter set.

[0010] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the picture width and picture height are measured in luma samples.

[0011] Optionally, in any of the aforementioned aspects, another implementation example of the aspect provides a step of determining whether bidirectional optical flow (BDOF) is valid for decoding the picture based on the picture width, picture height, and conformance window of the current picture and the reference picture of the current picture.

[0012] Optionally, in any of the aforementioned aspects, another implementation example of the aspect provides a step of determining whether decoder-side motion vector refinement (DMVR) is effective for decoding the picture based on the picture width, picture height, and conformance window of the current picture and the reference picture of the current picture.

[0013] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides displaying an image generated using the current block on a display of the electronic device.

[0014] A second aspect relates to a method of encoding a video bitstream performed by a video encoder, the method including: the video encoder generating a first picture parameter set and a second picture parameter set, each of which references the same sequence parameter set, where if the first picture parameter set and the second picture parameter set have the same values ​​for picture width and picture height, the first picture parameter set and the second picture parameter set have the same value for a conformance window; the video encoder encoding the first picture parameter set and the second picture parameter set into a video bitstream; and the video encoder storing the video bitstream for transmission to a video decoder.

[0015] The method provides a technique for constraining picture parameter sets with the same picture size to also have the same conformance window size (e.g., cropping the window size). By maintaining the same conformance window size for picture parameter sets with the same picture size, excessively complex processing can be avoided when reference picture resampling (RPR) is enabled. Therefore, processor, memory, and / or network resource usage may be reduced in both the encoder and the decoder. This improves the coder / decoder (also known as codec) in video coding compared to current codecs. In practical terms, improvements to the video coding process provide users with a more favorable user experience when transmitting, receiving, and / or watching video.

[0016] Optionally, in any of the aforementioned aspects, another implementation example of this aspect provides that the conformance window has a conformance window left offset, a conformance window right offset, a conformance window top offset, and a conformance window bottom offset.

[0017] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the picture width and picture height are measured in luma samples.

[0018] Optionally, in any of the above aspects, another implementation example of the aspect provides transmitting, to a video decoder, a video bitstream including the first picture parameter set and the second picture parameter set.

[0019] A third aspect relates to a decoding device including: a receiver configured to receive a coded video bitstream; a memory coupled to the receiver, the memory storing instructions; and a processor coupled to the memory, the processor configured to execute the instructions to cause the decoding device to receive a first picture parameter set and a second picture parameter set, each of which references the same sequence parameter set, where if the first picture parameter set and the second picture parameter set have the same values ​​for picture width and picture height, the first picture parameter set and the second picture parameter set have the same value for a conformance window, and to apply the conformance window to a current picture corresponding to the first picture parameter set or the second picture parameter set.

[0020] The decoding device provides a technique for constraining picture parameter sets with the same picture size to also have the same conformance window size (e.g., cropping the window size). By maintaining the same conformance window size for picture parameter sets with the same picture size, excessively complex processing can be avoided when reference picture resampling (RPR) is enabled. Therefore, processor, memory, and / or network resource usage may be reduced in both the encoder and the decoder. This improves the coder / decoder (also known as codec) in video coding compared to current codecs. In practical terms, improvements to the video coding process provide users with a more favorable user experience when transmitting, receiving, and / or watching video.

[0021] Optionally, in any of the aforementioned aspects, another implementation example of this aspect provides that the conformance window has a conformance window left offset, a conformance window right offset, a conformance window top offset, and a conformance window bottom offset.

[0022] Optionally, in any of the aforementioned aspects, another implementation example of the aspect provides decoding the current picture corresponding to the first picture parameter set or the second picture parameter set using inter prediction after applying a conformance window, where the inter prediction is based on a resampled reference picture.

[0023] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides a display configured to display an image generated based on the current picture.

[0024] A fourth aspect relates to an encoding device including: a memory including instructions; a processor coupled to the memory, the processor configured to execute the instructions to cause the encoding device to generate a first picture parameter set and a second picture parameter set that each reference a same sequence parameter set, where if the first picture parameter set and the second picture parameter set have the same values ​​for picture width and picture height, the first picture parameter set and the second picture parameter set have the same value for a conformance window, and encode the first picture parameter set and the second picture parameter set into a video bitstream; and a transmitter coupled to the processor, the transmitter configured to transmit the video bitstream including the first picture parameter set and the second picture parameter set to a video decoder.

[0025] The encoding device provides a technique for constraining picture parameter sets with the same picture size to also have the same conformance window size (e.g., cropping the window size). By maintaining the same conformance window size for picture parameter sets with the same picture size, excessively complex processing can be avoided when reference picture resampling (RPR) is enabled. Therefore, processor, memory, and / or network resource usage may be reduced in both the encoder and the decoder. This improves the coder / decoder (also known as codec) in video coding compared to current codecs. In practical terms, improvements to the video coding process provide users with a more favorable user experience when transmitting, receiving, and / or watching video.

[0026] Optionally, in any of the aforementioned aspects, another implementation example of this aspect provides that the conformance window has a conformance window left offset, a conformance window right offset, a conformance window top offset, and a conformance window bottom offset.

[0027] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the picture width and picture height are measured in luma samples.

[0028] A fifth aspect relates to an encoding device including: a receiver configured to receive a picture to encode or a bitstream to decode, a transmitter coupled to the receiver, the transmitter configured to transmit the bitstream to a decoder or to transmit the decoded image to a display, a memory coupled to at least one of the receiver or the transmitter, the memory configured to store instructions, and a processor coupled to the memory, the processor configured to execute the instructions stored in the memory to perform any of the methods disclosed herein.

[0029] The encoding device provides a technique for constraining picture parameter sets with the same picture size to also have the same conformance window size (e.g., cropping the window size). By maintaining the same conformance window size for picture parameter sets with the same picture size, excessively complex processing can be avoided when reference picture resampling (RPR) is enabled. Therefore, processor, memory, and / or network resource usage may be reduced in both the encoder and the decoder. This improves the coder / decoder (also known as codec) in video encoding compared to current codecs. In practical terms, improvements to the video encoding process provide users with a more favorable user experience when transmitting, receiving, and / or watching video.

[0030] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides a display configured to display the image.

[0031] A sixth aspect relates to a system, the system including an encoder and a decoder in communication with the encoder, the encoder or decoder including a decoding device, encoding device, or coding apparatus disclosed herein.

[0032] The system provides a technique for constraining picture parameter sets with the same picture size to also have the same conformance window size (e.g., cropping the window size). By maintaining the same conformance window size for picture parameter sets with the same picture size, excessively complex processing can be avoided when reference picture resampling (RPR) is enabled. Therefore, processor, memory, and / or network resource usage may be reduced in both the encoder and the decoder. This improves the coder / decoder (also known as codec) in video coding compared to current codecs. In practical terms, improvements to the video coding process provide users with a more favorable user experience when transmitting, receiving, and / or watching video.

[0033] A seventh aspect relates to a means for encoding, the means for encoding including receiving means configured to receive a picture to encode or to receive a bitstream to decode, transmitting means coupled to the receiving means, the transmitting means configured to transmit the bitstream to the decoding means or to transmit the decoded image to the display means, storage means coupled to at least one of the receiving means or the transmitting means, the storage means configured to store instructions, and processing means coupled to the storage means, the processing means configured to execute the instructions stored in the storage means to perform any of the methods disclosed herein.

[0034] The coding method provides a technique for constraining picture parameter sets with the same picture size to also have the same conformance window size (e.g., cropping the window size). By maintaining the same conformance window size for picture parameter sets with the same picture size, excessively complex processing can be avoided when reference picture resampling (RPR) is enabled. Therefore, processor, memory, and / or network resource usage may be reduced in both the encoder and the decoder. This improves the coder / decoder (also known as codec) in video coding compared to current codecs. In practical terms, improvements to the video coding process provide users with a more favorable user experience when transmitting, receiving, and / or watching video.

[0035] For purposes of clarity, any one of the above-described embodiments may be combined with any one or more of the other above-described embodiments to create new embodiments within the scope of the present disclosure.

[0036] These and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims. [Brief explanation of the drawings]

[0037] For a more complete understanding of the present disclosure, reference is now made to the following brief description taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals represent like parts.

[0038] [Figure 1] 1 is a flowchart of an example method for encoding a video signal.

[0039] [Figure 2] 1 is a schematic diagram of an example of a coding and decoding (codec) system for video encoding.

[0040] [Figure 3] FIG. 1 is a schematic diagram illustrating an example of a video encoder.

[0041] [Figure 4] FIG. 1 is a schematic diagram illustrating an example of a video decoder.

[0042] [Figure 5] 1 is a coded video sequence showing the relationship of Intra Random Access Point (IRAP) pictures to leading and trailing pictures in decoding order and display order.

[0043] [Figure 6] An example of multi-layer coding for spatial scalability is given below.

[0044] [Figure 7] FIG. 1 is a schematic diagram illustrating an example of unidirectional inter prediction.

[0045] [Figure 8] FIG. 1 is a schematic diagram illustrating an example of bidirectional inter prediction.

[0046] [Figure 9] 1 shows a video bitstream.

[0047] [Figure 10] 1 shows a picture division method.

[0048] [Figure 11] 1 is an embodiment of a method for decoding an encoded video bitstream.

[0049] [Figure 12] 1 is an embodiment of a method for encoding a coded video bitstream.

[0050] [Figure 13] FIG. 1 is a schematic diagram of a video encoding device.

[0051] [Figure 14] FIG. 1 is a schematic diagram of one embodiment of a means of encoding. DETAILED DESCRIPTION OF THE INVENTION

[0052] While exemplary implementations of one or more embodiments are provided below, it should be understood at the outset that the disclosed systems and / or methods may be implemented using any number of techniques, whether currently known or in existence. The present disclosure should in no way be limited to the exemplary implementations, drawings, and techniques shown below (including the exemplary designs and implementations shown and described herein), but may be modified within the scope of the appended claims together with their full range of equivalents.

[0053] The following terms are defined as follows, unless used herein in a contrary context. Specifically, the following definitions are intended to provide further clarity to the present disclosure. However, such terms may be explained differently in different contexts. Therefore, the following definitions should be considered supplemental to, and not limiting of, any other definitions provided for such terms herein.

[0054] A bitstream is a sequence of bits containing compressed video data for transmission between an encoder and a decoder. An encoder is a device configured to compress video data into a bitstream using an encoding process. A decoder is a device configured to reconstruct the video data in the bitstream for display using a decoding process. A picture is an array of luma samples and / or chroma samples that make up a frame or a field thereof. The picture being encoded or decoded may be referred to as the current picture for clarity.

[0055] A reference picture is a picture containing reference samples that can be used to encode other pictures by reference according to inter-prediction and / or inter-layer prediction. A reference picture list is a list of reference pictures used for inter-prediction and / or inter-layer prediction. Some video coding systems utilize two reference picture lists, which may be represented as Reference Picture List 1 and Reference Picture List 0. A reference picture list structure is an addressable syntax structure that contains multiple reference picture lists. Inter-prediction is a scheme in which samples of a current picture are coded by referencing indicated samples in a reference picture different from the current picture, where the reference picture and the current picture are in the same layer. A reference picture list structure entry is an addressable location in the reference picture list structure that indicates the reference picture associated with the reference picture list.

[0056] A slice header is a part of a coded slice that contains data elements related to all video data within the tile represented by the slice. A picture parameter set (PPS) is a parameter set containing data related to an entire picture. More specifically, a PPS is a syntax structure containing syntax elements that apply to zero or more entire coded pictures, and is determined by the syntax elements found in each picture header. A sequence parameter set (SPS) is a parameter set containing data related to a sequence of pictures. An access unit (AU) is a set of one or more coded pictures associated with the same display time (e.g., the same picture order count) for output from the decoded picture buffer (DPB) (e.g., for display to a user). A decoded video sequence is a sequence of pictures reconstructed by a decoder for display to a user.

[0057] A conformance cropping window (or simply a conformance window) refers to a window of picture samples included in a coded video sequence output from an encoding process. A bitstream may provide conformance window cropping parameters to indicate the output area of ​​a coded picture. Picture width is the width of the picture measured in luma samples. Picture height is the height of the picture measured in luma samples. Conformance window offsets (e.g., conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, conf_win_bottom_offset) specify the picture samples referenced by the PPS output from the decoding process, with a rectangular area specified in output picture coordinates.

[0058] Decoder-side motion vector refinement (DMVR) is a process, algorithm, or coding tool used to refine the motion or motion vector of a predicted block. DMVR allows for a motion vector to be determined based on two motion vectors seen for bi-prediction using a bilateral template matching process. DMVR allows for a weighted combination of the predictive coding unit generated using each of the two motion vectors, and the two motion vectors can be refined by replacing the combined predictive coding unit with a new motion vector that best points to it. Bidirectional optical flow (BDOF) is a process, algorithm, or coding tool used to refine the motion or motion vector of a predicted block. BDOF allows for a motion vector of a sub-coding unit to be determined based on the gradient of the difference between two reference pictures.

[0059] Reference picture resampling (RPR) is the ability to change the spatial resolution of a coded picture mid-bitstream without requiring intra-coding of the picture at the resolution change point. As used herein, resolution refers to the number of pixels contained in a video file. That is, resolution is the width and height of a video, measured in pixels. For example, a video may have a resolution of 1280 (horizontal pixels) by 720 (vertical pixels). This is usually written simply as 1280x720 or abbreviated as 720p.

[0060] Decoder-side motion vector refinement (DMVR) is a process, algorithm, or coding tool used to refine the motion or motion vectors of predicted blocks. Bidirectional optical flow (BDOF), also known as bidirectional optical flow (BIO), is a process, algorithm, or coding tool used to refine the motion or motion vectors of predicted blocks. Reference picture resampling (RPR) is the ability to change the spatial resolution of a coded picture mid-bitstream without requiring intra-coding of the picture at the resolution change point.

[0061] The following acronyms are used in this document: Coding Tree Block (CTB), Coding Tree Unit (CTU), Coding Unit (CU), Coded Video Sequence (CVS), Joint Video Experts Team (JVET), Motion Constrained Tile Set (MCTS), Maximum Transmission Unit (MTU), Network Abstraction Layer (NAL), Picture Order Count (POC), Raw Byte Sequence Payload (RBSP), Sequence Parameter Set (SPS), Versatile Video Coding (VVC), and Working Draft (WD).

[0062] FIG. 1 is a flowchart of an example method 100 of operation for encoding a video signal. Specifically, a video signal is encoded by an encoder. The encoding process compresses the video signal by reducing the video file size using various schemes. The smaller file size allows the compressed video file to be transmitted to a user while reducing the associated bandwidth overhead. A decoder then decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process is generally similar to the encoding process, allowing the decoder to consistently reconstruct the video signal.

[0063] In step 101, a video signal is input to the encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device such as a video camera and encoded to support live streaming of the video. The video file may include both an audio component and a video component. The video component includes a series of image frames that, when viewed in sequence, create the impression of visual movement. These frames include pixels represented by brightness, referred to herein as luma components (or luma samples), and color, referred to herein as chroma components (or color samples). In some examples, the frames may also include depth values ​​to support three-dimensional displays.

[0064] In step 103, the video is divided into blocks. The division includes subdividing pixels within each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC) (also known as H.265 and MPEG-H Part 2), a frame may first be divided into coding tree units (CTUs), which are blocks of a predetermined size (e.g., 64 pixels by 64 pixels). CTUs contain both luma samples and chroma samples. A coding tree may be used to divide the CTUs into blocks and then recursively subdivide these blocks until a configuration is obtained that supports further encoding. For example, the luma component of a frame may be subdivided until each block contains relatively uniform illumination values. Furthermore, the chroma component of a frame may be subdivided until each block contains relatively uniform color values. Thus, the division scheme varies depending on the content of the video frame.

[0065] In step 105, various compression schemes are used to compress the image blocks partitioned in step 103. For example, inter-prediction and / or intra-prediction may be used. Inter-prediction is designed to take advantage of the fact that objects in a common scene tend to appear in consecutive frames. Thus, a block depicting an object in a reference frame need not be repeatedly depicted in adjacent frames. Specifically, an object such as a table may remain in a fixed position across multiple frames. Thus, once the table is depicted, adjacent frames can refer back to the reference frame. Pattern matching techniques may be used to match objects across multiple frames. Furthermore, moving objects may be depicted across multiple frames due to, for example, object motion or camera motion. As a specific example, a video may show a car crossing the screen across multiple frames. Motion vectors may be used to depict such motion. A motion vector is a two-dimensional vector that provides the offset from the object's coordinates in a frame to the object's coordinates in a reference frame. Thus, in inter-prediction, image blocks in a current frame may be encoded as a series of motion vectors indicating the offset from corresponding blocks in a reference frame.

[0066] Intra prediction encodes blocks within a common frame. Intra prediction takes advantage of the fact that luma and chroma components tend to cluster together in a given frame. For example, green areas in a tree tend to be located adjacent to similar green areas. Intra prediction uses several directional prediction modes (e.g., 33 in HEVC), planar mode, and direct current (DC) mode. Directional mode indicates that the current block is similar to samples from neighboring blocks in the corresponding direction. Planar mode indicates that a series of blocks along a row / column (e.g., a plane) can be interpolated based on neighboring blocks at the end of the row. Planar mode effectively indicates a smooth transition in brightness / color along the row / column by using a relatively constant gradient for value change. DC mode is used for boundary smoothing and indicates that a block is similar to the average value associated with samples from all neighboring blocks associated with the angular direction of the directional prediction mode. Therefore, intra-predicted blocks can be represented as values ​​of various related prediction modes instead of actual values. Furthermore, inter-predicted blocks can be represented as values ​​of motion vectors instead of actual values. In either case, the prediction block may not represent the image block exactly: any differences are stored in a residual block, to which multiple transforms may be applied to further compress the file.

[0067] Various filtering techniques may be applied in stage 107. In HEVC, multiple filters are applied according to an in-loop filtering scheme. The block-based prediction described above may produce blocky images at the decoder. Furthermore, the block-based prediction scheme may encode a block and then reconstruct the encoded block for later use as a reference block. In-loop filtering schemes iteratively apply noise suppression filters, deblocking filters, adaptive loop filters, and sample adaptive offset (SAO) filters to blocks / frames. These filters reduce such blocking artifacts so that the encoded file can be accurately reconstructed. Furthermore, these filters reduce artifacts in the reconstructed reference block so that artifacts are less likely to produce further artifacts in subsequent blocks that are encoded based on the reconstructed reference block.

[0068] Once the video signal has been segmented, compressed, and filtered, the resulting data is encoded into a bitstream at step 109. The bitstream includes the data described above as well as any signaling data desired to support proper video signal reconstruction at the decoder. For example, such data may include segmentation data, prediction data, residual blocks, and various flags that provide coding instructions to the decoder. The bitstream may be stored in memory for transmission to the decoder upon request. The bitstream may be broadcast and / or multicast to multiple decoders. Creation of the bitstream is an iterative process. Thus, steps 101, 103, 105, 107, and 109 may be performed sequentially and / or simultaneously for multiple frames and blocks. The order shown in FIG. 1 is presented for clarity and ease of explanation, but is not intended to limit the video encoding process to any particular order.

[0069] The decoder receives the bitstream and begins the decoding process at step 111. Specifically, the decoder uses an entropy decoding scheme to convert the bitstream into corresponding syntax and video data. Using the syntax data from the bitstream, the decoder determines the frame partitioning at step 111. This partitioning should match the block partitioning results from step 103. We now describe the entropy encoding / decoding used in step 111. During the compression process, the encoder makes a number of choices, such as selecting a block partitioning scheme from several possible options based on the spatial arrangement of values ​​contained in the input image. A number of bins may be used to signal the exact choice. As used herein, a bin is a binary value treated as a variable (e.g., a bit value that can change depending on the situation). Entropy coding allows the encoder to discard any options that are clearly not feasible for a particular case, leaving a set of acceptable options. Each acceptable option is then assigned to a codeword. The length of the codeword is based on the number of allowable options (e.g., two options for one bin, three or four options for two bins, etc.). The encoder then encodes the codeword for the selected options. This reduces the size of the codeword because it is desirably large enough to uniquely represent a choice from a small subset of allowable options, as opposed to uniquely representing a choice from a potentially large set of all possible options. The decoder then decodes the selection by determining the set of allowable options in a similar manner to the encoder. By determining the set of allowable options, the decoder can read the codeword and determine the selection made by the encoder.

[0070] In step 113, the decoder performs block decoding. Specifically, the decoder uses an inverse transform to generate a residual block. The decoder then uses the residual block and the corresponding prediction block to reconstruct an image block according to the partition. The prediction block may include both intra-predicted blocks and inter-predicted blocks generated by the encoder in step 105. The reconstructed image block is then placed in a frame of the reconstructed video signal according to the partition data determined in step 111. The syntax of step 113 may be signaled in the bitstream via entropy coding as described above.

[0071] At step 115, the frames of the reconstructed video signal are filtered in a manner similar to step 107 in the encoder. For example, a noise suppression filter, a deblocking filter, an adaptive loop filter, and an SAO filter may be applied to the frames to remove blocking artifacts. Once the frames have been filtered, the video signal may be output to a display at step 117 for viewing by an end user.

[0072] FIG. 2 is a schematic diagram of an example coding-decoding (codec) system 200 for video encoding. Specifically, codec system 200 provides functionality to support an example implementation of operational method 100. Codec system 200 is generalized to show components used in both encoders and decoders. Codec system 200 receives and splits a video signal, as described with respect to steps 101 and 103 of operational method 100, resulting in split video signal 201. When acting as an encoder, codec system 200 then compresses split video signal 201 into an encoded bitstream, as described with respect to steps 105, 107, and 109 of method 100. When acting as a decoder, codec system 200 generates an output video signal from the bitstream, as described with respect to steps 111, 113, 115, and 117 of operational method 100. Codec system 200 includes an overall encoder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded picture buffer component 223, and a header format and context-adaptive binary arithmetic coding (CABAC) component 231. Such components are coupled as shown. In FIG. 2, black lines indicate the movement of data to be encoded / decoded, and dashed lines indicate the movement of control data that controls the operation of other components. All of the components of codec system 200 may be present in an encoder. A decoder may include a subset of the components of codec system 200. For example, the decoder may include an intra-picture prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded picture buffer component 223. These components are described herein.

[0073] The split video signal 201 is a captured video sequence that has been split into blocks of pixels by a coding tree. The coding tree uses various separation modes to subdivide blocks of pixels into smaller blocks of pixels. These blocks can then be further subdivided into smaller blocks. These blocks are sometimes referred to as nodes on the coding tree. Large parent nodes are separated into smaller child nodes. The number of times a node is subdivided is referred to as the depth of the node / coding tree. In some cases, the split blocks can be contained in a coding unit (CU). For example, a CU can be a sub-portion of a CTU that includes a luma block, a red-difference (Cr) block, and a blue-difference (Cb) block, along with corresponding syntax instructions for the CU. Separation modes can include a binary tree (BT), a triple tree (TT), and a quad tree (QT), which are used to split a node into two, three, or four child nodes, respectively, and vary in shape depending on the separation mode used. The split video signal 201 is forwarded to an overall encoder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, a filter control analysis component 227, and a motion estimation component 221 for compression.

[0074] The overall encoder control component 211 is configured to make decisions related to encoding images of a video sequence into a bitstream according to application constraints. For example, the overall encoder control component 211 manages the optimization of bitrate / bitstream size versus reconstruction quality. Such decisions may be based on storage space / bandwidth availability and image resolution requirements. The overall encoder control component 211 also manages buffer utilization, taking transmission speed into account, to reduce buffer underrun and overrun issues. To manage these issues, the overall encoder control component 211 manages segmentation, prediction, and filtering by other components. For example, the overall encoder control component 211 may dynamically increase compression complexity to increase resolution and bandwidth utilization, or decrease compression complexity to decrease resolution and bandwidth utilization. Thus, the overall encoder control component 211 controls other components of the codec system 200 to balance video signal reconstruction quality and bitrate issues. The overall encoder control component 211 generates control data, which controls the operation of other components. The control data is also forwarded to the header format and CABAC component 231 to be encoded into the bitstream signaling parameters for decoding at the decoder.

[0075] The partitioned video signal 201 is also sent to a motion estimation component 221 and a motion compensation component 219 for inter-prediction. A frame or slice of the partitioned video signal 201 may be divided into multiple video blocks. The motion estimation component 221 and the motion compensation component 219 perform inter-predictive coding of the received video blocks by comparing them with one or more blocks in one or more reference frames to provide temporal prediction. The codec system 200 may perform multiple coding passes, for example, to select an appropriate coding mode for each block of video data.

[0076] The motion estimation component 221 and the motion compensation component 219 may be highly integrated but are shown separately for conceptual purposes. Motion estimation, performed by the motion estimation component 221, is the process of generating motion vectors, which estimate the movement of video blocks. For example, a motion vector may indicate the displacement of an encoded object relative to a predictive block. A predictive block is a block that is found to closely match the block being encoded in terms of pixel differences. A predictive block may also be referred to as a reference block. Such pixel differences may be determined by sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics. HEVC uses several coded objects, including CTUs, coding tree blocks (CTBs), and CUs. For example, a CTU may be divided into multiple CTBs, which may then be divided into multiple CBs for inclusion in a CU. A CU may be encoded as a prediction unit (PU), which contains prediction data, and / or a transform unit (TU), which contains transformed residual data for the CU. The motion estimation component 221 generates motion vectors, PUs, and TUs by using rate-distortion analysis as part of a rate-distortion optimization process. For example, the motion estimation component 221 may determine multiple reference blocks, multiple motion vectors, etc. for the current block / frame and may select the reference block, motion vector, etc. with optimal rate-distortion characteristics. The optimal rate-distortion characteristics balance the quality of the video reconstruction (e.g., the amount of data loss due to compression) and the coding efficiency (e.g., the size of the final encoding).

[0077] In some examples, the codec system 200 may calculate values ​​for sub-integer pixel positions of a reference picture stored in the decoded picture buffer component 223. For example, the video codec system 200 may interpolate values ​​for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference picture. Accordingly, the motion estimation component 221 may perform motion search for full-pixel and fractional pixel positions and output fractional-pixel precision motion vectors. The motion estimation component 221 calculates motion vectors for PUs of video blocks in inter-coded slices by comparing the positions of the PUs with the positions of predictive blocks in the reference pictures. The motion estimation component 221 outputs the calculated motion vectors as motion data to the header format and CABAC component 231 for encoding, and also outputs motion to the motion compensation component 219.

[0078] Motion compensation is performed by motion compensation component 219 and may involve fetching or generating a prediction block based on a motion vector determined by motion estimation component 221. Again, in some examples, motion estimation component 221 and motion compensation component 219 may be functionally integrated. Upon receiving a motion vector for the PU of the current video block, motion compensation component 219 may locate the prediction block to which the motion vector points. A residual video block is then formed by subtracting pixel values ​​of the prediction block from pixel values ​​of the current video block being coded, thereby forming pixel difference values. Generally, motion estimation component 221 performs motion estimation on the luma component, and motion compensation component 219 uses the motion vector calculated based on the luma component for both the chroma and luma components. The prediction block and residual block are forwarded to transform scaling and quantization component 213.

[0079] The split video signal 201 is also sent to an intra-picture estimation component 215 and an intra-picture prediction component 217. Like the motion estimation component 221 and the motion compensation component 219, the intra-picture estimation component 215 and the intra-picture prediction component 217 may be highly integrated but are shown separately for conceptual purposes. The intra-picture estimation component 215 and the intra-picture prediction component 217 perform intra-prediction of the current block for each block in the current frame as an alternative to the inter-frame prediction performed by the motion estimation component 221 and the motion compensation component 219, as described above. Specifically, the intra-picture estimation component 215 determines the intra-prediction mode to use to encode the current block. In some examples, the intra-picture estimation component 215 selects an appropriate intra-prediction mode to encode the current block from multiple intra-prediction modes tested. The selected intra-prediction mode is then forwarded to the header format and CABAC component 231 for encoding.

[0080] For example, the intra picture estimation component 215 may use rate-distortion analysis to calculate rate-distortion values ​​for various intra prediction modes to be tested and select an intra prediction mode with optimal rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between an encoded block and the original unencoded block encoded to produce the encoded block, as well as the bit rate (e.g., number of bits) used to produce the encoded block. The intra picture estimation component 215 may calculate ratios from the distortions and rates of the various encoded blocks to determine which intra prediction mode exhibits the optimal rate-distortion value for each block. Furthermore, the intra picture estimation component 215 may be configured to encode depth blocks of the depth map using a rate-distortion optimization (RDO)-based depth modeling mode (DMM).

[0081] The intra-picture prediction component 217 may generate a residual block from the prediction block based on a selected intra-prediction mode determined by the intra-picture estimation component 215 if implemented in an encoder, or may read the residual block from the bitstream if implemented in a decoder. The residual block contains value differences between the prediction block and the original block and is represented as a matrix. The residual block is then forwarded to the transform scaling and quantization component 213. The intra-picture estimation component 215 and the intra-picture prediction component 217 may process both the luma and chroma components.

[0082] The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform, such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform, to the residual block to produce a video block containing residual transform coefficient values. A wavelet transform, an integer transform, a subband transform, or other types of transforms may also be used. This transform may convert the residual information from the pixel value domain to a transform domain, such as the frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling involves applying a scale factor to the residual information, which may cause different frequency information to be quantized with different granularity, thereby affecting the final display quality of the reconstructed video. The transform scaling and quantization component 213 is also configured to quantize the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, the transform scaling and quantization component 213 may then scan the matrix containing the quantized transform coefficients, which are forwarded to the header format and CABAC component 231 for encoding into a bitstream.

[0083] The scaling and inverse transform component 229 applies the inverse operations of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 applies inverse scaling, transform, and / or quantization to reconstruct residual blocks in the pixel domain for later use as reference blocks that may become prediction blocks for other current blocks, for example. The motion estimation component 221 and / or motion compensation component 219 may calculate reference blocks by adding the residual blocks to the corresponding prediction blocks for use in motion estimation of later blocks / frames. A filter is applied to the reconstructed reference blocks to reduce artifacts introduced during scaling, quantization, and transform. Such artifacts could otherwise cause erroneous predictions (and create further artifacts) when predicting subsequent blocks.

[0084] The filter control analysis component 227 and the in-loop filter component 225 apply filters to residual blocks and / or reconstructed image blocks. For example, a transformed residual block from the scaling and inverse transform component 229 may be combined with a corresponding prediction block from the intra-picture prediction component 217 and / or the motion compensation component 219 to reconstruct an original image block. A filter may then be applied to the reconstructed image block. In some examples, a filter may instead be applied to the residual block. Like the other components in FIG. 2, the filter control analysis component 227 and the in-loop filter component 225 are highly integrated and may be implemented together, but are shown separately for conceptual purposes. The filters applied to reconstructed reference blocks are applied to specific spatial regions and include multiple parameters for adjusting how such filters are applied. The filter control analysis component 227 analyzes the reconstructed reference blocks to determine where such filters should be applied and set the corresponding parameters. Such data is forwarded to the header format and CABAC component 231 as filter control data for encoding. The in-loop filter component 225 applies such filters based on the filter control data. The filters may include deblocking filters, noise suppression filters, SAO filters, and adaptive loop filters. Such filters may be applied in the spatial / pixel domain (e.g., to reconstructed pixel blocks) or in the frequency domain, depending on the case.

[0085] When operating as an encoder, the filtered reconstructed image blocks, residual blocks, and / or prediction blocks are stored in the decoded picture buffer component 223 for later use in motion estimation as described above. When operating as a decoder, the decoded picture buffer component 223 stores the reconstructed and filtered blocks and forwards them to a display as part of the output video signal. The decoded picture buffer component 223 may be any memory device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks.

[0086] The header format and CABAC component 231 receives data from various components of the codec system 200 and encodes such data into a coded bitstream for transmission to a decoder. Specifically, the header format and CABAC component 231 generates various headers and encodes control data, such as general control data and filter control data. Additionally, prediction data, including intra-prediction and motion data, and residual data in the form of quantized transform coefficient data are all encoded into the bitstream. The final bitstream contains all the information necessary for a decoder to reconstruct the original split video signal 201. Such information may also include an index table of intra-prediction modes (also called a codeword mapping table), definitions of the encoding contexts of various blocks, an indication of the most likely intra-picture mode, an indication of split information, and so on. Such data may be encoded using entropy coding. For example, such information may be encoded using context-adaptive variable length coding (CAVLC), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding technique. After entropy coding, the coded bitstream may be transmitted to another device (e.g., a video decoder) or archived for later transmission or retrieval.

[0087] 3 is a block diagram illustrating an example of a video encoder 300. The video encoder 300 may be used to implement the encoding functionality of the codec system 200 and / or to perform steps 101, 103, 105, 107, and / or 109 of the method of operation 100. The encoder 300 splits an input video signal to obtain split video signals 301, which are substantially similar to the split video signals 201. The split video signals 301 are then compressed and encoded into a bitstream by components of the encoder 300.

[0088] Specifically, the split video signal 301 is forwarded to an intra-picture prediction component 317 for intra prediction. The intra-picture prediction component 317 may be substantially similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. The split video signal 301 is also forwarded to a motion compensation component 321 for inter prediction based on reference blocks included in a decoded picture buffer component 323. The motion compensation component 321 may be substantially similar to the motion estimation component 221 and the motion compensation component 219. The prediction block and residual block from the intra-picture prediction component 317 and the motion compensation component 321 are forwarded to a transform and quantization component 313 for transforming and quantizing the residual block. The transform and quantization component 313 may be substantially similar to the transform scaling and quantization component 213. The transformed and quantized residual block and the corresponding prediction block (along with associated control data) are forwarded to an entropy coding component 331 for encoding into a bitstream. The entropy coding component 331 may be substantially similar to the header format and CABAC component 231 .

[0089] The transformed and quantized residual block, and / or the corresponding prediction block, are also forwarded from the transform and quantization component 313 to the inverse transform and quantization component 329 for reconstructing into a reference block for use by the motion compensation component 321. The inverse transform and quantization component 329 may be substantially similar to the scaling and inverse transform component 229. An in-loop filter included in the in-loop filter component 325 is also applied to the residual block and / or the reconstructed reference block, as the case may be. The in-loop filter component 325 may be substantially similar to the filter control analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 may include multiple filters, such as those described with respect to the in-loop filter component 225. The filtered block is then stored in the decoded picture buffer component 323 for use as a reference block by the motion compensation component 321. The decoded picture buffer component 323 may be substantially similar to the decoded picture buffer component 223.

[0090] 4 is a block diagram illustrating an example of a video decoder 400. Video decoder 400 may be used to implement the decoding functionality of codec system 200 and / or perform steps 111, 113, 115, and / or 117 of method of operation 100. Decoder 400 receives a bitstream, for example, from encoder 300, and generates a reconstructed output video signal based on the bitstream for display to an end user.

[0091] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 is configured to implement an entropy decoding scheme, such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component 433 may use header information to provide context and interpret additional data encoded as codewords in the bitstream. The decoded information includes any desired information for decoding the video signal, such as overall control data, filter control data, segmentation information, motion data, prediction data, and quantized transform coefficients from the residual blocks. The quantized transform coefficients are forwarded to the inverse transform and quantization component 429 for reconstruction into residual blocks. The inverse transform and quantization component 429 may be similar to the inverse transform and quantization component 329.

[0092] The reconstructed residual block and / or predictive block are forwarded to the intra-picture prediction component 417 for reconstruction into an image block based on an intra-prediction operation. The intra-picture prediction component 417 may be similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. Specifically, the intra-picture prediction component 417 uses the prediction mode to locate a reference block within a frame and applies the residual block to the result to reconstruct an intra-predicted image block. The reconstructed and intra-predicted image block and / or residual block, as well as the corresponding inter-prediction data, are forwarded to the decoded picture buffer component 423 via the in-loop filter component 425. These components may be substantially similar to the in-loop filter component 225 and the decoded picture buffer component 223, respectively. The in-loop filter component 425 filters the reconstructed image block, residual block, and / or predictive block, and such information is stored in the decoded picture buffer component 423. The reconstructed image blocks from the decoded picture buffer component 423 are forwarded to the motion compensation component 421 for inter-prediction. The motion compensation component 421 may be substantially similar to the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 generates a prediction block using a motion vector from a reference block and applies a residual block to the result to reconstruct an image block. The resulting reconstructed block may also be forwarded to the decoded picture buffer component 423 via an in-loop filter component 425. The decoded picture buffer component 423 continues to store further reconstructed image blocks, which may be reconstructed into frames according to the partition information. Such frames may be arranged in a sequence. This sequence is output to a display as a reconstructed output video signal.

[0093] In light of the above, these video compression techniques perform spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or remove redundancy inherent in video sequences. In block-based video coding, a video slice (i.e., a video picture or a portion of a video picture) may be divided into multiple video blocks, which may also be referred to as tree blocks, coding tree blocks (CTBs), coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are encoded using spatial prediction with respect to reference samples contained in neighboring blocks in the same picture. Video blocks in an inter-coded (P or B) slice of a picture may use spatial prediction with respect to reference samples contained in neighboring blocks in the same picture or temporal prediction with respect to reference samples contained in other reference pictures. These pictures may be referred to as frames, and reference pictures may be referred to as reference frames.

[0094] Spatial or temporal prediction results in a prediction block for the block to be coded. Residual data represents pixel differences between the original block to be coded and the prediction block. Inter-coded blocks are encoded according to a motion vector pointing to a block of reference samples forming the prediction block and the residual data indicating the difference between the coded block and the prediction block. Intra-coded blocks are encoded according to an intra-coding mode and the residual data. For further compression, the residual data can be transformed from the pixel domain to a transform domain to obtain residual transform coefficients. The residual transform coefficients may then be quantized. The quantized transform coefficients are initially arranged in a two-dimensional array and may be scanned to produce a one-dimensional vector of transform coefficients, and entropy coding may be applied to achieve further compression.

[0095] Image and video compression is experiencing rapid growth, resulting in a variety of coding standards. These video coding standards include ITU-T H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) MPEG-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, Advanced Video Coding (AVC) (also known as ITU-T H.264 or ISO / IEC MPEG-4 Part 10), and High Efficiency Video Coding (HEVC) (also known as ITU-T H.265 or MPEG-H Part 2). AVC includes extensions such as Scalable Video Coding (SVC), Multiview Video Coding (MVC), and Multiview Video Coding + Depth (MVC+D), as well as 3D AVC (3D-AVC). HEVC includes extension standards such as Scalable HEVC (SHVC), Multi-View HEVC (MV-HEVC), and 3D HEVC (3D-HEVC).

[0096] There is also an emerging video coding standard named Versatile Video Coding (VVC), which is being developed by the Joint Video Experts Team (JVET) of ITU-T and ISO / IEC. There are several working drafts of the VVC standard, but specifically, one working draft (WD) for VVC is referenced herein: "Versatile Video Coding (Draft 5)," JVET-N1001-v3, by B. Bross, J. Chen, and S. Liu (13th JVET Meeting, March 27, 2019) (VVC Draft 5). Each of the references in this and the preceding paragraphs is incorporated by reference in its entirety.

[0097] The techniques described herein are based on Versatile Video Coding (VVC), a video coding standard under development by the Joint Video Experts Team (JVET) of ITU-T and ISO / IEC, although the techniques also apply to other video codec specifications.

[0098] FIG. 5 is a representation 500 of the relationship of an Intra Random Access Point (IRAP) picture 502 to a leading picture 504 and a trailing picture 506 in decoding order 508 and display order 510. In one embodiment, the IRAP picture 502 is referred to as a Clean Random Access (CRA) picture or an Instant Decoder Refresh (IDR) picture with a Random Access Decodable (RADL) picture. In HEVC, IDR pictures, CRA pictures, and Broken Link Access (BLA) pictures are all considered IRAP pictures 502. For VVC, it was agreed at the 12th JVET meeting in October 2018 to have both IDR pictures and CRA pictures as IRAP pictures. In one embodiment, Broken Link Access (BLA) pictures and Gradual Decoder Refresh (GDR) pictures may also be considered IRAP pictures. The decoding process of a coded video sequence always begins with an IRAP.

[0099] 5, leading pictures 504 (e.g., pictures 2 and 3) come after the IRAP picture 502 in decoding order 508 but come before the IRAP picture 502 in display order 510. A trailing picture 506 comes after the IRAP picture 502 in both decoding order 508 and display order 510. Although two leading pictures 504 and one trailing picture 506 are shown in FIG. 5, those skilled in the art will understand that in practical applications, more or fewer leading pictures 504 and / or trailing pictures 506 may be present in the decoding order 508 and display order 510.

[0100] The leading pictures 504 in Figure 5 are divided into two types: random access skip leading (RASL) and RADL. If decoding starts with an IRAP picture 502 (e.g., picture 1), the RADL picture (e.g., picture 3) can be properly decoded. However, the RASL picture (e.g., picture 2) cannot be properly decoded. Therefore, the RASL picture is discarded. Considering the difference between RADL and RASL pictures, the type of the leading picture 504 associated with an IRAP picture 502 must be specified as either RADL or RASL for efficient and appropriate encoding. HEVC constrains that, if a RASL picture and a RADL picture exist, for the RASL picture and RADL picture associated with the same IRAP picture 502, the RASL picture must precede the RADL picture in the display order 510.

[0101] The IRAP picture 502 provides two important functions / advantages: First, the presence of the IRAP picture 502 indicates that the decoding process can start from this picture. This function enables a random access function in which the decoding process starts at a certain position in the bitstream as long as the IRAP picture 502 is present at that position, not necessarily at the beginning of the bitstream. Second, the presence of the IRAP picture 502 refreshes the decoding process so that coded pictures (except for RASL pictures) starting with the IRAP picture 502 are coded without any reference to previous pictures. Therefore, the presence of the IRAP picture 502 in the bitstream prevents any errors that may occur when decoding coded pictures before the IRAP picture 502 from propagating to the IRAP picture 502 and pictures that follow the IRAP picture 502 in decoding order 508.

[0102] While the IRAP picture 502 provides an important function, it comes at a cost to compression efficiency. The presence of the IRAP picture 502 causes a sudden increase in bitrate. This cost to compression efficiency is due to two reasons. First, because the IRAP picture 502 is an intra-predicted picture, the picture itself requires relatively more bits to represent compared to other pictures that are inter-predicted (e.g., the leading picture 504, the trailing picture 506). Second, because the presence of the IRAP picture 502 interrupts temporal prediction (because the decoder refreshes the decoding process, one of the steps of which is to remove previous reference pictures in the decoded picture buffer (DPB)), the IRAP picture 502 reduces the coding efficiency (i.e., requires more bits to represent) of pictures that come after the IRAP picture 502 in decoding order 508, because such pictures do not have reference pictures for inter-predictive coding.

[0103] Among the picture types considered to be IRAP pictures 502, IDR pictures in HEVC are signaled and derivated differently compared to other picture types. Some of the differences are as follows:

[0104] For the signaling and derivation of the Picture Order Count (POC) value of an IDR picture, the most significant bit (MSB) portion of the POC is simply set equal to 0 rather than being derived from the previous significant picture.

[0105] Regarding the signaling information required for reference picture management, the slice header of an IDR picture does not include information that needs to be signaled to support reference picture management. For other picture types (i.e., CRA, trailing, temporal sublayer access (TSA), etc.), information such as the reference picture set (RPS) described below or other forms of similar information (e.g., reference picture list) is required for the reference picture marking process (i.e., the process of determining the status of a reference picture included in the decoded picture buffer (DPB), i.e., whether it is used for reference or not). However, for an IDR picture, there is no need to signal such information, because the presence of the IDR indicates that the decoding process should simply mark all reference pictures included in the DPB as not used for reference.

[0106] In addition to the concept of an IRAP picture, there is also a leading picture associated with the IRAP picture, if one exists. A leading picture is a picture that comes after the associated IRAP picture in decoding order but precedes the IRAP picture in output order. Depending on the coding setting and picture reference structure, leading pictures are further specified into two types. The first type is a leading picture that may not be correctly decoded if the decoding process starts with its associated IRAP picture. This is possible because such a leading picture is coded with reference to a picture that comes before the IRAP picture in decoding order. Such a leading picture is called random access skip leading (RASL). The second type is a leading picture that must be correctly decoded even if the decoding process starts with its associated IRAP picture. This is possible because such a leading picture is coded without directly or indirectly referencing a picture that comes before the IRAP picture in decoding order. Such a leading picture is called random access decodable leading (RADL). In HEVC, if an RASL picture and a RADL picture exist, the constraint is that for RASL and RADL pictures associated with the same IRAP picture, the RASL picture must precede the RADL picture in output order.

[0107] In HEVC and VVC, the IRAP picture 502 and the leading picture 504 may each be included in a single Network Abstraction Layer (NAL) unit. A set of NAL units is sometimes referred to as an access unit. The IRAP picture 502 and the leading picture 504 are of different given NAL unit types that can be easily identified by system-level applications. For example, a video splicer needs to understand the coded picture types, particularly to identify the IRAP picture 502 from a non-IRAP picture and the leading picture 504 from a trailing picture 506 (including determining RASL and RADL pictures), without needing to understand excessive details of the syntax elements included in the coded bitstream. The trailing picture 506 is a picture associated with the IRAP picture 502 and comes after the IRAP picture 502 in display order 510. A picture may come after a particular IRAP picture 502 in decoding order 508 and may come before any other IRAP picture 502 in decoding order 508. In this regard, giving the IRAP picture 502 and the leading picture 504 their own NAL unit type is useful for such applications.

[0108] For HEVC, the NAL unit types for an IRAP picture include: · BLA with Leading Picture (BLA_W_LP): A NAL unit of a broken link access (BLA) picture that may be followed in decoding order by one or more leading pictures. · BLA with RADL (BLA_W_RADL): A NAL unit of a BLA picture that is followed in decoding order by one or more RADL pictures but may not have any RADL pictures. BLA without leading picture (BLA_N_LP): A NAL unit of a BLA picture that is not followed by a leading picture in decoding order. · IDR with RADL (IDR_W_RADL): A NAL unit of an IDR picture that is followed in decoding order by one or more RADL pictures, but may not have any RADL pictures. IDR without leading picture (IDR_N_LP): A NAL unit of an IDR picture that is not followed by a leading picture in decoding order. · CRA: A NAL unit of a clean random access (CRA) picture that may be followed by a leading picture (i.e., a RASL picture and / or a RADL picture). ·RADL: NAL unit of RADL picture. · RASL: NAL unit of RASL picture.

[0109] For VVC, the NAL unit types of the IRAP picture 502 and the leading picture 504 are as follows: · IDR with RADL (IDR_W_RADL): A NAL unit of an IDR picture that is followed in decoding order by one or more RADL pictures, but may not have any RADL pictures. IDR without leading picture (IDR_N_LP): A NAL unit of an IDR picture that is not followed by a leading picture in decoding order. · CRA: A NAL unit of a clean random access (CRA) picture that may be followed by a leading picture (i.e., a RASL picture and / or a RADL picture). ·RADL: NAL unit of RADL picture. · RASL: NAL unit of RASL picture.

[0110] The reference picture resampling (RPR) function is the ability to change the spatial resolution of a coded picture mid-bitstream without requiring intra-coding of the picture at the resolution change point. To enable this, a picture must be able to reference one or more reference pictures whose spatial resolution differs from that of the current picture for inter-prediction. Therefore, resampling of such reference pictures, or parts thereof, is necessary for encoding and decoding of the current picture. Hence the name RPR. This function is sometimes also called adaptive resolution change (ARC). There are use cases or application scenarios that would benefit from the RPR function, including the following:

[0111] Rate adaptation in video telephony and video conferencing, which allows the coded video to adapt to changing network conditions: when network conditions deteriorate and less bandwidth is available, the encoder may adapt to the conditions by encoding lower resolution pictures.

[0112] Switching the active speaker in a multi-way video conference. In a multi-way video conference, the video size of the active speaker is usually larger or wider than that of the other conference participants. When the active speaker switches, it may also be necessary to adjust the picture resolution for each participant. When active speaker switching occurs frequently, the need for the ARC function becomes more important.

[0113] Faster start of streaming. In streaming applications, it is common for the application to buffer a certain length of decoded picture before starting to display the picture. Starting the bitstream at a lower resolution allows the application to have enough picture in the buffer to start displaying it more quickly.

[0114] Adaptive Stream Switching in Streaming. The Dynamic Adaptive Streaming over HTTP (DASH) specification includes a feature named @mediaStreamStructureId. This feature enables switching between different views at random access points of an open group of pictures (GOP) using a non-decodable leading picture (e.g., in HEVC, a CRA picture with an associated RASL picture). If two different views of the same video have different bitrates but the same spatial resolution and have the same value for @mediaStreamStructureId, switching between the two views at a CRA picture with an associated RASL picture may occur, and the RASL picture associated with the switch at the CRA picture may be decoded with acceptable quality, thus enabling a seamless switch. With ARC, the @mediaStreamStructureId feature could also be used to switch between DASH views of different spatial resolutions.

[0115] Various methods facilitate the basic techniques to support RPR / ARC, such as signaling the list of picture resolutions, some constraints on the resampling of reference pictures in the DPB, etc.

[0116] One component of the approach needed to support RPR is a way to signal the picture resolutions that may be present in the bitstream. This is addressed in some examples by modifying the current signaling of picture resolutions in the list of picture resolutions in the SPS as follows: [Table 1]

[0117] num_pic_size_in_luma_samples_minus1+1 specifies the number of picture sizes (width and height) in units of luma samples that may be present in the coded video sequence.

[0118] pic_width_in_luma_samples[i] specifies the width of the ith decoded picture in units of luma samples that may be present in the coded video sequence. pic_width_in_luma_samples[i] must not be equal to 0 and must be an integer multiple of MinCbSizeY.

[0119] pic_height_in_luma_samples[i] specifies the height of the ith decoded picture in units of luma samples that may be present in the coded video sequence. pic_height_in_luma_samples[i] must not be equal to 0 and must be an integer multiple of MinCbSizeY.

[0120] At the 15th JVET meeting, another variant of signaling picture size and conformance window to support RPR was discussed. The signaling is as follows:

[0121] Signals the maximum picture size (i.e., picture width and picture height) of the SPS.

[0122] Signaling picture size in the Picture Parameter Set (PPS).

[0123] Move the current signaling of the conformance window from SPS to PPS. The conformance window information is used to crop the reconstructed / decoded picture in the process of preparing the picture for output. The cropped picture size is the picture size after the picture has been cropped using its associated conformance window.

[0124] The signaling of the picture size and conformance window is as follows: [Table 2]

[0125] max_width_in_luma_samples specifies that it is a bitstream conformance requirement that pic_width_in_luma_samples of any picture for which this SPS is active be less than or equal to max_width_in_luma_samples.

[0126] max_height_in_luma_samples specifies that it is a bitstream conformance requirement that pic_height_in_luma_samples of any picture for which this SPS is active be less than or equal to max_height_in_luma_samples. [Table 3]

[0127] pic_width_in_luma_samples specifies the width of each decoded picture, in units of luma samples, that references the PPS. pic_width_in_luma_samples must not be equal to 0 and must be an integer multiple of MinCbSizeY.

[0128] pic_height_in_luma_samples specifies the height of each decoded picture, in units of luma samples, referencing the PPS. pic_height_in_luma_samples must not be equal to 0 and must be an integer multiple of MinCbSizeY.

[0129] It is a bitstream conformance requirement that all of the following conditions be met for any active reference picture whose width and height are reference_pic_width_in_luma_samples and reference_pic_height_in_luma_samples: · 2 × pic_width_in_luma_samples ≥ reference_pic_width_in_luma_samples · 2 × pic_height_in_luma_samples ≥ reference_pic_height_in_luma_samples · pic_width_in_luma_samples ≤ 8 × reference_pic_width_in_luma_samples · pic_height_in_luma_samples ≤ 8 × reference_pic_height_in_luma_samples

[0130] The variables PicWidthInCtbsY, PicHeightInCtbsY, PicSizeInCtbsY, PicWidthInMinCbsY, PicHeightInMinCbsY, PicSizeInMinCbsY, PicSizeInSamplesY, PicWidthInSamplesC, and PicHeightInSamplesC are derived as follows. · PicWidthInCtbsY = Ceil(pic_width_in_luma_samples / CtbSizeY) (1) · PicHeightInCtbsY = Ceil(pic_height_in_luma_samples / CtbSizeY) (2) · PicSizeInCtbsY = PicWidthInCtbsY × PicHeightInCtbsY (3) · PicWidthInMinCbsY = pic_width_in_luma_samples / MinCbSizeY (4) · PicHeightInMinCbsY = pic_height_in_luma_samples / MinCbSizeY (5) · PicSizeInMinCbsY = PicWidthInMinCbsY × PicHeightInMinCbsY (6) ·PicSizeInSamplesY=pic_width_in_luma_samples×pic_height_in_luma_samples (7) ·PicWidthInSamplesC=pic_width_in_luma_samples / SubWidthC (8) ·PicHeightInSamplesC=pic_height_in_luma_samples / SubHeightC (9)

[0131] conformance_window_flag equal to 1 indicates that conformance cropping window offset parameters follow in the PPS. conformance_window_flag equal to 0 indicates that conformance cropping window offset parameters are not present.

[0132] conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset refer to the PPS and specify the picture samples output from the decoding process by a rectangular area specified in picture coordinates for output. If conformance_window_flag is equal to 0, the values ​​of conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset are inferred to be equal to 0.

[0133] The conformance cropping window contains luma samples with horizontal picture coordinates (including borders) from [SubWidthC × conf_win_left_offset] to [pic_width_in_luma_samples-(SubWidthC × conf_win_right_offset+1)] and vertical picture coordinates (including borders) from [SubHeightC × conf_win_top_offset] to [pic_height_in_luma_samples-(SubHeightC × conf_win_bottom_offset+1)].

[0134] The value of [SubWidthC × (conf_win_left_offset + conf_win_right_offset)] must be less than pic_width_in_luma_samples, and the value of [SubHeightC × (conf_win_top_offset + conf_win_bottom_offset)] must be less than pic_height_in_luma_samples.

[0135] The variables PicOutputWidthL and PicOutputHeightL are derived as follows: ·PicOutputWidthL=pic_width_in_luma_samples-SubWidthC×(conf_win_right_offset+conf_win_left_offset) (10) ·PicOutputHeightL=pic_height_in_pic_size_units-SubHeightC×(conf_win_bottom_offset+conf_win_top_offset) (11)

[0136] If ChromaArrayType is not equal to 0, the corresponding specified samples of the two chroma arrays are the samples with picture coordinates (x / SubWidthC, y / SubHeightC), where (x, y) are the picture coordinates of the specified luma sample.

[0137] Note: The conformance cropping window offset parameters are only applied on output. All internal decoding processes are applied to the uncropped picture size.

[0138] Signaling picture size and conformance window within the PPS causes the following problems:

[0139] Since multiple PPSs can exist in a Coded Video Sequence (CVS), it is possible that two PPSs can contain signaling for the same picture size but different conformance windows. This leads to a situation where two pictures referencing different PPSs have the same picture size but different cropping sizes.

[0140] For RPR support, it is proposed to turn off some coding tools to code a block if its current and reference pictures have different picture sizes. However, now it is possible that the two pictures have the same picture size but the cropping size may be different, so an additional check based on the cropping size is required.

[0141] Disclosed herein is a technique for constraining picture parameter sets having the same picture size to also have the same conformance window size (e.g., cropping the window size). By maintaining the same conformance window size for picture parameter sets having the same picture size, excessively complex processing can be avoided when reference picture resampling (RPR) is enabled. Therefore, processor, memory, and / or network resource usage may be reduced in both the encoder and the decoder. This improves the coder / decoder (also known as codec) in video coding compared to current codecs. In practical terms, improvements to the video coding process provide users with a more favorable user experience when transmitting, receiving, and / or watching video.

[0142] Scalability in video coding is typically supported using a multi-layer coding scheme. A multi-layer bitstream includes a reference layer (BL) and one or more enhancement layers (EL). Examples of scalability include spatial scalability, quality / signal-to-noise (SNR) scalability, and multiview scalability. When using a multi-layer coding scheme, a picture or a portion thereof may be coded (1) without using a reference picture (i.e., using intra prediction), (2) by referencing a reference picture in the same layer (i.e., using inter prediction), or (3) by referencing a reference picture in another layer (i.e., using inter-layer prediction). A reference picture used for inter-layer prediction of a current picture is called an inter-layer reference picture (ILRP).

[0143] 6 is a schematic diagram illustrating an example of layer-based prediction 600 performed to determine MVs, for example, in block compression stage 105, block decoding stage 113, motion estimation component 221, motion compensation component 219, motion compensation component 321, and / or motion compensation component 421. Layer-based prediction 600 is compatible with unidirectional inter prediction and / or bidirectional inter prediction, but also between pictures of different layers.

[0144] Layer-based prediction 600 is applied between pictures 611, 612, 613, and 614, and pictures 615, 616, 617, and 618, which are in different layers. In the illustrated example, pictures 611, 612, 613, and 614 are part of layer [N+1] 632, and pictures 615, 616, 617, and 618 are part of layer [N] 631. A layer, such as layer [N] 631 and / or layer [N+1] 632, is a group of pictures that are all associated with similar values ​​of characteristics such as similar size, quality, resolution, signal-to-noise ratio, capacity, etc. In the illustrated example, layer [N+1] 632 is associated with a larger image size than layer [N] 631. Thus, in this example, pictures 611, 612, 613, and 614 in layer [N+1] 632 have larger picture sizes (e.g., larger heights and widths, and therefore more samples) than pictures 615, 616, 617, and 618 in layer [N] 631. However, such pictures may be separated between layer [N+1] 632 and layer [N] 631 by other characteristics. Although only two layers, layer [N+1] 632 and layer [N] 631, are shown, a set of pictures may be separated into any number of layers based on associated characteristics. Layer [N+1] 632 and layer [N] 631 may be indicated by a layer ID. A layer ID is an item of data associated with a picture that indicates that the picture is part of the indicated layer. Therefore, each picture 611 to 618 is associated with a corresponding layer ID, which can indicate whether the corresponding picture is included in layer [N+1] 632 or layer [N] 631.

[0145] The pictures 611-618 in different layers 631-632 are configured to be displayed differently. Thus, the pictures 611-618 in different layers 631-632 may share the same temporary identifier (ID) and may be included in the same AU. As used herein, an AU is a set of one or more coded pictures associated with the same display time for output from the DPB. For example, if a small picture is desired, the decoder may decode picture 615 and display it at the current display time. Alternatively, if a large picture is desired, the decoder may decode picture 611 and display it at the current display time. Thus, the pictures 611-614 in the upper layer [N+1] 632 contain substantially the same image data as the corresponding pictures 615-618 in the lower layer [N] 631 (despite the difference in picture size). Specifically, picture 611 contains substantially the same image data as picture 615, picture 612 contains substantially the same image data as picture 616, and so on.

[0146] Pictures 611-618 may be coded by referencing other pictures 611-618 in the same layer [N] 631 or [N+1] 632. Coding a picture with reference to another picture in the same layer results in inter-prediction 623, which is compatible with unidirectional inter-prediction and / or bidirectional inter-prediction. Inter-prediction 623 is indicated by a solid arrow. For example, picture 613 may be coded using inter-prediction 623 with one or two of pictures 611, 612, and / or 614 in layer [N+1] 632 as references, where one picture is referenced for unidirectional inter-prediction and / or two pictures are referenced for bidirectional inter-prediction. Furthermore, picture 617 may be coded using inter-prediction 623 with one or two of pictures 615, 616, and / or 618 in layer [N] 631 as references. Here, one picture is referenced for unidirectional inter prediction, and / or two pictures are referenced for bidirectional inter prediction. When performing inter prediction 623, a picture may be called a reference picture when it is used as a reference for another picture in the same layer. For example, picture 612 may be a reference picture used to encode picture 613 according to inter prediction 623. Inter prediction 623 may also be called intra-layer prediction in a multi-layer context. Thus, inter prediction 623 is a scheme for encoding samples of a current picture by referencing indicated samples in a reference picture different from the current picture, where the reference picture and the current picture are in the same layer.

[0147] Pictures 611-618 may be coded by referencing other pictures 611-618 in different layers. This process is known as inter-layer prediction 621 and is indicated by dashed arrows. Inter-layer prediction 621 is a method of coding samples of a current picture by referencing indicated samples in a reference picture, where the current picture and the reference picture are in different layers and therefore have different layer IDs. For example, a picture included in a lower layer [N] 631 may be used as a reference picture to code a corresponding picture in an upper layer [N+1] 632. As a specific example, picture 611 may be coded by referencing picture 615 according to inter-layer prediction 621. In such a case, picture 615 is used as an inter-layer reference picture. An inter-layer reference picture is a reference picture used in inter-layer prediction 621. In most cases, inter-layer prediction 621 is constrained such that a current picture, such as picture 611, can only use inter-layer reference pictures that are contained in the same AU and at a lower layer, such as picture 615. When multiple layers (e.g., more than two) are available, inter-layer prediction 621 can encode / decode the current picture based on multiple inter-layer reference pictures at a lower level than the current picture.

[0148] A video encoder can use layer-based prediction 600 to encode pictures 611-618 with many different combinations and / or permutations of inter-prediction 623 and inter-layer prediction 621. For example, picture 615 may be coded according to intra-prediction. Pictures 616-618 may then be coded according to inter-prediction 623 by using picture 615 as a reference picture. Furthermore, picture 611 may be coded according to inter-layer prediction 621 by using picture 615 as an inter-layer reference picture. Pictures 612-614 may then be coded according to inter-prediction 623 by using picture 611 as a reference picture. Thus, a reference picture can serve as both a single-layer reference picture and an inter-layer reference picture for different coding schemes. By coding pictures in the higher layer [N+1] 632 based on pictures in the lower layer [N] 631, the use of intra prediction can be avoided in the higher layer [N+1] 632. Intra prediction has much lower coding efficiency than inter prediction 623 and inter layer prediction 621. Therefore, intra prediction, with its low coding efficiency, may be limited to minimum / lowest quality pictures and, therefore, to coding a minimum amount of video data. Pictures used as reference pictures and / or inter layer reference pictures may be indicated in entries of a reference picture list included in a reference picture list structure.

[0149] Previous H.26x-based video coding schemes provide support for scalability in profiles different from those for single-layer coding. Scalable Video Coding (SVC) is a scalable extension of AVC / H.264 that provides support for spatial, temporal, and quality scalability. In SVC, a flag is signaled in each macroblock (MB) in an EL picture to indicate whether the EL MB is predicted using co-located blocks from a lower layer. Predictions from co-located blocks may include texture, motion vectors, and / or coding mode. An SVC implementation cannot directly reuse an unmodified H.264 / AVC implementation design. The EL macroblock syntax and decoding process for SVC differ from those of H.264 / AVC.

[0150] Scalable HEVC (SHVC) is an extension of the HEVC / H.265 standard that provides support for spatial and quality scalability. Multiview HEVC (MV-HEVC) is an extension of HEVC / H.265 that provides support for multiview scalability. 3D HEVC (3D-HEVC) is an extension of HEVC / H.264 that provides support for more advanced and efficient three-dimensional (3D) video coding than MV-HEVC. Note that temporal scalability is included as an integer part of the single-layer HEVC codec. The design of multi-layer extensions to HEVC uses the idea that decoded pictures used for inter-layer prediction only come from the same access unit (AU), are treated as long-term reference pictures (LTRPs), and are assigned reference indices in a reference picture list along with other temporal reference pictures in the current layer. Inter-layer prediction (ILP) is achieved at the prediction unit (PU) level by setting the value of the reference index to refer to an inter-layer reference picture in the reference picture list.

[0151] In particular, both the reference picture resampling function and the spatial scalability function require resampling of the reference picture or a part thereof. Reference picture resampling may be realized at the picture level or the coding block level. However, when RPR is referred to as a coding function, it is a function of single-layer coding. Nevertheless, from a codec design point of view, it is possible, or even preferable, to use the same resampling filter for both the RPR function of single-layer coding and the spatial scalability function of multi-layer coding.

[0152] 7 is a schematic diagram illustrating an example of unidirectional inter prediction 700. The unidirectional inter prediction 700 may be used to determine motion vectors for encoding and / or decoding blocks produced when dividing a picture.

[0153] Unidirectional inter prediction 700 predicts a current block 711 contained in a current frame 710 using a reference frame 730 having a reference block 731. The reference frame 730 may be located temporally after the current frame 710 as shown (e.g., as a subsequent reference frame), but in some examples, it may be located temporally before the current frame 710 (e.g., as a preceding reference frame). The current frame 710 is an example of a frame / picture being encoded / decoded at a particular time. The current frame 710 includes an object in the current block 711 that matches an object in the reference block 731 of the reference frame 730. The reference frame 730 is a frame used as a reference for encoding the current frame 710, and the reference block 731 is a block in the reference frame 730 that includes an object that is also included in the current block 711 of the current frame 710.

[0154] The current block 711 is any coding unit being encoded / decoded at a particular stage of the encoding process. The current block 711 may be an entire partitioned block or a sub-block when using affine inter-prediction mode. The current frame 710 is separated from the reference frame 730 by a certain temporal distance (TD) 733. The TD 733 indicates the time between the current frame 710 and the reference frame 730 in the video sequence and may be measured in units of frames. The prediction information for the current block 711 may reference the reference frame 730 and / or the reference block 731 by a reference index that indicates the direction and temporal distance between the frames. During the time period represented by the TD 733, an object in the current block 711 moves from one position in the current frame 710 to another position in the reference frame 730 (e.g., the position of the reference block 731). For example, the object may move along a motion trajectory 713, which is the direction of the object's movement over time. A motion vector 735 represents the direction and magnitude of the object's movement along the motion trajectory 713 at the TD 733. Therefore, the encoded motion vector 735, the reference block 731, and the residual, which includes the difference between the current block 711 and the reference block 731, provide sufficient information to reconstruct the current block 711 and place it in the current frame 710.

[0155] 8 is a schematic diagram illustrating an example of bidirectional inter prediction 800. The bidirectional inter prediction 800 may be used to determine motion vectors for encoding and / or decoding blocks produced when dividing a picture.

[0156] Bidirectional inter prediction 800 is similar to unidirectional inter prediction 700, but uses a pair of reference frames to predict a current block 811 contained in a current frame 810. Thus, current frame 810 and current block 811 are substantially similar to current frame 710 and current block 711, respectively. Current frame 810 is temporally positioned between a previous reference frame 820, which precedes current frame 810 in the video sequence, and a subsequent reference frame 830, which follows current frame 810 in the video sequence. Apart from that, previous reference frame 820 and subsequent reference frame 830 are substantially similar to reference frame 730.

[0157] A current block 811 matches a previous reference block 821 in a previous reference frame 820 and a subsequent reference block 831 in a subsequent reference frame 830. Such a match indicates that an object moves from a location in the previous reference block 821 to a location in the subsequent reference block 831 along a motion trajectory 813 via the current block 811 over the course of the video sequence. The current frame 810 is separated from the previous reference frame 820 by a certain previous time distance (TD0) 823 and from the subsequent reference frame 830 by a certain subsequent time distance (TD1) 833. TD0 (823) indicates the time in frames between the previous reference frame 820 and the current frame 810 in the video sequence. TD1 (833) indicates the time in frames between the current frame 810 and the subsequent reference frame 830 in the video sequence. Thus, the object moves from the previous reference block 821 to the current block 811 along the motion trajectory 813 over the period indicated by TD0 (823). The object also moves from the current block 811 to the subsequent reference block 831 along the motion trajectory 813 over a period indicated by TD1 (833). The prediction information of the current block 811 may reference a previous reference frame 820 and / or a previous reference block 821, and a subsequent reference frame 830 and / or a subsequent reference block 831, by a pair of reference indices indicating the direction and time distance between the frames.

[0158] The previous motion vector (MV0) 825 represents the direction and magnitude of object motion across TD0 (823) (e.g., between the previous reference frame 820 and the current frame 810) along the motion trajectory 813. The subsequent motion vector (MV1) 835 represents the direction and magnitude of object motion across TD1 (833) (e.g., between the current frame 810 and the subsequent reference frame 830) along the motion trajectory 813. Thus, in bidirectional inter prediction 800, the current block 811 can be coded and reconstructed using the previous reference block 821 and / or the subsequent reference block 831, MV0 (825), and MV1 (835).

[0159] In one embodiment, inter prediction and / or bidirectional inter prediction may be performed sample by sample (e.g., pixel by pixel) instead of block by block. That is, a motion vector pointing to each sample included in the previous reference block 821 and / or the subsequent reference block 831 may be determined for each sample included in the current block 811. In such an embodiment, the motion vector 825 and the motion vector 835 shown in FIG. 8 represent multiple motion vectors corresponding to multiple samples included in the current block 811, the previous reference block 821, and the subsequent reference block 831.

[0160] In both merge mode and advanced motion vector prediction (AMVP) mode, a candidate list is generated by adding candidate motion vectors to the candidate list in an order defined by a candidate list determination pattern. Such candidate motion vectors may include motion vectors generated by unidirectional inter prediction 700, bidirectional inter prediction 800, or a combination thereof. Specifically, these motion vectors are generated for neighboring blocks when such blocks are encoded. Such motion vectors are added to a candidate list for a current block, from which a motion vector for the current block is selected. The motion vector may then be signaled as the index of the selected motion vector in the candidate list. A decoder can build the candidate list using the same process as an encoder and determine the selected motion vector from the candidate list based on the signaled index. Thus, the candidate motion vectors include motion vectors generated according to unidirectional inter prediction 700 and / or bidirectional inter prediction 800, depending on which technique is used when encoding such neighboring blocks.

[0161] 9 illustrates a video bitstream 900. As used herein, the video bitstream 900 may be referred to as a coded video bitstream, a bitstream, or variations thereof. As shown in FIG. 9, the bitstream 900 includes a sequence parameter set (SPS) 902, a picture parameter set (PPS) 904, a slice header 906, and image data 908.

[0162] The SPS 902 contains data common to all pictures in a sequence of pictures (SOP). In contrast, the PPS 904 contains data common to the entire picture. The slice header 906 contains information about the current slice, such as the slice type and which reference pictures are used. The SPS 902 and PPS 904 are sometimes commonly referred to as parameter sets. The SPS 902, PPS 904, and slice header 906 are types of Network Abstraction Layer (NAL) units. A NAL unit is a syntax structure that contains an indication of the type of data (e.g., coded video data) that follows. NAL units are classified as video coding layer (VCL) NAL units and non-VCL NAL units. VCL NAL units contain data representing the values ​​of samples contained in a video picture, and non-VCL NAL units contain any associated additional information, such as parameter sets (important header data that may apply to many VCL NAL units) and auxiliary enhancement information (timing information and other auxiliary data that may increase the usefulness of the decoded video signal but are not necessary for decoding the sample values ​​contained in a video picture). Those skilled in the art will understand that bitstream 900 may contain other parameters and information as appropriate for practical applications.

[0163] The image data 908 in Figure 9 includes data associated with the image or video being encoded or decoded. The image data 908 may simply be referred to as the payload or data being carried in the bitstream 900. In one embodiment, the image data 908 includes a CVS 914 (or CLVS) that includes multiple pictures 910. The CVS 914 is a coded video sequence of every coded layer video sequence (CLVS) included in the video bitstream 900. Notably, the CVS and CLVS are the same if the video bitstream 900 includes a single layer. The CVS and CLVS are different only if the video bitstream 900 includes multiple layers.

[0164] 9, each slice of a picture 910 may be contained in its own VCL NAL unit 912. The set of VCL NAL units 912 contained in a CVS 914 may be referred to as an access unit.

[0165] FIG. 10 illustrates a division technique 1000 for a picture 1010. The picture 1010 may be similar to any of the pictures 910 illustrated in FIG. 9. As illustrated, the picture 1010 may be divided into multiple slices 1012. A slice is a spatially distinct region of a frame (e.g., a picture) that is encoded separately from any other region within the same frame. Although three slices 1012 are illustrated in FIG. 10, more or fewer slices may be used in practical applications. Each slice 1012 may be divided into multiple blocks 1014. The blocks 1014 in FIG. 10 may be similar to the current block 811, the previous reference block 821, and the subsequent reference block 831 in FIG. 8. The blocks 1014 may represent CUs. Although four blocks 1014 are illustrated in FIG. 10, more or fewer blocks may be used in practical applications.

[0166] Each block 1014 may be divided into multiple samples 1016 (e.g., pixels). In one embodiment, the size of each block 1014 is measured in luma samples. Although 16 samples 1016 are shown in Figure 10, more or fewer samples may be used in practical applications.

[0167] In one embodiment, a conformance window 1060 is applied to the picture 1010. As described above, the conformance window 1060 is used to crop, reduce, or otherwise change the size of the picture 1010 (e.g., a reconstructed / decoded picture) in the process of preparing the picture for output. For example, a decoder may apply the conformance window 1060 to the picture 1010 to crop, trim, reduce, or otherwise change the size of the picture 1010 before the picture is output for display to a user. The size of the conformance window 1060 is determined by applying a conformance window top offset 1062, a conformance window bottom offset 1064, a conformance window left offset 1066, and a conformance window right offset 1068 to the picture 1010 to reduce the size of the picture 1010 before output. That is, only the portion of picture 1010 that lies within conformance window 1060 is output. Thus, picture 1010 is cropped in size before being output. In one embodiment, the first picture parameter set and the second picture parameter set each reference the same sequence parameter set and have the same values ​​for picture width and picture height. Thus, the first picture parameter set and the second picture parameter set also have the same value for the conformance window.

[0168] FIG. 11 illustrates one embodiment of a decoding method 1100 implemented by a video decoder (e.g., video decoder 400). Method 1100 may be performed after a bitstream to be decoded is received directly or indirectly from a video encoder (e.g., video encoder 300). Method 1100 improves the decoding process by maintaining the same conformance window size for picture parameter sets with the same picture size. Thus, reference picture resampling (RPR) may remain enabled or on for the entire CVS. Maintaining a consistent conformance window size for picture parameter sets with the same picture size can improve coding efficiency. Therefore, in practice, codec performance improves, which translates into a more desirable user experience.

[0169] In block 1102, a video decoder receives a first picture parameter set (e.g., ppsA) and a second picture parameter set (e.g., ppsB), each of which references the same sequence parameter set. If the first picture parameter set and the second picture parameter set have the same values ​​for picture width and picture height, the first picture parameter set and the second picture parameter set have the same value for conformance window. In one embodiment, the picture width and picture height are measured in luma samples.

[0170] In one embodiment, the picture width is specified as pic_width_in_luma_samples. In one embodiment, the picture height is specified as pic_height_in_luma_samples. In one embodiment, pic_width_in_luma_samples specifies the width, in units of luma samples, of each decoded picture that references a PPS. In one embodiment, pic_height_in_luma_samples specifies the height, in units of luma samples, of each decoded picture that references a PPS.

[0171] In one embodiment, the conformance window has a conformance window left offset, a conformance window right offset, a conformance window top offset, and a conformance window bottom offset, which collectively represent the conformance window size. In one embodiment, the conformance window left offset is specified as pps_conf_win_left_offset. In one embodiment, the conformance window right offset is specified as pps_conf_win_right_offset. In one embodiment, the conformance window top offset is specified as pps_conf_win_top_offset. In one embodiment, the conformance window bottom offset is specified as pps_conf_win_bottom_offset. In one embodiment, the conformance window size or values ​​are signaled in the PPS.

[0172] In block 1104, the video decoder applies a conformance window to the current picture corresponding to the first picture parameter set or the second picture parameter set. By doing so, the video encoder crops the current picture to the size of the conformance window.

[0173] In one embodiment, the method further comprises decoding the current picture based on the resampled reference picture using inter prediction. In one embodiment, the method further comprises resampling the reference picture corresponding to the current picture using reference picture resampling (RPS). In one embodiment, the resampling of the reference picture changes the resolution of the reference picture.

[0174] In one embodiment, the method further comprises determining whether bidirectional optical flow (BDOF) is enabled for decoding the picture based on the picture width, picture height, and conformance window of the current picture and the reference pictures of the current picture. In one embodiment, the method further comprises determining whether decoder-side motion vector refinement (DMVR) is enabled for decoding the picture based on the picture width, picture height, and conformance window of the current picture and the reference pictures of the current picture.

[0175] In one embodiment, the method further comprises displaying the image generated using the current block on a display of an electronic device (e.g., a smartphone, a tablet, a laptop, a personal computer, etc.).

[0176] 12 illustrates one embodiment of a method 1200 for encoding a video bitstream, performed by a video encoder (e.g., video encoder 300). Method 1200 may be performed when pictures (e.g., from a video) are encoded into a video bitstream and then transmitted to a video decoder (e.g., video decoder 400). Method 1200 improves the encoding process by maintaining the same conformance window size for picture parameter sets with the same picture size. Thus, reference picture resampling (RPR) may remain enabled or on for the entire CVS. Maintaining a consistent conformance window size for picture parameter sets with the same picture size can improve coding efficiency. Therefore, in practice, codec performance improves, which translates into a more desirable user experience.

[0177] In block 1202, a video encoder generates a first picture parameter set and a second picture parameter set, each of which references the same sequence parameter set. If the first picture parameter set and the second picture parameter set have the same values ​​for picture width and picture height, the first picture parameter set and the second picture parameter set have the same value for conformance window. In one embodiment, the picture width and picture height are measured in luma samples.

[0178] In one embodiment, the picture width is specified as pic_width_in_luma_samples. In one embodiment, the picture height is specified as pic_height_in_luma_samples. In one embodiment, pic_width_in_luma_samples specifies the width, in units of luma samples, of each decoded picture that references a PPS. In one embodiment, pic_height_in_luma_samples specifies the height, in units of luma samples, of each decoded picture that references a PPS.

[0179] In one embodiment, the conformance window has a conformance window left offset, a conformance window right offset, a conformance window top offset, and a conformance window bottom offset, which collectively represent the conformance window size. In one embodiment, the conformance window left offset is specified as pps_conf_win_left_offset. In one embodiment, the conformance window right offset is specified as pps_conf_win_right_offset. In one embodiment, the conformance window top offset is specified as pps_conf_win_top_offset. In one embodiment, the conformance window bottom offset is specified as pps_conf_win_bottom_offset. In one embodiment, the conformance window size or values ​​are signaled in the PPS.

[0180] At block 1204, the video encoder encodes the first picture parameter set and the second picture parameter set into a video bitstream. At block 1206, the video encoder stores the video bitstream for transmission to the video decoder. In one embodiment, the video encoder transmits the video bitstream including the first picture parameter set and the second picture parameter set to the video decoder.

[0181] In one embodiment, a method for encoding a video bitstream is provided. The bitstream has a plurality of parameter sets and a plurality of pictures. Each picture of the plurality of pictures includes a plurality of slices. Each slice of the plurality of slices includes a plurality of coding blocks. The method includes generating and writing to the bitstream a parameter set, parameterSetA, including information including a picture size, picSizeA, and a conformance window, confWinA. The parameters may be a Picture Parameter Set (PPS). The method further includes generating and writing to the bitstream another parameter set, parameterSetB, including information including a picture size, picSizeB, and a conformance window, confWinB. The parameters may be a Picture Parameter Set (PPS). The method further includes constraining the values ​​of conformance windows confWinA included in parameterSetA and confWinB included in parameterSetB to be the same when the values ​​of picSizeA included in parameterSetA and picSizeB included in parameterSetB are the same, and constraining the values ​​of picture sizes picSizeA included in parameterSetA and picSizeB included in parameterSetB to be the same when the values ​​of confWinA included in parameterSetA and confWinB included in parameterSetB are the same. The method further includes encoding the bitstream.

[0182] In one embodiment, a method for decoding a video bitstream is provided. The bitstream has a plurality of parameter sets and a plurality of pictures. Each picture of the plurality of pictures includes a plurality of slices. Each slice of the plurality of slices includes a plurality of coding blocks. The method includes analyzing the parameter set to obtain a picture size and a conformance window size associated with a current picture, currPic. The obtained information is used to derive a picture size and a cropped size of the current picture. The method further includes analyzing another parameter set to obtain a picture size and a conformance window size associated with a reference picture, refPic. The obtained information is used to derive a picture size and a cropped size of the reference picture. The method further includes determining refPic as a reference picture for decoding a current block curBlock located within a current picture currPic, determining whether bidirectional optical flow (BDOF) is used or enabled to decode the current coding block based on the picture sizes and conformance windows of the current picture and the reference picture, and decoding the current block.

[0183] In one embodiment, if the picture sizes and conformance windows of the current and reference pictures are different, BDOF is not used or is invalid for decoding the current coding block.

[0184] In one embodiment, a method for decoding a video bitstream is provided. The bitstream has a plurality of parameter sets and a plurality of pictures. Each picture of the plurality of pictures includes a plurality of slices. Each slice of the plurality of slices includes a plurality of coding blocks. The method includes analyzing the parameter set to obtain a picture size and a conformance window size associated with a current picture, currPic. The obtained information is used to derive a picture size and a cropped size of the current picture. The method further includes analyzing another parameter set to obtain a picture size and a conformance window size associated with a reference picture, refPic. The obtained information is used to derive a picture size and a cropped size of the reference picture. The method further includes determining refPic as a reference picture for decoding a current block curBlock located within a current picture currPic, determining whether decoder-side motion vector refinement (DMVR) is used or enabled to decode the current coding block based on the picture sizes and conformance windows of the current picture and the reference picture, and decoding the current block.

[0185] In one embodiment, if the picture sizes and conformance windows of the current and reference pictures are different, the DMVR is not used or is invalid for decoding the current coding block.

[0186] In one embodiment, a method for encoding a video bitstream is provided. In one embodiment, the bitstream has a plurality of parameter sets and a plurality of pictures. Each picture of the plurality of pictures includes a plurality of slices. Each slice of the plurality of slices includes a plurality of coding blocks. The method includes generating a parameter set including a picture size and a conformance window size associated with a current picture, currPic. This information is used to derive a picture size and a cropped size of the current picture. The method further includes generating another parameter set including a picture size and a conformance window size associated with a reference picture, refPic. The obtained information is used to derive a picture size and a cropped size of the reference picture. The method further includes constraining temporal motion vector prediction (TMVP) of all slices belonging to the current picture, currPic, such that the reference picture, refPic, must not be used as a co-located reference picture if the picture sizes and conformance windows of the current picture and the reference picture are different. That is, if a reference picture refPic is a co-located reference picture for encoding a block included in a current picture currPic for TMVP, the picture size and conformance window of the current picture and the reference picture must be the same. The method further includes decoding the bitstream.

[0187] In one embodiment, a method for decoding a video bitstream is provided. The bitstream has a plurality of parameter sets and a plurality of pictures. Each picture of the plurality of pictures includes a plurality of slices. Each slice of the plurality of slices includes a plurality of coding blocks. The method includes analyzing the parameter set to obtain a picture size and a conformance window size associated with a current picture, currPic. The obtained information is used to derive a picture size and a cropped size of the current picture. The method further includes analyzing another parameter set to obtain a picture size and a conformance window size associated with a reference picture, refPic. The obtained information is used to derive a picture size and a cropped size of the reference picture. The method further includes determining refPic as a reference picture for decoding a current block curBlock located in a current picture currPic, and analyzing a syntax element (slice_DVMR_BDOF_enable_flag) to determine whether decoder-side motion vector refinement (DMVR) and / or bidirectional optical flow (BDOF) are used or enabled for decoding the current coded picture and slice. The method further includes constraining the value of the syntax element (slice_DVMR_BDOF_enable_flag) to be zero if a conformance window confWinA included in parameterSetA and a conformance window confWinB included in parameterSetB are not the same, or if the values ​​of picSizeA included in parameterSetA and picSizeB included in parameterSetB are not the same.

[0188] The following explanations relate to the base text, which is a VVC Working Draft, i.e., only increments are shown, and any text not mentioned below that is included in the base text applies as is. Deleted text is shown in italics, and added text is in bold.

[0189] The syntax and semantics of the sequence parameter set are provided. [Table 4]

[0190] max_width_in_luma_samples specifies that it is a bitstream conformance requirement that pic_width_in_luma_samples of any picture for which this SPS is active be less than or equal to max_width_in_luma_samples.

[0191] max_height_in_luma_samples specifies that it is a bitstream conformance requirement that pic_height_in_luma_samples of any picture for which this SPS is active be less than or equal to max_height_in_luma_samples.

[0192] The syntax and semantics of picture parameter sets are provided. [Table 5]

[0193] pic_width_in_luma_samples specifies the width of each decoded picture, in units of luma samples, that references the PPS. pic_width_in_luma_samples must not be equal to 0 and must be an integer multiple of MinCbSizeY.

[0194] pic_height_in_luma_samples specifies the height of each decoded picture, in units of luma samples, referencing the PPS. pic_height_in_luma_samples must not be equal to 0 and must be an integer multiple of MinCbSizeY.

[0195] It is a bitstream conformance requirement that all of the following conditions be met for any active reference picture whose width and height are reference_pic_width_in_luma_samples and reference_pic_height_in_luma_samples:

[0196] ·2×pic_width_in_luma_samples≧reference_pic_width_in_luma_samples

[0197] ·2×pic_height_in_luma_samples≧reference_pic_height_in_luma_samples

[0198] ·pic_width_in_luma_samples≦8×reference_pic_width_in_luma_samples

[0199] ·pic_height_in_luma_samples≦8×reference_pic_height_in_luma_samples

[0200] The variables PicWidthInCtbsY, PicHeightInCtbsY, PicSizeInCtbsY, PicWidthInMinCbsY, PicHeightInMinCbsY, PicSizeInMinCbsY, PicSizeInSamplesY, PicWidthInSamplesC, and PicHeightInSamplesC are derived as follows:

[0201] ·PicWidthInCtbsY=Ceil(pic_width_in_luma_samples / CtbSizeY) (1)

[0202] ·PicHeightInCtbsY=Ceil(pic_height_in_luma_samples / CtbSizeY) (2)

[0203] ·PicSizeInCtbsY=PicWidthInCtbsY×PicHeightInCtbsY (3)

[0204] ·PicWidthInMinCbsY=pic_width_in_luma_samples / MinCbSizeY (4)

[0205] ·PicHeightInMinCbsY=pic_height_in_luma_samples / MinCbSizeY (5)

[0206] ·PicSizeInMinCbsY=PicWidthInMinCbsY×PicHeightInMinCbsY (6)

[0207] ·PicSizeInSamplesY=pic_width_in_luma_samples×pic_height_in_luma_samples (7)

[0208] ·PicWidthInSamplesC=pic_width_in_luma_samples / SubWidthC (8)

[0209] ·PicHeightInSamplesC=pic_height_in_luma_samples / SubHeightC (9)

[0210] conformance_window_flag equal to 1 indicates that conformance cropping window offset parameters follow in the PPS. conformance_window_flag equal to 0 indicates that conformance cropping window offset parameters are not present.

[0211] conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset refer to the PPS and specify the picture samples output from the decoding process by a rectangular area specified in picture coordinates for output. If conformance_window_flag is equal to 0, the values ​​of conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset are inferred to be equal to 0.

[0212] The conformance cropping window contains luma samples with horizontal picture coordinates (including borders) from [SubWidthC × conf_win_left_offset] to [pic_width_in_luma_samples-(SubWidthC × conf_win_right_offset+1)] and vertical picture coordinates (including borders) from [SubHeightC × conf_win_top_offset] to [pic_height_in_luma_samples-(SubHeightC × conf_win_bottom_offset+1)].

[0213] The value of [SubWidthC × (conf_win_left_offset + conf_win_right_offset)] must be less than pic_width_in_luma_samples, and the value of [SubHeightC × (conf_win_top_offset + conf_win_bottom_offset)] must be less than pic_height_in_luma_samples.

[0214] The variables PicOutputWidthL and PicOutputHeightL are derived as follows:

[0215] ·PicOutputWidthL=pic_width_in_luma_samples-SubWidthC×(conf_win_right_offset+conf_win_left_offset) (10)

[0216] ·PicOutputHeightL=pic_height_in_pic_size_units-SubHeightC×(conf_win_bottom_offset+conf_win_top_offset) (11)

[0217] If ChromaArrayType is not equal to 0, the corresponding specified samples of the two chroma arrays are the samples with picture coordinates (x / SubWidthC, y / SubHeightC), where (x, y) are the picture coordinates of the specified luma sample.

[0218] Note: The conformance cropping window offset parameters are only applied on output. All internal decoding processes are applied to the uncropped picture size.

[0219] If PPS_A and PPS_B are picture parameter sets that reference the same sequence parameter set, and the values ​​of pic_width_in_luma_samples contained in PPS_A and PPS_B are the same, and the values ​​of pic_height_in_luma_samples contained in PPS_A and PPS_B are the same, then the bitstream conformance requirement is that all of the following conditions must be true: The values ​​of conf_win_left_offset included in PPS_A and PPS_B are the same. The values ​​of conf_win_right_offset included in PPS_A and PPS_B are the same. The values ​​of conf_win_top_offset contained in PPS_A and PPS_B are the same. -The values ​​of conf_win_bottom_offset included in PPS_A and PPS_B are the same.

[0220] The following constraints are imposed on the semantics of collocated_ref_idx:

[0221] collocated_ref_idx specifies the reference index of the co-located picture used for temporal motion vector prediction.

[0222] If slice_type is equal to P, or if slice_type is equal to B and collocated_from_l0_flag is equal to 1, collocated_ref_idx refers to a picture in list 0, and the value of collocated_ref_idx must be in the range 0 to (NumRefIdxActive[0]-1) (inclusive).

[0223] If slice_type is equal to B and collocated_from_l0_flag is equal to 0, collocated_ref_idx refers to a picture in list 1 and the value of collocated_ref_idx must be in the range 0 to (NumRefIdxActive[1]-1), inclusive.

[0224] If collocated_ref_idx is not present, the value of collocated_ref_idx is inferred to be equal to 0.

[0225] It is a bitstream conformance requirement that the picture referenced by collocated_ref_idx must be common to all slices of a coded picture.

[0226] It is a bitstream conformance requirement that the resolution of the reference picture referenced by collocated_ref_idx and the current picture must be the same.

[0227] It is a bitstream conformance requirement that the picture size and conformance window of the reference picture referenced by collocated_ref_idx and the current picture must be the same.

[0228] The following conditions for setting dmvrFlag to 1 are modified:

[0229] dmvrFlag is set equal to 1 if all of the following conditions are true:

[0230] sps_dmvr_enabled_flag is equal to 1.

[0231] general_merge_flag[xCb][yCb] is equal to 1.

[0232] · predFlagL0[0][0] and predFlagL1[0][0] are both equal to 1.

[0233] mmvd_merge_flag[xCb][yCb] is equal to 0.

[0234] ·DiffPicOrderCnt(currPic,RefPicList[0][refIdxL0]) is equal to DiffPicOrderCnt(RefPicList[1][refIdxL1],currPic).

[0235] ·BcwIdx[xCb][yCb] is equal to 0.

[0236] ·luma_weight_l0_flag[refIdxL0] and luma_weight_l1_flag[refIdxL1] are both equal to 0.

[0237] cbWidth is greater than or equal to 8.

[0238] ·cbHeight is greater than or equal to 8.

[0239] cbHeight × cbWidth is greater than or equal to 128.

[0240] · If X is 0 and 1, respectively, pic_width_in_luma_samples and pic_height_in_luma_samples of the reference picture refPicLX associated with refIdxLX are equal to pic_width_in_luma_samples and pic_height_in_luma_samples of the current picture, respectively.

[0241] · If X is 0 and 1, respectively, pic_width_in_luma_samples, pic_height_in_luma_samples, conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset of the reference picture refPicLX associated with refIdxLX are equal to pic_width_in_luma_samples, pic_height_in_luma_samples, conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset of the current picture, respectively.

[0242] The following conditions for setting dmvrFlag to 1 are modified:

[0243] bdofFlag is set equal to TRUE if all of the following conditions are true:

[0244] sps_bdof_enabled_flag is equal to 1.

[0245] · predFlagL0[xSbIdx][ySbIdx] and predFlagL1[xSbIdx][ySbIdx] are both equal to 1.

[0246] DiffPicOrderCnt(currPic,RefPicList[0][refIdxL0]) × DiffPicOrderCnt(currPic,RefPicList[1][refIdxL1]) is less than 0.

[0247] MotionModelIdc[xCb][yCb] is equal to 0.

[0248] merge_subblock_flag[xCb][yCb] is equal to 0.

[0249] sym_mvd_flag[xCb][yCb] is equal to 0.

[0250] ·BcwIdx[xCb][yCb] is equal to 0.

[0251] ·luma_weight_l0_flag[refIdxL0] and luma_weight_l1_flag[refIdxL1] are both equal to 0.

[0252] ·cbHeight is greater than or equal to 8.

[0253] · If X is 0 and 1, respectively, pic_width_in_luma_samples and pic_height_in_luma_samples of the reference picture refPicLX associated with refIdxLX are equal to pic_width_in_luma_samples and pic_height_in_luma_samples of the current picture, respectively.

[0254] · If X is 0 and 1, respectively, pic_width_in_luma_samples, pic_height_in_luma_samples, conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset of the reference picture refPicLX associated with refIdxLX are equal to pic_width_in_luma_samples, pic_height_in_luma_samples, conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset of the current picture, respectively.

[0255] cIdx is equal to 0.

[0256] 13 is a schematic diagram of a video encoding device 1300 (e.g., video encoder 20 or video decoder 30) according to one embodiment of the present disclosure. The video encoding device 1300 is suitable for implementing the disclosed embodiments as described herein. The video encoding device 1300 includes an ingress port 1310 and a receiving unit (Rx) 1320 for receiving data, a processor, logic unit, or central processing unit (CPU) 1330 for processing the data, a transmitting unit (Tx) 1340 and an egress port 1350 for transmitting the data, and a memory 1360 for storing the data. The video encoding device 1300 may also include optical-to-electrical (OE) and electrical-to-optical (EO) conversion components for the egress or ingress of optical or electrical signals coupled to the ingress port 1310, the receiving unit 1320, the transmitting unit 1340, and the egress port 1350.

[0257] The processor 1330 is implemented by hardware and software. The processor 1330 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor 1330 communicates with the ingress port 1310, the receiving unit 1320, the transmitting unit 1340, the egress port 1350, and the memory 1360. The processor 1330 includes an encoding module 1370. The encoding module 1370 implements the disclosed embodiments described above. For example, the encoding module 1370 implements, processes, prepares, or provides various codec functions. Thus, the inclusion of the encoding module 1370 significantly improves the functionality of the video encoding device 1300 and changes the video encoding device 1300 to another state. Alternatively, the encoding module 1370 is implemented as instructions stored in the memory 1360 and executed by the processor 1330 .

[0258] Video encoding device 1300 may also include input and / or output (I / O) devices 1380 for communicating data to and from a user. I / O devices 1380 may include output devices such as a display for displaying video data, speakers for outputting audio data, etc. I / O devices 1380 may also include input devices such as a keyboard, mouse, trackball, etc., and / or corresponding interfaces for interacting with such output devices.

[0259] Memory 1360 may include one or more disks, tape drives, and solid state drives, which may be used as overflow data storage devices to store programs when such programs are selected for execution, and to store instructions and data read during program execution. Memory 1360 may be volatile and / or non-volatile, and may be read-only memory (ROM), random access memory (RAM), ternary content addressable memory (TCAM), and / or static random access memory (SRAM).

[0260] 14 is a schematic diagram of one embodiment of a means for encoding 1400. In one embodiment, the means for encoding 1400 is implemented in a video encoding device 1402 (e.g., video encoder 20 or video decoder 30). The video encoding device 1402 includes a means for receiving 1401. The means for receiving 1401 is configured to receive a picture to encode or a bitstream to decode. The video encoding device 1402 includes a means for transmitting 1407 coupled to the means for receiving 1401. The means for transmitting 1407 is configured to transmit the bitstream to a decoder or transmit a decoded image to a display means (e.g., one of the plurality of I / O devices 1380).

[0261] The video encoding device 1402 includes a storage means 1403. The storage means 1403 is coupled to at least one of the receiving means 1401 or the transmitting means 1407. The storage means 1403 is configured to store instructions. The video encoding device 1402 also includes a processing means 1405. The processing means 1405 is coupled to the storage means 1403. The processing means 1405 is configured to execute the instructions stored in the storage means 1403 to perform the methods disclosed herein.

[0262] It should also be understood that the steps of the exemplary methods described herein do not necessarily have to be performed in the order described, and the order of the steps of such methods should be understood to be exemplary only. Similarly, additional steps may be included in such methods, and certain steps may be omitted or combined, in methods consistent with various embodiments of the present disclosure.

[0263] While the present disclosure provides several embodiments, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. The present examples should be considered illustrative and not restrictive, and the intent is not to limit the details presented herein. For example, various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.

[0264] Furthermore, techniques, systems, subsystems, and methods described and shown in various embodiments as separate or distinct may be combined or integrated with other systems, modules, techniques, or methods without departing from the scope of the present disclosure. Other things shown or described as coupled or directly coupled to each other or in communication with each other may also be indirectly coupled or communicate through some interface, device, or intermediate component, whether electrical, mechanical, or otherwise. Other examples of changes, substitutions, and alterations will be ascertainable by those skilled in the art, and such examples may be made without departing from the spirit and scope disclosed herein. [Other possible items] (Item 1) 1. A method of decoding performed by a video decoder, comprising: receiving, by the video decoder, a first picture parameter set and a second picture parameter set, each of which references the same sequence parameter set, wherein if the first picture parameter set and the second picture parameter set have the same values ​​for picture width and picture height, the first picture parameter set and the second picture parameter set have the same value for a conformance window; applying, by the video decoder, the conformance window to a current picture corresponding to the first picture parameter set or the second picture parameter set; A method for providing (Item 2) Item 10. The method of item 1, wherein the conformance window has a conformance window left offset, a conformance window right offset, a conformance window top offset, and a conformance window bottom offset. (Item 3) 3. The method of claim 1, further comprising: decoding the current picture corresponding to the first picture parameter set or the second picture parameter set using inter prediction after the conformance window is applied, wherein the inter prediction is based on a resampled reference picture. (Item 4) 4. The method of any one of items 1 to 3, further comprising resampling a reference picture associated with the current picture corresponding to the first picture set or the second picture set using reference picture resampling (RPS). (Item 5) 5. The method of any of items 1 to 4, wherein the resampling of the reference pictures changes the resolution of the reference pictures used to inter-predict the current picture corresponding to the first picture set or the second picture set. (Item 6) 6. The method of any of items 1 to 5, wherein the picture width and the picture height are measured in luma samples. (Item 7) 7. The method according to any one of items 1 to 6, further comprising determining whether bidirectional optical flow (BDOF) is valid for decoding the picture based on the picture width, the picture height, and the conformance window of the current picture and reference pictures of the current picture. (Item 8) 7. The method according to any one of items 1 to 6, further comprising determining whether decoder-side motion vector refinement (DMVR) is enabled for decoding the picture based on the picture width, the picture height, and the conformance window of the current picture and a reference picture of the current picture. (Item 9) 7. The method according to any one of items 1 to 6, further comprising displaying an image generated using the current block on a display of an electronic device. (Item 10) 1. A method of encoding performed by a video encoder, comprising: generating, by the video encoder, a first picture parameter set and a second picture parameter set, each of which references the same sequence parameter set, wherein if the first picture parameter set and the second picture parameter set have the same values ​​for picture width and picture height, the first picture parameter set and the second picture parameter set have the same value for a conformance window; the video encoder encoding the first picture parameter set and the second picture parameter set into a video bitstream; storing the video bitstream by the video encoder for transmission to a video decoder; A method for providing (Item 11) Item 11. The method of item 10, wherein the conformance window has a conformance window left offset, a conformance window right offset, a conformance window top offset, and a conformance window bottom offset. (Item 12) 12. The method of any of items 10 to 11, wherein the picture width and the picture height are measured in luma samples. (Item 13) 13. The method of any of items 10 to 12, further comprising transmitting the video bitstream including the first picture parameter set and the second picture parameter set to the video decoder. (Item 14) a receiver configured to receive an encoded video bitstream; a memory coupled to the receiver, the memory storing instructions; and a processor coupled to the memory, the processor executing the instructions to cause the decoding device to: receiving a first picture parameter set and a second picture parameter set, each of which references the same sequence parameter set, wherein if the first picture parameter set and the second picture parameter set have the same values ​​for picture width and picture height, the first picture parameter set and the second picture parameter set have the same value for a conformance window; applying the conformance window to a current picture corresponding to the first picture parameter set or the second picture parameter set; a processor configured to cause A decoding device comprising: (Item 15) Item 15. The decoding device of item 14, wherein the conformance window has a conformance window left offset, a conformance window right offset, a conformance window top offset, and a conformance window bottom offset. (Item 16) 16. The decoding device of claim 14, further comprising: decoding the current picture corresponding to the first picture parameter set or the second picture parameter set using inter prediction after the conformance window is applied, wherein the inter prediction is based on a resampled reference picture. (Item 17) 17. The decoding device of any of items 15 to 16, wherein the decoding device further comprises a display configured to display an image generated based on the current picture. (Item 18) 1. An encoding device, comprising: a memory containing instructions; a processor coupled to the memory, the processor executing the instructions to cause the encoding device to: generating a first picture parameter set and a second picture parameter set each referencing the same sequence parameter set, wherein if the first picture parameter set and the second picture parameter set have the same values ​​for picture width and picture height, the first picture parameter set and the second picture parameter set have the same value for conformance window; encoding the first picture parameter set and the second picture parameter set into a video bitstream; a processor configured to cause a transmitter coupled to the processor, the transmitter configured to transmit the video bitstream including the first picture parameter set and the second picture parameter set to a video decoder; and An encoding device comprising: (Item 19) Item 19. The encoding device of item 18, wherein the conformance window has a conformance window left offset, a conformance window right offset, a conformance window top offset, and a conformance window bottom offset. (Item 20) 20. The encoding device of any of items 18 to 19, wherein the picture width and the picture height are measured in luma samples. (Item 21) a receiver configured to receive pictures to encode or to receive a bitstream to decode; a transmitter coupled to the receiver, the transmitter configured to transmit the bitstream to a decoder or to transmit a decoded image to a display; a memory coupled to at least one of the receiver or the transmitter, the memory configured to store instructions; and a processor coupled to the memory, the processor configured to execute the instructions stored in the memory to perform the method in any of items 1 to 9 and any of items 10 to 13; An encoding device comprising: (Item 22) Item 21. The encoding device of item 20, wherein the encoding device further comprises a display configured to display an image. (Item 23) An encoder; a decoder in communication with the encoder; 23. A system comprising: the encoder or decoder comprising a decoding device, encoding device, or coding apparatus according to any one of items 15 to 22. (Item 24) a means for encoding, receiving means configured to receive pictures to encode or to receive a bitstream to decode; transmitting means coupled to said receiving means, said transmitting means configured to transmit said bitstream to a decoding means or to transmit a decoded image to a display means; a storage means coupled to at least one of the receiving means or the transmitting means, the storage means configured to store instructions; processing means coupled to said storage means, said processing means configured to execute the instructions stored in said storage means to perform the method in any one of items 1 to 9 and any one of items 10 to 13; A means for encoding comprising:

Claims

1. 1. A method for encoding a bitstream, the method comprising: generating a first picture parameter set (PPS) and a second PPS, each of which references the same sequence parameter set (SPS), wherein the first PPS includes a first picture width parameter, a first picture height parameter, and a first conformance window flag, and the second PPS includes a second picture width parameter, a second picture height parameter, and a second conformance window flag, wherein the first conformance window flag equal to 1 indicates that the first PPS further includes a first conformance window parameter, and the second conformance window flag equal to 1 indicates that the second PPS further includes a second conformance window parameter; encoding the first PPS and the second PPS into the bitstream; Equipped with If the first picture width parameter has the same value as the second picture width parameter and the first picture height parameter has the same value as the second picture height parameter, the first conformance window parameter is constrained to have the same value as the second conformance window parameter. method.

2. The first conformance window parameter or the second conformance window parameter includes a conformance window left offset, a conformance window right offset, a conformance window upper offset, and a conformance window lower offset. The method of claim 1.

3. The first picture width parameter or the second picture width parameter is measured in luma samples, and the first picture height parameter or the second picture height parameter is measured in luma samples.

3. The method according to claim 1 or 2.

4. 1. A method for decoding a bitstream, comprising: receiving a bitstream, the bitstream including a first picture parameter set (PPS) and a second PPS, each PPS referencing the same sequence parameter set (SPS), the first PPS including a first picture width parameter, a first picture height parameter, and a first conformance window flag, the second PPS including a second picture width parameter, a second picture height parameter, and a second conformance window flag, the first conformance window flag equal to 1 indicating that the first PPS further includes a first conformance window parameter, and the second conformance window flag equal to 1 indicating that the second PPS further includes a second conformance window parameter; parsing the first picture width parameter and the first picture height parameter from the first PPS of the bitstream; parsing the second picture width parameter and the second picture height parameter from the second PPS of the bitstream; Equipped with If the first picture width parameter has the same value as the second picture width parameter and the first picture height parameter has the same value as the second picture height parameter, the first conformance window parameter is constrained to have the same value as the second conformance window parameter. method.

5. The first conformance window parameter or the second conformance window parameter includes a conformance window left offset, a conformance window right offset, a conformance window upper offset, and a conformance window lower offset. The method of claim 4.

6. The first picture width parameter or the second picture width parameter is measured in luma samples, and the first picture height parameter or the second picture height parameter is measured in luma samples.

6. The method according to claim 4 or 5.

7. a processor and a memory, The memory is configured to store instructions, and the processor is configured to execute the instructions in the memory to perform the method of any one of claims 1 to 3. Encoding device.

8. a processor and a memory, The memory is configured to store instructions, and the processor is configured to execute the instructions in the memory to perform the method of any one of claims 4 to 6. Decoding device.

9. 1. A device for storing a bitstream, the device comprising: at least one storage medium and at least one communication interface; the at least one communication interface is configured to receive or transmit the bitstream; the at least one storage medium is configured to store the bitstream; the bitstream includes a first picture parameter set (PPS) and a second PPS, each of which references the same sequence parameter set (SPS), the first PPS including a first picture width parameter, a first picture height parameter, and a first conformance window flag, the second PPS including a second picture width parameter, a second picture height parameter, and a second conformance window flag, the first conformance window flag equal to 1 indicating that the first PPS further includes a first conformance window parameter, and the second conformance window flag equal to 1 indicating that the second PPS further includes a second conformance window parameter; If the first picture width parameter has the same value as the second picture width parameter and the first picture height parameter has the same value as the second picture height parameter, the first conformance window parameter is constrained to have the same value as the second conformance window parameter. device.

10. The first conformance window parameter or the second conformance window parameter includes a conformance window left offset, a conformance window right offset, a conformance window upper offset, and a conformance window lower offset.

10. The device of claim 9.

11. The first picture width parameter or the second picture width parameter is measured in luma samples, and the first picture height parameter or the second picture height parameter is measured in luma samples.

11. A device according to claim 9 or 10.

12. When the first conformance window flag is equal to 0, it indicates that the first PPS does not include the first conformance window parameter, and when the second conformance window flag is equal to 0, it indicates that the second PPS does not include the second conformance window parameter. A device according to any one of claims 9 to 11.

13. 1. A method for storing a bitstream, comprising: receiving or transmitting a bitstream via a communication interface; storing the bitstream on one or more storage media, the bitstream including a first picture parameter set (PPS) and a second PPS, each of which references the same sequence parameter set (SPS), the first PPS including a first picture width parameter, a first picture height parameter, and a first conformance window flag, the second PPS including a second picture width parameter, a second picture height parameter, and a second conformance window flag, the first conformance window flag equal to 1 indicating that the first PPS further includes a first conformance window parameter, and the second conformance window flag equal to 1 indicating that the second PPS further includes a second conformance window parameter; Equipped with If the first picture width parameter has the same value as the second picture width parameter and the first picture height parameter has the same value as the second picture height parameter, the first conformance window parameter is constrained to have the same value as the second conformance window parameter. method.

14. The first conformance window parameter or the second conformance window parameter includes a conformance window left offset, a conformance window right offset, a conformance window upper offset, and a conformance window lower offset. The method of claim 13.

15. The first picture width parameter or the second picture width parameter is measured in luma samples, and the first picture height parameter or the second picture height parameter is measured in luma samples.

15. The method of claim 13 or 14.

16. When the first conformance window flag is equal to 0, it indicates that the first PPS does not include the first conformance window parameter, and when the second conformance window flag is equal to 0, it indicates that the second PPS does not include the second conformance window parameter.

16. The method according to any one of claims 13 to 15.

17. A device for transmitting a bitstream, said device comprising:

1. At least one storage medium configured to store at least one bitstream, the bitstream including a first picture parameter set (PPS) and a second PPS each referencing a same sequence parameter set (SPS), the first PPS including a first picture width parameter, a first picture height parameter, and a first conformance window flag, the second PPS including a second picture width parameter, a second picture height parameter, and a second conformance window flag, the first conformance window flag equal to 1 being at least one storage medium, wherein a second conformance window flag equal to 1 indicates that the first PPS further includes a first conformance window parameter, and the second conformance window flag equal to 1 indicates that the second PPS further includes a second conformance window parameter, and the first conformance window parameter is constrained to have the same value as the second conformance window parameter if the first picture width parameter has the same value as the second picture width parameter and the first picture height parameter has the same value as the second picture height parameter; at least one processor configured to obtain one or more bitstreams from one of the at least one storage medium; a transmitter configured to transmit the one or more bitstreams to a destination device; 1. A device comprising:

18. The first conformance window parameters or the second conformance window parameters include a conformance window left offset, a conformance window right offset, a conformance window upper offset, and a conformance window lower offset.

18. The device of claim 17.

19. The first picture width parameter or the second picture width parameter is measured in luma samples, and the first picture height parameter or the second picture height parameter is measured in luma samples.

19. A device according to claim 17 or 18.

20. When the first conformance window flag is equal to 0, it indicates that the first PPS does not include the first conformance window parameter, and when the second conformance window flag is equal to 0, it indicates that the second PPS does not include the second conformance window parameter.

20. A device according to any one of claims 17 to 19.

21. 1. A method for transmitting a bitstream, comprising: Storing at least one bitstream on at least one storage medium, the bitstream including a first picture parameter set (PPS) and a second PPS each referencing the same sequence parameter set (SPS), the first PPS including a first picture width parameter, a first picture height parameter, and a first conformance window flag, the second PPS including a second picture width parameter, a second picture height parameter, and a second conformance window flag, the first conformance window flag being equal to 1. a second conformance window flag equal to 1 indicates that the second PPS further includes a second conformance window parameter, and the first conformance window parameter is constrained to have the same value as the second conformance window parameter if the first picture width parameter has the same value as the second picture width parameter and the first picture height parameter has the same value as the second picture height parameter; obtaining one or more bitstreams from one of the at least one storage medium; transmitting the one or more bitstreams to a destination device; A method for providing

22. The first conformance window parameters or the second conformance window parameters include a conformance window left offset, a conformance window right offset, a conformance window upper offset, and a conformance window lower offset.

22. The method of claim 21.

23. The first picture width parameter or the second picture width parameter is measured in luma samples, and the first picture height parameter or the second picture height parameter is measured in luma samples.

23. The method of claim 21 or 22.

24. When the first conformance window flag is equal to 0, it indicates that the first PPS does not include the first conformance window parameter, and when the second conformance window flag is equal to 0, it indicates that the second PPS does not include the second conformance window parameter.

24. The method of any one of claims 21 to 23.

25. 1. A system for processing a bitstream, comprising: an encoding device, one or more storage devices, and a decoding device; the encoding device is configured to obtain a video signal and encode the video signal to obtain one or more bitstreams, at least one of the one or more bitstreams including a first picture parameter set (PPS) and a second PPS that each refer to a same sequence parameter set (SPS), the first PPS including a first picture width parameter, a first picture height parameter, and a first conformance window flag, and the second PPS including a second picture width parameter, a second picture height parameter, and a second conformance window flag; the first conformance window flag equal to 1 indicates that the first PPS further includes a first conformance window parameter, the second conformance window flag equal to 1 indicates that the second PPS further includes a second conformance window parameter, and if the first picture width parameter has the same value as the second picture width parameter and the first picture height parameter has the same value as the second picture height parameter, the first conformance window parameter is constrained to have the same value as the second conformance window parameter; the one or more storage devices are used to store the one or more bitstreams; The decoding device is used to decode the one or more bitstreams. system.

26. The first conformance window parameters or the second conformance window parameters include a conformance window left offset, a conformance window right offset, a conformance window upper offset, and a conformance window lower offset.

26. The system of claim 25.

27. The first picture width parameter or the second picture width parameter is measured in luma samples, and the first picture height parameter or the second picture height parameter is measured in luma samples.

27. A system according to claim 25 or 26.

28. a video capture device configured to capture the video signal; 28. The system of any one of claims 25 to 27, further comprising:

29. When the first conformance window flag is equal to 0, it indicates that the first PPS does not include the first conformance window parameter, and when the second conformance window flag is equal to 0, it indicates that the second PPS does not include the second conformance window parameter.

29. A system according to any one of claims 25 to 28.