Processing multiple picture sizes and conformance windows for reference picture resampling in video encoding.

By ensuring consistent conformance window sizes for picture parameter sets with the same picture size, the method optimizes video encoding and decoding processes, reducing resource usage and enhancing user experience.

JP2026053411APending Publication Date: 2026-03-25HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Existing video encoding technologies face challenges in efficiently compressing video data without sacrificing picture quality, particularly when dealing with varying picture sizes and conformance windows, leading to complex processing and resource inefficiencies in encoders and decoders.

Method used

Constraining picture parameter sets with the same picture size to have the same conformance window size, such as through cropping, to maintain consistent processing and reduce resource usage in both the encoder and decoder, thereby optimizing video encoding and decoding processes.

Benefits of technology

This approach reduces processor, memory, and network resource usage, enhancing the user experience by improving video transmission, reception, and viewing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026053411000001_ABST
    Figure 2026053411000001_ABST
Patent Text Reader

Abstract

This document provides a method, device, and program for encoding / decoding a bitstream. [Solution] The decoding method includes the step of receiving a first picture parameter set and a second picture parameter set, each referencing the same sequence parameter set. If the first and second picture parameter sets have the same values ​​with respect to picture width and picture height, then the first and second picture parameter sets have the same values ​​with respect to conformance window. The method also includes the step of applying a conformance window to the current picture corresponding to the first or second picture parameter set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross - Reference to Related Applications] This patent application claims the benefit of U.S. Provisional Patent Application No. 62 / 871,493, filed Jul. 8, 2019, by Jianle Chen et al., which is hereby incorporated herein by reference in its entirety.

[0002] The present disclosure generally describes techniques for supporting multiple picture sizes and conformance windows in video coding. More specifically, the present disclosure ensures that picture parameter sets having the same picture size also have the same conformance window.

Background Art

[0003] The amount of video data required to depict even relatively short videos can be quite large, thus causing difficulties when streaming data over a communication network with limited bandwidth capacity or communicating in another way. Therefore, video data is generally compressed before being communicated over modern telecommunications networks. The size of a video can also be a problem when the video is stored on a memory device, as memory resources may be limited. In video compression devices, the source often uses software and / or hardware to encode video data before transmission or storage, thereby reducing the amount of data required to represent a digital video image. The compressed data is then received at the destination by a video decompression device that decodes the video data. Due to limited network resources and the increasing demand for high video quality, improvements in compression / decompression techniques that improve the compression ratio without sacrificing all or most of the picture quality are desired.

Summary of the Invention

[0004] The first aspect relates to a method for decoding an encoded video bitstream performed by a video decoder. The method includes the steps of: receiving a first picture parameter set and a second picture parameter set, each referring to the same sequence parameter set, wherein the first and second picture parameter sets have the same values ​​with respect to picture width and picture height, and the first and second picture parameter sets have the same values ​​with respect to conformance window; and applying a conformance window to the current picture corresponding to the first or second picture parameter set.

[0005] This method provides a technique for constraining picture parameter sets with the same picture size to also have the same conformance window size (e.g., by cropping the window size). By maintaining the same conformance window size for picture parameter sets with the same picture size, overly complex processing can be avoided when reference picture resampling (RPR) is enabled. Consequently, the usage of processor, memory, and / or network resources can be reduced in both the encoder and decoder. In this way, the coder / decoder (also known as the codec) in video encoding is improved compared to the current codec. In practical terms, improvements in the video encoding process provide users with a more desirable user experience when transmitting, receiving, and / or viewing video.

[0006] Optionally, in any of the embodiments described above, another implementation example of this embodiment specifies that the conformance window has a conformance window left offset, a conformance window right offset, a conformance window upper offset, and a conformance window lower offset.

[0007] Optionally, in any of the embodiments described above, another implementation of this embodiment provides a step of decoding the current picture corresponding to a first or second picture parameter set using interpretation after a conformance window has been applied, where interpretation is based on a resampled reference picture.

[0008] Optionally, in any of the embodiments described above, another implementation example of this embodiment provides a step of resampling a reference picture associated with the current picture corresponding to a first picture parameter set or a second picture parameter set using reference picture resampling (RPS).

[0009] Optionally, in any of the embodiments described above, another implementation of this embodiment specifies that the resampling of the reference picture changes the resolution of the reference picture used to interpret the current picture corresponding to the first or second picture parameter set.

[0010] Optionally, in any of the embodiments described above, another implementation of this embodiment specifies that the picture width and picture height are measured using lumens.

[0011] Optionally, in any of the embodiments described above, another implementation example of this embodiment provides a step of determining whether bidirectional optical flow (BDOF) is effective for decoding the picture based on the picture width, picture height, conformance window, and reference picture of the current picture.

[0012] Optionally, in any of the embodiments described above, another implementation example of this embodiment provides a step of determining whether decoder-side motion vector fine-tuning (DMVR) is effective for decoding a picture based on the picture width, picture height, conformance window, and reference picture of the current picture.

[0013] Optionally, in any of the embodiments described above, another implementation example of this embodiment provides a step of displaying an image generated using the current block on the display of an electronic device.

[0014] A second aspect relates to a method for encoding a video bitstream performed by a video encoder. The method includes the steps of: generating a first picture parameter set and a second picture parameter set, each referencing the same sequence parameter set, wherein the first and second picture parameter sets have the same values ​​with respect to picture width and picture height, and the first and second picture parameter sets have the same values ​​with respect to conformance window; encoding the first and second picture parameter sets into a video bitstream; and storing the video bitstream for transmission to a video decoder.

[0015] This method provides a technique for constraining picture parameter sets with the same picture size to also have the same conformance window size (e.g., by cropping the window size). By maintaining the same conformance window size for picture parameter sets with the same picture size, overly complex processing can be avoided when reference picture resampling (RPR) is enabled. Consequently, the usage of processor, memory, and / or network resources can be reduced in both the encoder and decoder. In this way, the coder / decoder (also known as the codec) in video encoding is improved compared to the current codec. In practical terms, improvements in the video encoding process provide users with a more desirable user experience when transmitting, receiving, and / or viewing video.

[0016] Optionally, in any of the embodiments described above, another implementation example of this embodiment specifies that the conformance window has a conformance window left offset, a conformance window right offset, a conformance window upper offset, and a conformance window lower offset.

[0017] Optionally, in any of the embodiments described above, another implementation of this embodiment specifies that the picture width and picture height are measured using lumens.

[0018] Optionally, in any of the embodiments described above, another implementation of this embodiment provides a step of transmitting a video bitstream including a first picture parameter set and a second picture parameter set to a video decoder.

[0019] A third aspect relates to a decoding device. The decoding device includes a receiver configured to receive an encoded video bitstream, a memory coupled to the receiver which stores instructions, and a processor coupled to the memory which executes instructions and causes the decoding device to receive a first picture parameter set and a second picture parameter set, each referring to the same sequence parameter set, wherein if the first and second picture parameter sets have the same values ​​with respect to picture width and picture height, the first and second picture parameter sets have the same values ​​with respect to conformance window, and the processor applies a conformance window to the current picture corresponding to the first or second picture parameter set.

[0020] The decoding device provides a technique to constrain picture parameter sets with the same picture size to also have the same conformance window size (e.g., by cropping the window size). By maintaining the same conformance window size for picture parameter sets with the same picture size, overly complex processing can be avoided when reference picture resampling (RPR) is enabled. Consequently, the usage of processor, memory, and / or network resources can be reduced in both the encoder and decoder. In this way, the coder / decoder (also known as the codec) in video encoding is improved compared to the current codec. In practical terms, improvements in the video encoding process provide users with a more desirable user experience when transmitting, receiving, and / or viewing video.

[0021] Optionally, in any of the embodiments described above, another implementation example of this embodiment specifies that the conformance window has a conformance window left offset, a conformance window right offset, a conformance window upper offset, and a conformance window lower offset.

[0022] Optionally, in any of the embodiments described above, another implementation example of this embodiment provides decoding the current picture corresponding to a first or second picture parameter set using interpretation after a conformance window has been applied, where interpretation is based on a resampled reference picture.

[0023] Optionally, in any of the embodiments described above, another implementation example of this embodiment is provided, which is configured to display an image generated based on the current picture.

[0024] A fourth aspect relates to an encoding device. The encoding device includes a memory containing instructions; a processor coupled to the memory, the processor is configured to execute instructions to cause the encoding device to generate a first picture parameter set and a second picture parameter set, each referencing the same sequence parameter set, wherein if the first and second picture parameter sets have the same values ​​with respect to picture width and picture height, the first and second picture parameter sets have the same values ​​with respect to conformance window; and encode the first and second picture parameter sets into a video bitstream; and a transmitter coupled to the processor, the transmitter is configured to transmit a video bitstream containing the first and second picture parameter sets to a video decoder.

[0025] The encoding device provides a technique for constraining (e.g., cropping the window size) a picture parameter set having the same picture size to also have the same compliance window size. By maintaining the compliance window at the same size for picture parameter sets having the same picture size, when reference picture resampling (RPR) is effective, overly complex processing can be avoided. Thus, the usage amounts of the processor, memory, and / or network resources can be reduced in both the encoder and the decoder. In this way, the coder / decoder (also known as codec) in video encoding is improved compared to the current codec. As a practical matter, the improvement of the video encoding process provides a more desirable user experience for the user when sending, receiving, and / or viewing video.

[0026] Optionally, in any of the foregoing aspects, another implementation example of this aspect defines that the compliance window has a compliance window left offset, a compliance window right offset, a compliance window top offset, and a compliance window bottom offset.

[0027] Optionally, in any of the foregoing aspects, another implementation example of this aspect defines that the picture width and picture height are measured in lumasamples.

[0028] Aspect 5 relates to an encoding device. The encoding device includes a receiver configured to receive a picture to be encoded or a bitstream to be decoded, a transmitter coupled to the receiver, the transmitter being configured to transmit the bitstream to a decoder or to transmit a decoded picture to a display, a memory coupled to at least one of the receiver or the transmitter, the memory being configured to store instructions, and a processor coupled to the memory, the processor being configured to execute the instructions stored in the memory to perform any of the methods disclosed herein.

[0029] The encoding device provides a technique for constraining (e.g., cropping the window size) a picture parameter set having the same picture size to also have the same compliance window size. By maintaining the compliance window at the same size for picture parameter sets having the same picture size, when reference picture resampling (RPR) is effective, overly complex processing can be avoided. Thus, the usage of processor, memory, and / or network resources can be reduced in both the encoder and the decoder. In this way, the coder / decoder (also known as codec) in video encoding is improved compared to the current codec. As a practical matter, the improvement of the video encoding process provides a more desirable user experience for users when sending, receiving, and / or viewing video.

[0030] Optionally, in any of the foregoing aspects, another implementation of this aspect provides a display configured to display an image.

[0031] Aspect 6 relates to a system. The system includes an encoder and a decoder that communicates with the encoder, where the encoder or the decoder includes a decoding device, an encoding device, or an encoding device disclosed herein.

[0032] The system provides a technique for constraining picture parameter sets with the same picture size to also have the same conformance window size (e.g., cropping the window size). By maintaining the same conformance window size for picture parameter sets with the same picture size, overly complex processing can be avoided when reference picture resampling (RPR) is enabled. Consequently, the usage of processor, memory, and / or network resources can be reduced in both the encoder and decoder. In this way, the coder / decoder (also known as the codec) in video encoding is improved compared to the current codec. In practical terms, improvements in the video encoding process provide users with a more desirable user experience when transmitting, receiving, and / or viewing video.

[0033] The seventh aspect relates to an encoding means. The encoding means includes a receiving means configured to receive a picture to encode or a bitstream to decode; a transmitting means coupled to the receiving means, the transmitting means configured to transmit a bitstream to a decoding means or a decoded image to a display means; a storage means coupled to at least one of the receiving means or the transmitting means, the storage means configured to store instructions; and a processing means coupled to the storage means, the processing means configured to execute instructions stored in the storage means and perform one of the methods disclosed herein.

[0034] The encoding method provides a technique to constrain picture parameter sets with the same picture size to also have the same conformance window size (e.g., by cropping the window size). By maintaining the same conformance window size for picture parameter sets with the same picture size, overly complex processing can be avoided when reference picture resampling (RPR) is enabled. Consequently, the usage of processor, memory, and / or network resources can be reduced in both the encoder and decoder. In this way, the coder / decoder (also known as the codec) in video encoding is improved compared to the current codec. In practical terms, improvements in the video encoding process provide users with a more desirable user experience when transmitting, receiving, and / or viewing video.

[0035] For the purpose of clarification, one of the embodiments described above may be combined with one or more of the other embodiments described above to create new embodiments that fall within the scope of this disclosure.

[0036] These and other features will be more clearly understood from the detailed description below, which will be used in conjunction with the attached drawings and claims. [Brief explanation of the drawing]

[0037] For a better understanding of this disclosure, the following brief description, which is used together with the accompanying drawings and detailed description, is referred to here. Here, like reference numbers represent like parts.

[0038] [Figure 1] This is a flowchart illustrating one example of a method for encoding video signals.

[0039] [Figure 2] This is a schematic diagram illustrating an example of a video coding / decoding (codec) system.

[0040] [Figure 3] This is a schematic diagram showing an example of a video encoder.

[0041] [Figure 4] This is a schematic diagram showing an example of a video decoder.

[0042] [Figure 5] This is an encoded video sequence that shows the relationship between the leading picture, the trailing picture, and the intra-random access point (IRAP) picture in terms of decoding order and display order.

[0043] [Figure 6] An example of multilayer coding for spatial scalability is presented.

[0044] [Figure 7] This is a schematic diagram showing an example of unidirectional interpretation.

[0045] [Figure 8] This is a schematic diagram showing an example of bidirectional interface prediction.

[0046] [Figure 9] This shows the video bitstream.

[0047] [Figure 10] This shows a method for splitting a picture.

[0048] [Figure 11] This is one embodiment of a method for decoding an encoded video bitstream.

[0049] [Figure 12] This is one embodiment of a method for encoding an encoded video bitstream.

[0050] [Figure 13] This is a schematic diagram of a video encoding device.

[0051] [Figure 14] This is a schematic diagram of one embodiment relating to a means of encoding. [Modes for carrying out the invention]

[0052] While exemplary implementations of one or more embodiments are provided below, it should be understood from the outset that any number of techniques, whether currently known or existing, may be used to implement the disclosed systems and / or methods. This disclosure should not be limited to the exemplary implementations, drawings, and techniques shown below (including exemplary designs and implementations shown and described herein), but may be modified within the scope of the appended claims combined with their equivalents.

[0053] The terms below are defined as follows, unless used in the opposite context herein. Specifically, the following definitions are intended to provide further clarity to this disclosure. However, these terms may be described differently in different contexts. Therefore, the following definitions should be considered supplementary and should not be considered to limit any other definitions of such terms provided herein.

[0054] A bitstream is a sequence of bits containing compressed video data for transmission between an encoder and a decoder. An encoder is a device configured to use an encoding process to compress video data into a bitstream. A decoder is a device configured to use a decoding process to reconstruct the video data from the bitstream for display. A picture is an array of lumens and / or chroma samples that make up a frame or its fields. A picture being encoded or decoded may be referred to as the current picture for clarity.

[0055] A reference picture is a picture containing reference samples that can be used to encode other pictures by reference according to interpretation and / or interhierarchical prediction. A reference picture list is a list of reference pictures used for interpretation and / or interhierarchical prediction. Some video coding systems use two reference picture lists, which may be represented as reference picture list 1 and reference picture list 0. A reference picture list structure is an addressable syntax structure that contains multiple reference picture lists. Interpretation is a method of encoding samples in the current picture by referencing indicated samples in a reference picture that is different from the current picture, and the reference picture and the current picture are at the same hierarchical level. A reference picture list structure entry is an addressable location in the reference picture list structure that indicates a reference picture associated with a reference picture list.

[0056] A slice header is a part of an encoded slice that contains data elements related to all the video data within a tile represented by a slice. A picture parameter set (PPS) is a set of parameters that contains data related to the entire picture. More specifically, a PPS is a syntax structure that contains syntax elements that apply to zero or more entire encoded pictures, determined by the syntax elements found in each picture header. A sequence parameter set (SPS) is a set of parameters that contains data related to a sequence of pictures. An access unit (AU) is a set of one or more encoded pictures associated with the same display time (e.g., the same picture sequence count) for output from the decode picture buffer (DPB) (e.g., for display to the user). A decoded video sequence is a sequence of pictures reconstructed by the decoder in preparation for display to the user.

[0057] A conformance cropping window (or simply a conformance window) refers to a window of picture samples contained in the encoded video sequence output from the encoding process. The bitstream may provide conformance window cropping parameters to indicate the output region of the encoded picture. Picture width is the width of the picture as measured in lumen samples. Picture height is the height of the picture as measured in lumen samples. Conformance window offsets (e.g., conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, conf_win_bottom_offset) define the picture samples that refer to the PPS output from the decoding process by a rectangular region defined by the output picture coordinates.

[0058] Decoder-side motion vector fine-tuning (DMVR) is a process, algorithm, or coding tool used to fine-tune the predicted block motion or motion vector. DMVR allows for the determination of motion vectors based on two motion vectors observed in bilateral predictions using a bilateral template matching process. DMVR can determine weighted combinations of predictive coding units generated using each of the two motion vectors, and the two motion vectors can be fine-tuned by replacing them with a new motion vector that best points to the combined predictive coding unit. Bidirectional optical flow (BDOF) is a process, algorithm, or coding tool used to fine-tune the predicted block motion or motion vector. BDOF allows for the determination of sub-coding unit motion vectors based on the gradient of the difference between two reference pictures.

[0059] Reference picture resampling (RPR) is the ability to change the spatial resolution of an encoded picture midway through the bitstream without requiring intra-encoding of the picture at the resolution change point. As used herein, resolution refers to the number of pixels contained in a video file. That is, resolution is the width and height of the video and is measured in pixels. For example, a video may have a resolution of 1280 (horizontal pixels) × 720 (vertical pixels). This is usually simply written as 1280 × 720 or abbreviated as 720p.

[0060] Decoder-side motion vector fine-tuning (DMVR) is a process, algorithm, or coding tool used to fine-tune the predicted block motion or motion vector. Bidirectional optical flow (BDOF), also known as bidirectional optical flow (BIO), is a process, algorithm, or coding tool used to fine-tune the predicted block motion or motion vector. Reference picture resampling (RPR) is the ability to change the spatial resolution of an encoded picture midway through the bitstream without requiring intra-encoding of the picture at the resolution change point.

[0061] In this specification, the following acronyms are used: Encoded Tree Block (CTB), Encoded Tree Unit (CTU), Encoded Unit (CU), Encoded Video Sequence (CVS), Joint Video Experts Team (JVET), Motion-Constrained Tile Set (MCTS), Maximum Transmission Unit (MTU), Network Abstraction Layer (NAL), Picture Order Count (POC), Low-Byte Sequence Payload (RBSP), Sequence Parameter Set (SPS), Multipurpose Video Coding (VVC), and Working Draft (WD).

[0062] Figure 1 is a flowchart of operation method 100 as an example of video signal encoding. Specifically, the video signal is encoded by an encoder. The encoding process compresses the video signal by reducing the video file size using various methods. A smaller file size makes it possible to send the compressed video file to the user while reducing the associated bandwidth overhead. The decoder then decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process is generally very similar to the encoding process, enabling the decoder to reconstruct the video signal without inconsistencies.

[0063] In step 101, a video signal is input to the encoder. For example, the video signal may be an uncompressed video file stored in memory. Alternatively, the video file may be captured by a video recording device such as a video camera and encoded to support live streaming of the video. The video file may contain both audio and video components. The video component contains a series of image frames that give the impression of visual motion when viewed in sequence. These frames contain pixels represented by brightness, which is referred to herein as the lumen component (or lumen sample), and color, which is referred to herein as the chromen component (or color sample). In some examples, the frames may also include depth values ​​to support three-dimensional display.

[0064] In step 103, the video is divided into multiple blocks. The division step includes subdividing the pixels within each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC) (also known as H.265 and MPEG-H Part 2), a frame can first be divided into coding tree units (CTUs), which are blocks of a predetermined size (e.g., 64 pixels × 64 pixels). A CTU contains both lumen and chroma samples. The coding tree may be used to divide the CTU into multiple blocks, and then to recursively subdivide these blocks until a configuration supporting further encoding is obtained. For example, the lumen component of a frame may be subdivided until the individual blocks contain relatively homogeneous illumination values. Furthermore, the chroma component of a frame may be subdivided until the individual blocks contain relatively homogeneous color values. Thus, the division method varies depending on the contents of the video frame.

[0065] In step 105, the image blocks divided in step 103 are compressed using various compression methods. For example, interpretation and / or intrapretation may be used. Interpretation is designed to take advantage of the fact that objects in a common scene tend to appear in consecutive frames. Therefore, a block depicting an object in a reference frame does not need to be repeatedly depicted in adjacent frames. Specifically, an object such as a table may remain in the same position across multiple frames. Therefore, once the table is depicted, adjacent frames can refer back to the reference frame. Pattern matching methods may be used to match objects across multiple frames. Furthermore, moving objects may be depicted across multiple frames, for example, by the movement of the object or the movement of the camera. As a concrete example, a video may show a car traversing the screen across multiple frames. Motion vectors can be used to depict such movement. A motion vector is a two-dimensional vector that provides an offset from the coordinates of an object in a frame to the coordinates of the object in a reference frame. Therefore, in interpretation, an image block in the current frame can be encoded as a series of motion vectors indicating the offset from the corresponding block in the reference frame.

[0066] Intra-prediction encodes blocks within a common frame. It leverages the fact that lumen and chroma components tend to cluster in a given frame. For example, green areas in a tree tend to be adjacent to similar green areas. Intra-prediction uses multiple directional prediction modes (e.g., 33 in HEVC), planar mode, and DC mode. Directional modes indicate that the current block is similar to / identical to samples of adjacent blocks in the corresponding direction. Planar mode indicates that a series of blocks along a row / column (e.g., a plane) can be interpolated based on adjacent blocks at the ends of the row. Planar mode actually shows smooth brightness / color transitions along the row / column by using a relatively constant gradient for value changes. DC mode is used for boundary smoothing, indicating that a given block is similar to / identical to the mean value associated with samples of all adjacent blocks associated with the angular direction of the directional prediction mode. Thus, intra-predicted blocks can represent image blocks as values ​​of various related prediction modes instead of actual values. Furthermore, intra-predicted blocks can represent image blocks as values ​​of motion vectors instead of actual values. In both cases, the prediction block does not necessarily have to accurately represent the image block. All differences are stored in the residual block. Multiple transformations may be applied to this residual block to further compress the file.

[0067] In step 107, various filtering techniques may be applied. In HEVC, multiple filters are applied according to an in-loop filtering scheme. The block-based prediction described above may produce images with block noise in the decoder. Furthermore, the block-based prediction scheme may encode blocks and then reconstruct the encoded blocks for later use as reference blocks. In an in-loop filtering scheme, noise suppression filters, deblocking filters, adaptive loop filters, and sample-adaptive offset (SAO) filters are repeatedly applied to blocks / frames. These filters reduce such blocking artifacts so that the encoded file can be accurately reconstructed. Furthermore, these filters reduce artifacts in the reconstructed reference blocks so that artifacts are less likely to create further artifacts in subsequent blocks encoded based on the reconstructed reference blocks.

[0068] Once the video signal has been split, compressed, and filtered, the resulting data is encoded into a bitstream in step 109. The bitstream includes, in addition to the data described above, any signaling data desirable to support proper video signal reconstruction in the decoder. For example, such data may include split data, prediction data, residual blocks, and various flags that provide encoding instructions to the decoder. The bitstream may be stored in memory for transmission to the decoder upon request. The bitstream may be broadcast to multiple decoders and / or multicast. The creation of the bitstream is an iterative process; therefore, steps 101, 103, 105, 107, and 109 may be performed sequentially and / or simultaneously for a large number of frames and blocks. The order shown in Figure 1 is presented for clarity and ease of explanation, but is not intended to restrict the video encoding process to a specific order.

[0069] The decoder receives the bitstream and begins the decoding process in step 111. Specifically, the decoder uses an entropy decoding scheme to convert the bitstream into corresponding syntax and image data. Using the syntax data from the bitstream, the decoder determines the frame divisions in step 111. These divisions should coincide with the result of the block division in step 103. Herein, we describe the entropy encoding / decoding used in step 111. During the compression process, the encoder makes numerous choices, such as selecting a block division scheme from several possible options based on the spatial arrangement of values ​​contained in the input image. Numerous bins may be used to signal the correct selection. As used herein, bins are binary values ​​treated as variables (e.g., bit values ​​that can change depending on the context). Entropy coding allows the encoder to discard any options that are obviously not feasible for a particular case, leaving a set of acceptable options. Each of the acceptable options is then assigned a codeword. The length of a codeword is based on the number of acceptable options (e.g., two options in one bin, three to four options in two bins, etc.). The encoder then encodes the codeword of the selected option. This reduces the size of the codeword because it is desirable that the codeword uniquely represents a selection from a small subset of acceptable options, rather than uniquely representing a selection from a potentially large set of all possible options. The decoder then decodes the selection by determining the set of acceptable options, in a similar manner to the encoder. By determining the set of acceptable options, the decoder can read the codeword and determine the selection made by the encoder.

[0070] In step 113, the decoder performs block decoding. Specifically, the decoder uses the inverse transform to generate residual blocks. The decoder then uses the residual blocks and the corresponding prediction blocks to reconstruct the image blocks according to the partitioning. The prediction blocks may include both intra-prediction blocks and inter-prediction blocks generated by the encoder in step 105. The reconstructed image blocks are then placed in the frames of the reconstructed video signal according to the partitioning data determined in step 111. The syntax of step 113 may be signaled in the bitstream via the entropy coding described above.

[0071] In step 115, the reconstructed video signal frames are filtered in a manner similar to that of step 107 in the encoder. For example, noise suppression filters, deblocking filters, adaptive loop filters, and SAO filters may be applied to the frames to remove blocking artifacts. After filtering the frames, the video signal may be output to a display in step 117 for viewing by the end user.

[0072] Figure 2 is a schematic diagram of an example of a video coding / decoding (codec) system 200. Specifically, the codec system 200 provides functions to support an implementation example of operation method 100. The codec system 200 is generalized to show components used for both the encoder and the decoder. The codec system 200 receives and divides the video signal as described in steps 101 and 103 of operation method 100. This yields the divided video signal 201. The codec system 200 then compresses the divided video signal 201 into an encoded bitstream when acting as an encoder, as described in steps 105, 107, and 109 of method 100. When acting as a decoder, the codec system 200 generates an output video signal from the bitstream, as described in steps 111, 113, 115, and 117 of operation method 100. The codec system 200 includes a general encoder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decode picture buffer component 223, and a header format and context-adaptive binary arithmetic coding (CABAC) component 231. Such components are combined as shown in the figure. In Figure 2, the black lines show the motion of the data being encoded / decoded, and the dashed lines show the motion of the control data that controls the operation of the other components. All components of the codec system 200 may be present in the encoder. The decoder may contain a subset of the components of the codec system 200. For example, the decoder may include an intra-picture prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded picture buffer component 223. These components are described here.

[0073] The segmented video signal 201 is a captured video sequence, which is divided into blocks of pixels by an encoding tree. The encoding tree uses various decoupling modes to subdivide blocks of pixels into smaller blocks of pixels. These blocks can then be further subdivided into even smaller blocks. These blocks are sometimes called nodes on the encoding tree. A large parent node is separated into smaller child nodes. The number of times a node is subdivided is called the node / encoding tree depth. Segmented blocks may also be contained within an encoding unit (CU). For example, a CU may be a sub-part of a CTU that includes a lumen block, a red difference (Cr) block, and a blue difference (Cb) block, along with corresponding syntax instructions for the CU. Decoupling modes may include binary trees (BT), triple trees (TT), and quad trees (QT), which are used to divide a node into two, three, or four child nodes, respectively, and their shapes vary depending on the decoupling mode used. The segmented video signal 201 is transferred for compression to the integrated encoder control component 211, the transformation scaling and quantization component 213, the intrapicture estimation component 215, the filter control analysis component 227, and the motion estimation component 221.

[0074] The integrated encoder control component 211 is configured to make decisions related to encoding the video sequence images into a bitstream, in accordance with application constraints. For example, the integrated encoder control component 211 manages the optimization of bitrate / bitstream size versus reconstruction quality. Such decisions may be based on memory space / bandwidth availability and image resolution requirements. The integrated encoder control component 211 also manages buffer utilization, taking transmission speed into consideration, to reduce buffer underrun and buffer overrun problems. To manage these problems, the integrated encoder control component 211 manages partitioning, prediction, and filtering by other components. For example, the integrated encoder control component 211 may dynamically increase compression complexity to increase resolution and bandwidth utilization, or decrease compression complexity to decrease resolution and bandwidth utilization. Thus, the integrated encoder control component 211 controls other components of the codec system 200 to balance the quality of video signal reconstruction with bitrate issues. The integrated encoder control component 211 generates control data, thereby controlling the operation of other components. The control data is also transferred to the header format and CABAC component 231, where it is encoded into a bitstream that signals the parameters for decoding by the decoder.

[0075] The segmented video signal 201 is also sent to the motion estimation component 221 and the motion compensation component 219 for interprediction. A frame or slice of the segmented video signal 201 may be divided into multiple video blocks. The motion estimation component 221 and the motion compensation component 219 perform interpredictive coding of the received video blocks by comparing them with one or more blocks in one or more reference frames to provide time predictions. The codec system 200 may perform multiple coding passes, for example, to select an appropriate coding mode for each block of video data.

[0076] The motion estimation component 221 and the motion compensation component 219 may be highly integrated, but are shown separately for conceptual purposes. Motion estimation performed by the motion estimation component 221 is the process of generating motion vectors, which estimate the motion of the image blocks. For example, the motion vectors may show the displacement of an encoded object by comparing it to a predicted block. A predicted block is a block that, from the point of pixel difference, is known to be a very good match to the block being encoded. Predicted blocks are sometimes also called reference blocks. Such pixel differences may be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference metrics. HEVC uses several encoded objects, including CTUs, encoded tree blocks (CTBs), and CUs. For example, a CTU may be divided into multiple CTBs, which may then be divided into multiple CBs to be included in a CU. A CU may be encoded as a prediction unit (PU) containing prediction data and / or a transform unit (TU) containing transformed residual data for the CU. The motion estimation component 221 generates motion vectors, PUs, and TUs by using rate distortion analysis as part of the rate distortion optimization process. For example, the motion estimation component 221 may determine multiple reference blocks and multiple motion vectors for the current block / frame, and may also select reference blocks and motion vectors with optimal rate distortion characteristics. Optimal rate distortion characteristics balance the quality of the video reconstruction (e.g., the amount of data loss due to compression) and the encoding efficiency (e.g., the size of the final encoding).

[0077] In some examples, the codec system 200 may calculate the values ​​of sub-integer pixel positions of the reference picture stored in the decode picture buffer component 223. For example, the video codec system 200 may interpolate the values ​​of quarter-pixel, eighth-pixel, or other fractional pixel positions of the reference picture. Thus, the motion estimation component 221 may perform motion search on the complete pixel positions and fractional pixel positions and output a motion vector with fractional pixel precision. The motion estimation component 221 calculates the motion vector of the video block in the inter-encoded slice relative to the PU by comparing the PU position with the predicted block position of the reference picture. The motion estimation component 221 outputs the calculated motion vector as motion data to the header format and CABAC component 231 for encoding, and also outputs the motion to the motion compensation component 219.

[0078] Motion compensation is performed by the motion compensation component 219, which may involve fetching or generating predicted blocks based on motion vectors determined by the motion estimation component 221. Here again, in some examples, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated. Upon receiving the motion vector of the PU of the current image block, the motion compensation component 219 may identify the position of the predicted block pointed to by the motion vector. Next, a residual image block is formed and pixel difference values ​​are formed by subtracting the pixel values ​​of the current image block being encoded from the pixel values ​​of the predicted block. Generally, the motion estimation component 221 performs motion estimation on the lumen component, and the motion compensation component 219 uses motion vectors calculated based on the lumen component for both the chromen and lumen components. The predicted and residual blocks are transferred to the transformation scaling and quantization component 213.

[0079] The segmented video signal 201 is also sent to the intra-picture estimation component 215 and the intra-picture prediction component 217. Similar to the motion estimation component 221 and the motion compensation component 219, the intra-picture estimation component 215 and the intra-picture prediction component 217 may be highly integrated, but are shown separately for conceptual purposes. As described above, the intra-picture estimation component 215 and the intra-picture prediction component 217 perform intra-prediction of the current block for each block within the current frame, as an alternative to the inter-frame prediction performed by the motion estimation component 221 and the motion compensation component 219. Specifically, the intra-picture estimation component 215 determines the intra-prediction mode to use for encoding the current block. In some examples, the intra-picture estimation component 215 selects an appropriate intra-prediction mode for encoding the current block from several intra-prediction modes being tested. The selected intra-prediction mode is then forwarded to the header format and CABAC component 231 for encoding.

[0080] For example, the intra-picture estimation component 215 calculates rate distortion values ​​for various intra-prediction modes under test using rate distortion analysis and selects the intra-prediction mode with the best rate distortion characteristics among the tested modes. Rate distortion analysis generally determines the amount of distortion (or error) between an encoded block and the original unencoded block encoded to produce the encoded block, and the bitrate (e.g., the number of bits) used to produce the encoded block. The intra-picture estimation component 215 calculates ratios from the distortion and rates of various encoded blocks to determine which intra-prediction mode exhibits the best rate distortion value for each block. Furthermore, the intra-picture estimation component 215 may be configured to encode depth blocks of a depth map using a depth modeling mode (DMM) based on rate distortion optimization (RDO).

[0081] The intrapicture prediction component 217, when implemented in an encoder, may generate residual blocks from prediction blocks based on a selected intrapicture prediction mode determined by the intrapicture estimation component 215, or, when implemented in a decoder, may read residual blocks from a bitstream. The residual blocks contain the difference in values ​​between the prediction blocks and the original blocks and are represented as a matrix. The residual blocks are then transferred to the transformation scaling and quantization component 213. The intrapicture estimation component 215 and the intrapicture prediction component 217 may process both lumen and chroma components.

[0082] The transformation scaling and quantization component 213 is configured to further compress the residual block. The transformation scaling and quantization component 213 applies a transformation, such as a discrete cosine transform (DCT), discrete sine transform (DST), or a conceptually similar transformation, to the residual block to create a video block containing residual transformation coefficient values. Wavelet transforms, integer transforms, subband transforms, or other types of transformations may also be used. This transformation may convert the residual information from the pixel value domain to a transformation domain, such as the frequency domain. The transformation scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling involves applying a scale factor to the residual information, which can result in different frequency information being quantized at different granularities, potentially affecting the final display quality of the reconstructed video. The transformation scaling and quantization component 213 is also configured to quantize the transformation coefficients to further reduce the bitrate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting the quantization parameters. In some examples, the transformation scaling and quantization component 213 may then scan a matrix containing the quantized transformation coefficients. The quantized transformation coefficients are then transferred to the header format and CABAC component 231 and encoded into a bitstream.

[0083] The scaling and inverse transform component 229 supports motion estimation by applying the inverse operation of the transform scaling and quantization component 213. The scaling and inverse transform component 229 applies inverse scaling, transform, and / or quantization to reconstruct the residual block in the pixel region for later use as a reference block that may become a prediction block for another current block, for example. The motion estimation component 221 and / or motion compensation component 219 may compute the reference block by adding the residual block to the corresponding prediction block for use in motion estimation of a later block / frame. A filter is applied to the reconstructed reference block to reduce artifacts generated during scaling, quantization, and transform. Otherwise, such artifacts could lead to false predictions (and generate further artifacts) when predicting subsequent blocks.

[0084] The filter-controlled analysis component 227 and the in-loop filter component 225 apply filters to residual blocks and / or reconstructed image blocks. For example, the transformed residual blocks from the scaling and inverse transform component 229 may be combined with the corresponding predicted blocks from the intra-picture prediction component 217 and / or the motion compensation component 219 to reconstruct the original image block. Filters may then be applied to the reconstructed image block. In some examples, instead, filters may be applied to the residual blocks. As with the other components in Figure 2, the filter-controlled analysis component 227 and the in-loop filter component 225 are highly integrated and may be implemented together, but are shown separately for conceptual purposes. Filters applied to the reconstructed reference block are applied to a specific spatial region and include several parameters for adjusting how such filters are applied. The filter-controlled analysis component 227 analyzes the reconstructed reference block to determine where such filters should be applied and sets the corresponding parameters. Such data is transferred to the header format and CABAC component 231 as filter-controlled data for encoding. The in-loop filter component 225 applies such filters based on filter control data. These filters may include deblocking filters, noise suppression filters, SAO filters, and adaptive loop filters. Such filters may be applied in the spatial / pixel domain (e.g., on reconstructed pixel blocks) or in the frequency domain, depending on the case.

[0085] When operating as an encoder, the filtered reconstructed image blocks, residual blocks, and / or prediction blocks are stored in the decode picture buffer component 223 for later use in the motion estimation described above. When operating as a decoder, the decode picture buffer component 223 stores the reconstructed and filtered blocks and transfers them to the display as part of the output video signal. The decode picture buffer component 223 can be any memory device capable of storing the prediction blocks, residual blocks, and / or reconstructed image blocks.

[0086] The header format and CABAC component 231 receive data from various components of the codec system 200, encode such data, and convert it into an encoded bitstream for transmission to the decoder. Specifically, the header format and CABAC component 231 generate various headers and encode control data such as overall control data and filter control data. Furthermore, prediction data, including intra-prediction and motion data, as well as residual data in the form of quantized transformation coefficient data, are all encoded into a bitstream. The final bitstream contains all the information desirable for the decoder to reconstruct the original segmented video signal 201. Such information may include an index table of intra-prediction modes (also called a codeword mapping table), definitions of encoding contexts for various blocks, indications of the most likely intra-picture modes, indications of segmentation information, etc. Such data may be encoded using entropy coding. For example, such information may be encoded using context-adaptive variable-length coding (CAVLC), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC), stochastic interval-partitioned entropy (PIPE) coding, or another entropy coding technique. After entropy coding, the encoded bitstream may be transmitted to another device (e.g., a video decoder) or archived for later transmission or retrieval.

[0087] Figure 3 is a block diagram showing an example of a video encoder 300. The video encoder 300 may be used to implement the encoding function of the codec system 200 and / or to carry out steps 101, 103, 105, 107, and / or 109 of the operation method 100. The encoder 300 splits the input video signal to obtain a split video signal 301, which is substantially the same as the split video signal 201. The split video signal 301 is then compressed by the components of the encoder 300 and encoded into a bitstream.

[0088] Specifically, the segmented video signal 301 is transferred to the intra-picture prediction component 317 for intra-prediction. The intra-picture prediction component 317 may be substantially the same as the intra-picture estimation component 215 and the intra-picture prediction component 217. The segmented video signal 301 is also transferred to the motion compensation component 321 for inter-prediction based on the reference block contained in the decoded picture buffer component 323. The motion compensation component 321 may be substantially the same as the motion estimation component 221 and the motion compensation component 219. The prediction block and residual block from the intra-picture prediction component 317 and the motion compensation component 321 are transferred to the transformation and quantization component 313 for transformation and quantization of the residual block. The transformation and quantization component 313 may be substantially the same as the transformation scaling and quantization component 213. The transformed and quantized residual block, and the corresponding prediction block (along with the associated control data), are transferred to the entropy coding component 331 for encoding into a bitstream. The entropy coding component 331 may be substantially the same as the header format and the CABAC component 231.

[0089] The transformed and quantized residual blocks, and / or the corresponding prediction blocks, are also transferred from the transform and quantization component 313 to the inverse transform and quantization component 329 for reconstruction into reference blocks for use by the motion compensation component 321. The inverse transform and quantization component 329 may be substantially the same as the scaling and inverse transform component 229. The in-loop filters included in the in-loop filter component 325 are also applied to the residual blocks and / or the reconstructed reference blocks, as applicable. The in-loop filter component 325 may be substantially the same as the filter control analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 may include multiple filters as described with respect to the in-loop filter component 225. The filtered blocks are then stored in the decode picture buffer component 323 for use as reference blocks by the motion compensation component 321. The decode picture buffer component 323 may be substantially the same as the decode picture buffer component 223.

[0090] Figure 4 is a block diagram showing an example of a video decoder 400. The video decoder 400 may be used to implement the decoding function of the codec system 200 and / or to carry out steps 111, 113, 115, and / or 117 of the operation method 100. The decoder 400 receives a bitstream from, for example, the encoder 300 and generates a reconstructed output video signal based on this bitstream for display to the end user.

[0091] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 is configured to implement an entropy decoding scheme such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component 433 may use header information to provide context and interpret additional data encoded as codewords in the bitstream. The decoded information includes any desirable information for decoding the video signal, such as overall control data, filter control data, segmentation information, motion data, prediction data, and quantized transformation coefficients from residual blocks. The quantized transformation coefficients are transferred to the inverse transform and quantization component 429 for reconstruction into residual blocks. The inverse transform and quantization component 429 may be similar to the inverse transform and quantization component 329.

[0092] The reconstructed residual blocks and / or predicted blocks are transferred to the intra-picture prediction component 417 for reconstruction into image blocks based on the intra-prediction operation. The intra-picture prediction component 417 may be similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. Specifically, the intra-picture prediction component 417 uses a prediction mode to locate the reference block in the frame and applies the residual block to the result to reconstruct the intra-predicted image block. The reconstructed and intra-predicted image block and / or residual block, as well as the corresponding intra-prediction data, are transferred to the decode picture buffer component 423 via the in-loop filter component 425. These components may be substantially similar to the in-loop filter component 225 and the decode picture buffer component 223, respectively. The in-loop filter component 425 filters the reconstructed image block, residual block, and / or predicted block, and such information is stored in the decode picture buffer component 423. The reconstructed image blocks from the decode picture buffer component 423 are transferred to the motion compensation component 421 for interpretation. The motion compensation component 421 may be substantially the same as the motion estimation component 221 and / or motion compensation component 219. Specifically, the motion compensation component 421 generates a prediction block using a motion vector from a reference block and reconstructs the image block by applying a residual block to the result. The resulting reconstructed block may also be transferred to the decode picture buffer component 423 via the in-loop filter component 425. The decode picture buffer component 423 continues to store further reconstructed image blocks, which may be reconstructed into frames by segmentation information. Such frames may be arranged in a sequence, which is output to a display as a reconstructed output video signal.

[0093] Based on the above, these video compression techniques reduce or eliminate redundancy inherent in a video sequence by performing spatial (intra-picture) prediction and / or temporal (inter-picture) prediction. In block-based video coding, a video slice (i.e., a video picture or part of a video picture) may be divided into multiple video blocks, which may also be called tree blocks, coded tree blocks (CTBs), coded tree units (CTUs), coded units (CUs), and / or coded nodes. A video block in an intra-coded (I) slice of a picture is encoded using spatial prediction with respect to a reference sample contained in an adjacent block within the same picture. A video block in an inter-coded (P or B) slice of a picture may use spatial prediction with respect to a reference sample contained in an adjacent block within the same picture, or temporal prediction with respect to a reference sample contained in another reference picture. These pictures may be called frames, and reference pictures may be called reference frames.

[0094] Spatial or temporal prediction yields a predicted block for the block to be encoded. Residual data represents the pixel difference between the original block to be encoded and the predicted block. Inter-encoded blocks are encoded according to motion vectors pointing to the reference sample blocks forming the predicted block, and residual data showing the difference between the encoded block and the predicted block. Intra-encoded blocks are encoded according to the intra-encoded mode and residual data. For further compression, the residual data can be transformed from the pixel domain to the transformation domain to obtain residual transformation coefficients. These residual transformation coefficients may then be quantized. The quantized transformation coefficients are initially arranged in a two-dimensional array and may be scanned to produce a one-dimensional vector of transformation coefficients, and entropy coding may be applied to achieve further compression.

[0095] Image and video compression has grown rapidly, resulting in a variety of encoding standards. Such video encoding standards include ITU-T H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) MPEG-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, Next Generation Video Coding (AVC) (also known as ITU-T H.264 or ISO / IEC MPEG-4 Part 10), and High Efficiency Video Coding (HEVC) (also known as ITU-T H.265 or MPEG-H Part 2). AVC includes extensions such as Scalable Video Coding (SVC), Multi-View Video Coding (MVC) and Multi-View Video Coding + Depth (MVC+D), as well as 3D AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multi-view HEVC (MV-HEVC), and 3D HEVC (3D-HEVC).

[0096] There is also a new video coding standard called Versatile Video Coding (VVC), which is under development by the Joint Team of Video Experts (JVET) of the ITU-T and ISO / IEC. There are several working drafts of the VVC standard, but specifically, one working draft (WD) on VVC, namely "Versatile Video Coding (Draft 5)" JVET-N1001-v3 (13th JVET Meeting, March 27, 2019) (VVC Draft 5) by B. Bross, J. Chen, and S. Liu, is referred to herein. Each of the references contained in this paragraph and the preceding paragraphs is incorporated in whole by reference.

[0097] The descriptions of each technique disclosed herein are based on Multipurpose Video Coding (VVC), a video coding standard under development by the Joint Team of Video Experts (JVET) of the ITU-T and ISO / IEC. However, these techniques are also applicable to other video codec specifications.

[0098] Figure 5 is a representation 500 relating the relationship between the intra-random access point (IRAP) picture 502 and the leading picture 504 and trailing picture 506 in the decoding order 508 and display order 510. In one embodiment, the IRAP picture 502 is called a clean random access (CRA) picture or an immediate decoder refresh (IDR) picture with a random access decodeable (RADL) picture. In HEVC, the IDR picture, CRA picture, and broken access (BLA) picture are all considered IRAP pictures 502. For VVC, at the 12th JVET meeting in October 2018, it was agreed that both the IDR picture and the CRA picture should be considered IRAP pictures. In one embodiment, the broken access (BLA) picture and the gradual decoder refresh (GDR) picture may also be considered IRAP pictures. The decoding process of an encoded video sequence always begins with IRAP.

[0099] As shown in Figure 5, the leading pictures 504 (e.g., pictures 2 and 3) come after the IRAP picture 502 in the decoding order 508, but before the IRAP picture 502 in the display order 510. The trailing picture 506 comes after the IRAP picture 502 in both the decoding order 508 and the display order 510. Although two leading pictures 504 and one trailing picture 506 are shown in Figure 5, those skilled in the art will understand that in actual application, more or fewer leading pictures 504 and / or trailing pictures 506 may be present in the decoding order 508 and the display order 510.

[0100] The reading picture 504 in Figure 5 is divided into two types: Random Access Skip Reading (RASL) and RADL. If decoding begins with IRAP picture 502 (e.g., picture 1), a RADL picture (e.g., picture 3) can be properly decoded. However, a RASL picture (e.g., picture 2) cannot be properly decoded. Therefore, the RASL picture is discarded. Considering the difference between RADL and RASL pictures, the type of reading picture 504 associated with IRAP picture 502 must be specified as either RADL or RASL for efficient and proper encoding. HEVC is constrained that, if both RASL and RADL pictures exist, for RASL and RADL pictures associated with the same IRAP picture 502, the RASL picture must come before the RADL picture in the display order 510.

[0101] IRAP picture 502 provides the following two important functions / benefits. First, the presence of IRAP picture 502 indicates that the decoding process can be started from that picture. This function enables a random access function, where as long as IRAP picture 502 is present at that location, the decoding process starts at that location in the bitstream, not necessarily at the beginning of the bitstream. Second, the presence of IRAP picture 502 refreshes the decoding process, so that encoded pictures (excluding RASL pictures) starting with IRAP picture 502 are encoded without referencing the previous picture at all. Therefore, the presence of IRAP picture 502 in the bitstream prevents any errors that may occur when decoding the encoded picture before IRAP picture 502 from propagating to IRAP picture 502 and the pictures that come after IRAP picture 502 in the decoding order 508.

[0102] IRAP picture 502 provides important functionality, but this comes at a cost to compression efficiency. The presence of IRAP picture 502 causes a sharp increase in bitrate. This cost to compression efficiency is due to two reasons. Firstly, because IRAP picture 502 is an intra-predictive picture, it itself requires relatively more bits to represent compared to other pictures that are inter-predictive pictures (e.g., leading picture 504, trailing picture 506). Secondly, because the presence of IRAP picture 502 interrupts time prediction (because the decoder refreshes the decoding process, and for this reason one of the actions of the decoding process is to remove the previous reference picture in the decode picture buffer (DPB)), IRAP picture 502 reduces the encoding efficiency of pictures that come after IRAP picture 502 in the decoding order 508 (i.e., they require more bits to represent). This is because these pictures do not have a reference picture for inter-predictive coding.

[0103] Among the picture types considered IRAP Picture 502, HEVC IDR pictures differ in signaling and derivation compared to other picture types. Some of these differences are as follows:

[0104] For the signaling and derivation of the Picture Order Count (POC) value of an IDR picture, the most significant bit (MSB) of the POC is set to simply equal to 0, rather than being derived from the previous significant picture.

[0105] Regarding the signaling information required for reference picture management, the slice header of an IDR picture does not contain any information that needs to be signaled to assist in reference picture management. For other picture types (i.e., CRA, trailing, time sublayer access (TSA), etc.), information such as the Reference Picture Set (RPS) or other forms of similar information (e.g., a reference picture list) described below is required for the reference picture marking process (i.e., the process of determining the status of reference pictures contained in the Decode Picture Buffer (DPB), i.e., whether they are used for reference or not). However, for IDR pictures, there is no need to signal such information. This is because the presence of the IDR indicates that all reference pictures contained in the DPB must simply be marked by the decoding process as not being used for reference.

[0106] In addition to the concept of IRAP pictures, there are also reading pictures associated with IRAP pictures, if any exist. A reading picture is a picture that comes after the associated IRAP picture in the decoding order, but before the IRAP picture in the output order. Depending on the encoding configuration and picture reference structure, reading pictures can be further categorized into two types. The first type is a reading picture that may not be correctly decoded if the decoding process starts with its associated IRAP picture. This can happen because such reading pictures are encoded by referencing a picture that comes before the IRAP picture in the decoding order. Such reading pictures are called Random Access Skip Readings (RASLs). The second type is a reading picture that must be correctly decoded even if the decoding process starts with its associated IRAP picture. This is possible because such reading pictures are encoded without directly or indirectly referencing a picture that comes before the IRAP picture in the decoding order. Such reading pictures are called Random Access Decodeable Readings (RADLs). In HEVC, if both RASL and RADL pictures exist, there is a constraint that, for RASL and RADL pictures associated with the same IRAP picture, the RASL picture must come before the RADL picture in the output order.

[0107] In HEVC and VVC, the IRAP picture 502 and the leading picture 504 may each be contained within a single Network Abstraction Layer (NAL) unit. A set of NAL units is sometimes called an access unit. The IRAP picture 502 and the leading picture 504 are given different NAL unit types that are easily identifiable by system-level applications. For example, a video splicer needs to understand the encoded picture type without needing to over-understand the details of the syntax elements contained in the encoded bitstream, and in particular, identify the IRAP picture 502 from a non-IRAP picture and the leading picture 504 from a trailing picture 506 (including determining RASL and RADL pictures). The trailing picture 506 is a picture associated with the IRAP picture 502 and comes after the IRAP picture 502 in display order 510. A certain picture may come after a specific IRAP picture 502 in the decoding order 508, and may come before any other IRAP picture 502 in the decoding order 508. For this purpose, assigning each IRAP picture 502 and reading picture 504 its own NAL unit type is helpful in such applications.

[0108] For HEVC, the NAL unit types of IRAP pictures include the following: • BLA with Reading Picture (BLA_W_LP): A NAL unit of a broken link access (BLA) picture where one or more reading pictures may come later in the decoding order. • BLA with RADL (BLA_W_RADL): A NAL unit of a BLA picture in which one or more RADL pictures come later in the decoding order, but RASL pictures are not required. • BLA without a leading picture (BLA_N_LP): A NAL unit of a BLA picture where the leading picture does not come after the leading picture in the decoding order. • IDR with RADL (IDR_W_RADL): A NAL unit of an IDR picture in which one or more RADL pictures come later in the decoding order, but RASL pictures are not required. • IDR without a reading picture (IDR_N_LP): A NAL unit of an IDR picture where the reading picture does not come after the reading picture in the decoding order. • CRA: A NAL unit of a clean random access (CRA) picture, which may be followed by a leading picture (i.e., a RASL picture or a RADL picture or both). • RADL: NAL unit for RADL pictures. • RASL: NAL unit for RASL pictures.

[0109] For VVC, the NAL unit types for IRAP picture 502 and reading picture 504 are as follows: • IDR with RADL (IDR_W_RADL): A NAL unit of an IDR picture in which one or more RADL pictures come later in the decoding order, but RASL pictures are not required. • IDR without a reading picture (IDR_N_LP): A NAL unit of an IDR picture where the reading picture does not come after the reading picture in the decoding order. • CRA: A NAL unit of a clean random access (CRA) picture, which may be followed by a leading picture (i.e., a RASL picture or a RADL picture or both). • RADL: NAL unit for RADL pictures. • RASL: NAL unit for RASL pictures.

[0110] Reference Picture Resampling (RPR) is the ability to change the spatial resolution of an encoded picture midway through the bitstream without requiring intra-encoding of the picture at the resolution change point. For this to work, the picture must be able to reference one or more reference pictures whose spatial resolution differs from the current picture for inter-prediction. Therefore, resampling, or part thereof, of such reference pictures is necessary for encoding and decoding the current picture. Hence the name RPR. This feature is sometimes also called Adaptive Resolution Change (ARC), among other names. Examples of use cases or application scenarios that would benefit from the RPR feature include the following:

[0111] Rate adaptation in video conferencing and video telephony. This involves adapting encoded video to changing network conditions. When network conditions deteriorate and available bandwidth decreases, the encoder may adapt to this by encoding a lower-resolution picture.

[0112] Switching active speakers in multi-site video conferences. In multi-site video conferences, the active speaker's video size is typically larger or wider than that of other conference participants. When the active speaker switches, it may be necessary to adjust the picture resolution for each participant. If active speaker switching occurs frequently, the need for ARC (Automatic Channel Resolution) functionality becomes even more important.

[0113] Faster streaming initiation. Streaming applications typically buffer a picture until it has been decoded to a certain length before they begin displaying it. Starting the bitstream at a lower resolution allows the application to hold enough picture in the buffer to begin displaying it more quickly.

[0114] Adaptive stream switching in streaming. The Dynamic Adaptive Streaming (DASH) specification over HTTP includes a feature called @mediaStreamStructureId. This feature allows switching between different displays at random access points of an open GOP (group of picture) using an undecodeable reading picture (e.g., a CRA picture with an associated RASL picture in HEVC). If two different displays of the same video have different bitrates but the same spatial resolution and the same value for @mediaStreamStructureId, then switching between the two displays in the CRA picture with the associated RASL picture may occur, and the RASL picture associated with the CRA picture switching may be decoded with acceptable quality, thus enabling seamless switching. With ARC, the @mediaStreamStructureId feature could also be used to switch between DASH displays of various spatial resolutions.

[0115] Various methods facilitate the implementation of fundamental techniques to support RPR / ARC, such as signaling for the list of picture resolutions and imposing several constraints on resampling reference pictures within the DPB.

[0116] One component of the techniques required to support RPR is a way to signal the possible picture resolutions present in the bitstream. This is addressed in some examples by modifying the current signaling of picture resolutions in the SPS, as shown below. [Table 1]

[0117] num_pic_size_in_luma_samples_minus1+1 specifies the number of picture sizes (width and height) in units of luma samples that may exist in the encoded video sequence.

[0118] pic_width_in_luma_samples[i] specifies the width of the i-th decoded picture in units of luma samples that may exist in the encoded video sequence. pic_width_in_luma_samples[i] must not be equal to 0 and must be an integer multiple of MinCbSizeY.

[0119] pic_height_in_luma_samples[i] specifies the height of the i-th decoded picture in units of luma samples that may exist in the encoded video sequence. pic_height_in_luma_samples[i] must not be equal to 0 and must be an integer multiple of MinCbSizeY.

[0120] At the 15th JVET meeting, another modification of signaling picture size and conformance window to support RPR was discussed. The signaling is as follows:

[0121] • Signals the maximum picture size (i.e., picture width and picture height) for SPS.

[0122] • Signals the picture size in the Picture Parameter Set (PPS).

[0123] • Move the current conformance window signaling from SPS to PPS. Conformance window information is used to crop the reconstructed / decoded picture during the process of preparing the picture for output. The cropped picture size is the size of the picture after it has been cropped using its associated conformance window.

[0124] The signaling for picture size and conformance window is as follows: [Table 2]

[0125] The `max_width_in_luma_samples` setting specifies that the `pic_width_in_luma_samples` of any picture in which this SPS is active must be less than or equal to `max_width_in_luma_samples` as a requirement for bitstream conformance.

[0126] The `max_height_in_luma_samples` parameter specifies that the `pic_height_in_luma_samples` of any picture in which this SPS is active must be less than or equal to `max_height_in_luma_samples` as a requirement for bitstream conformance. [Table 3]

[0127] `pic_width_in_luma_samples` specifies the width of each decoded picture referencing the PPS in units of luma samples. `pic_width_in_luma_samples` must not be equal to 0 and must be an integer multiple of `MinCbSizeY`.

[0128] `pic_height_in_luma_samples` specifies the height of each decoded picture referencing the PPS in units of luma samples. `pic_height_in_luma_samples` must not be equal to 0 and must be an integer multiple of `MinCbSizeY`.

[0129] For any active referenced picture whose width and height are reference_pic_width_in_luma_samples and reference_pic_height_in_luma_samples, all of the following conditions must be met for bitstream conformance to be valid. · 2 × pic_width_in_luma_samples ≥ reference_pic_width_in_luma_samples · 2 × pic_height_in_luma_samples ≥ reference_pic_height_in_luma_samples · pic_width_in_luma_samples ≤ 8 × reference_pic_width_in_luma_samples · pic_height_in_luma_samples ≤ 8 × reference_pic_height_in_luma_samples

[0130] The variables PicWidthInCtbsY, PicHeightInCtbsY, PicSizeInCtbsY, PicWidthInMinCbsY, PicHeightInMinCbsY, PicSizeInMinCbsY, PicSizeInSamplesY, PicWidthInSamplesC, and PicHeightInSamplesC are derived as follows. · PicWidthInCtbsY = Ceil(pic_width_in_luma_samples / CtbSizeY) (1) · PicHeightInCtbsY = Ceil(pic_height_in_luma_samples / CtbSizeY) (2) · PicSizeInCtbsY = PicWidthInCtbsY × PicHeightInCtbsY (3) · PicWidthInMinCbsY = pic_width_in_luma_samples / MinCbSizeY (4) · PicHeightInMinCbsY = pic_height_in_luma_samples / MinCbSizeY (5) · PicSizeInMinCbsY = PicWidthInMinCbsY × PicHeightInMinCbsY (6) ·PicSizeInSamplesY=pic_width_in_luma_samples×pic_height_in_luma_samples (7) ·PicWidthInSamplesC=pic_width_in_luma_samples / SubWidthC (8) ·PicHeightInSamplesC=pic_height_in_luma_samples / SubHeightC (9)

[0131] A conformance_window_flag equal to 1 indicates that the conformance cropping window offset parameter is followed in PPS. A conformance_window_flag equal to 0 indicates that the conformance cropping window offset parameter does not exist.

[0132] conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset define the sample of the picture output from the decoding process, which references the PPS, by a rectangular area defined by picture coordinates for output. If conformance_window_flag is equal to 0, the values ​​of conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset are presumed to be equal to 0.

[0133] The conformance cropping window contains luma samples that have horizontal picture coordinates (including boundaries) from [SubWidthC×conf_win_left_offset] to [pic_width_in_luma_samples-(SubWidthC×conf_win_right_offset+1)] and vertical picture coordinates (including boundaries) from [SubHeightC×conf_win_top_offset] to [pic_height_in_luma_samples-(SubHeightC×conf_win_bottom_offset+1)].

[0134] The value of [SubWidthC × (conf_win_left_offset + conf_win_right_offset)] must be less than pic_width_in_luma_samples. Also, the value of [SubHeightC × (conf_win_top_offset + conf_win_bottom_offset)] must be less than pic_height_in_luma_samples.

[0135] The variables PicOutputWidthL and PicOutputHeightL are derived as follows: ·PicOutputWidthL=pic_width_in_luma_samples-SubWidthC×(conf_win_right_offset+conf_win_left_offset) (10) ·PicOutputHeightL=pic_height_in_pic_size_units-SubHeightC×(conf_win_bottom_offset+conf_win_top_offset) (11)

[0136] If ChromaArrayType is not equal to 0, the corresponding specified sample of the two chroma arrays is the sample with picture coordinates (x / SubWidthC, y / SubHeightC), where (x, y) are the picture coordinates of the specified chroma sample.

[0137] Note: The offset parameter of the conformance cropping window is only applied to the output. All internal decoding processes are applied to the uncropped picture size.

[0138] The signaling of picture size and conformance window within PPS causes the following problems:

[0139] Because multiple PPSs can exist in an encoded video sequence (CVS), two PPSs can have the same picture size signaling but contain different conformance window signalings. This can result in a situation where two pictures referencing different PPSs have the same picture size but different cropping sizes.

[0140] Regarding RPR support, it has been suggested that if the current picture and reference picture of a block have different picture sizes, some encoding tools should be turned off to encode that block. However, nowadays, even if two pictures have the same picture size, their cropping sizes may differ, so additional checks based on cropping size are necessary.

[0141] This specification discloses a technique for restricting picture parameter sets having the same picture size to also have the same conformance window size (e.g., cropping the window size). By maintaining the same conformance window size for picture parameter sets having the same picture size, overly complex processing can be avoided when reference picture resampling (RPR) is enabled. Consequently, the usage of processor, memory, and / or network resources can be reduced in both the encoder and decoder. Thus, the coder / decoder (also known as the codec) in video encoding is improved compared to the current codec. In practical terms, improvements in the video encoding process provide users with a more desirable user experience when transmitting, receiving, and / or viewing video.

[0142] Scalability in video coding is typically supported by multilayer coding techniques. A multilayer bitstream has a base layer (BL) and one or more extension layers (EL). Examples of scalability include spatial scalability, signal-to-noise (SNR) scalability, and multi-view scalability. When using multilayer coding techniques, a picture or a part of a picture may be coded by (1) not using a reference picture (i.e., using intra-prediction), (2) referencing a reference picture at the same hierarchical level (i.e., using inter-prediction), or (3) referencing a reference picture at another hierarchical level (i.e., using inter-hierarchical prediction). A reference picture used for inter-hierarchical prediction of the current picture is called an inter-hierarchical reference picture (ILRP).

[0143] Figure 6 is a schematic diagram illustrating an example of a hierarchical prediction 600 performed to determine the MV in, for example, the block compression stage 105, the block decoding stage 113, the motion estimation component 221, the motion compensation component 219, the motion compensation component 321, and / or the motion compensation component 421. The hierarchical prediction 600 is compatible with unidirectional interpretation and / or bidirectional interpretation, but can also be performed between pictures of different hierarchies.

[0144] The hierarchical prediction 600 is applied between pictures 611, 612, 613, and 614, which are in different hierarchies, and between pictures 615, 616, 617, and 618. In the example shown, pictures 611, 612, 613, and 614 are part of hierarchical[N+1]632, and pictures 615, 616, 617, and 618 are part of hierarchical[N]631. Hierarchies such as hierarchical[N]631 and / or hierarchical[N+1]632 are groups of pictures associated with similar values ​​of characteristics such as similar size, quality, resolution, signal-to-noise ratio, and capability. In the example shown, hierarchical[N+1]632 is associated with larger image sizes than hierarchical[N]631. Therefore, in this example, pictures 611, 612, 613, and 614 in hierarchy [N+1]632 have larger picture sizes than pictures 615, 616, 617, and 618 in hierarchy [N]631 (e.g., larger height and width, and therefore more samples). However, such pictures can be separated between hierarchy [N+1]632 and hierarchy [N]631 based on other characteristics. Although only two hierarchies, hierarchy [N+1]632 and hierarchy [N]631, are shown, a set of pictures can be divided into any number of hierarchies based on relevant characteristics. Hierarchies [N+1]632 and hierarchy [N]631 may be represented by hierarchy IDs. A hierarchy ID is a data item associated with a picture that indicates that the picture belongs to the hierarchy shown. Therefore, each picture 611-618 is associated with a corresponding hierarchy ID, indicating whether the corresponding picture belongs to hierarchy [N+1]632 or hierarchy [N]631.

[0145] Pictures 611-618 in different tiers 631-632 are configured to be displayed in different ways. Thus, pictures 611-618 in different tiers 631-632 can share the same temporary identifier (ID) and be included in the same AU. As used herein, an AU is a set of one or more encoded pictures associated with the same display time for output from the DPB. For example, if a smaller picture is desired, the decoder may decode picture 615 and display it at the current display time. Alternatively, if a larger picture is desired, the decoder may decode picture 611 and display it at the current display time. Thus, pictures 611-614 in the higher tier [N+1]632 contain substantially the same image data as the corresponding pictures 615-618 in the lower tier [N]631 (despite the difference in picture size). Specifically, picture 611 contains essentially the same image data as picture 615, and picture 612 contains essentially the same image data as picture 616.

[0146] Pictures 611-618 may be encoded by referencing other pictures 611-618 in the same hierarchy [N]631 or [N+1]632. The encoding of a picture that references another picture in the same hierarchy becomes an interpretation 623, which is compatible with unidirectional interpretations and / or bidirectional interpretations. Interpretations 623 are indicated by solid arrows. For example, picture 613 may be encoded using an interpretation 623 that references one or two of pictures 611, 612, and / or 614 in hierarchy [N+1]632, where one picture is referenced for unidirectional interpretations and / or two pictures are referenced for bidirectional interpretations. Furthermore, picture 617 may be encoded using an interpretation 623 that references one or two of pictures 615, 616, and / or 618 in hierarchy [N]631. Here, one picture is referenced for unidirectional interprediction, and / or two pictures are referenced for bidirectional interprediction. When performing interprediction 623, a picture is sometimes called a reference picture when it is used as a reference to another picture at the same hierarchical level. For example, picture 612 may be a reference picture used to encode picture 613 according to interprediction 623. Interprediction 623 may also be called intrahierarchical prediction in a multilayer context. Thus, interprediction 623 is a method of encoding a sample of the current picture by referencing an indicated sample in a reference picture that is different from the current picture, and the reference picture and the current picture are at the same hierarchical level.

[0147] Pictures 611-618 may be encoded by referencing other pictures 611-618 located in different hierarchies. This process is known as interhierarchical prediction 621 and is indicated by a dashed arrow. Interhierarchical prediction 621 is a method of encoding a sample of the current picture by referencing an indicated sample in a reference picture, where the current picture and the reference picture are in different hierarchies and therefore have different hierarchical IDs. For example, a picture in a lower hierarchy [N]631 may be used as the reference picture to encode the corresponding picture in a higher hierarchy [N+1]632. Specifically, picture 611 may be encoded by referencing picture 615 according to interhierarchical prediction 621. In such a case, picture 615 is used as the interhierarchical reference picture. An interhierarchical reference picture is a reference picture used in interhierarchical prediction 621. In most cases, the interhierarchical prediction 621 has constraints such that the current picture, such as picture 611, can only use interhierarchical reference pictures, such as picture 615, that are in the same AU and are in a lower hierarchy. If multiple hierarchies (e.g., more than two) are available, the interhierarchical prediction 621 can encode / decode the current picture based on multiple interhierarchical reference pictures that are at a lower level than the current picture.

[0148] The video encoder can encode pictures 611-618 using a hierarchical prediction 600 and numerous different combinations and / or changes in the order of inter-prediction 623 and inter-hierarchical prediction 621. For example, picture 615 may be encoded according to intra-prediction. Pictures 616-618 may then be encoded according to inter-prediction 623 by using picture 615 as a reference picture. Furthermore, picture 611 may be encoded according to inter-hierarchical prediction 621 by using picture 615 as an inter-hierarchical reference picture. Pictures 612-614 may then be encoded according to inter-prediction 623 by using picture 611 as a reference picture. Thus, a reference picture can serve as both a single-hierarchical reference picture and an inter-hierarchical reference picture for different encoding schemes. By encoding the pictures of the upper layer [N+1]632 based on the pictures of the lower layer [N]631, the use of intra prediction can be avoided in the upper layer [N+1]632. Intra prediction is far less efficient in encoding than inter-prediction 623 and inter-layer prediction 621. Therefore, intra prediction, with its low encoding efficiency, may be limited to the smallest / lowest quality pictures and thus limited to encoding the smallest amount of video data. Pictures used as reference pictures and / or inter-layer reference pictures may be indicated in the reference picture list entries contained in the reference picture list structure.

[0149] Previous H.26x video coding systems provided support for scalability in profiles different from those used for single-layer coding. Scalable Video Coding (SVC) is a scalable extension of AVC / H.264 that provides support for spatial, temporal, and quality scalability. In SVC, each macroblock (MB) in an EL picture is signaled with a flag to indicate whether the MB in the EL can be predicted using a block at the same location from a lower layer. Predictions from blocks at the same location may include textures, motion vectors, and / or coding modes. An example implementation of SVC cannot directly reuse an unmodified example H.264 / AVC implementation. The EL macroblock syntax and decoding process of SVC differ from the H.264 / AVC syntax and decoding process.

[0150] Scalable HEVC (SHVC) is an extension of the HEVC / H.265 standard that provides support for spatial and quality scalability, multi-view HEVC (MV-HEVC) is an extension of HEVC / H.265 that provides support for multi-view scalability, and 3D HEVC (3D-HEVC) is an extension of HEVC / H.264 that provides support for more advanced and efficient three-dimensional (3D) video coding than MV-HEVC. Note that temporal scalability is included as an integer part of the single-layer HEVC codec. The design of the multi-layer extension of HEVC uses the idea that the decode picture used for inter-layer prediction comes only from the same access unit (AU) and is treated as a long-term reference picture (LTRP), and is assigned a reference index in the reference picture list along with other temporal reference pictures in the current layer. Inter-layer prediction (ILP) is implemented at the prediction unit (PU) level by setting the value of the reference index and referencing the inter-layer reference picture in the reference picture list.

[0151] In particular, both the reference picture resampling function and the spatial scalability function require resampling of the reference picture or a portion thereof. Reference picture resampling may be implemented at the picture level or the coded block level. However, when RPR is referred to as a coding function, it is a function of single-layer coding. Even so, from a codec design perspective, it is possible, or rather preferable, to use the same resampling filter for both the RPR function of single-layer coding and the spatial scalability function of multi-layer coding.

[0152] Figure 7 is a schematic diagram showing an example of unidirectional interpretation 700. Unidirectional interpretation 700 can be used to determine the motion vectors of the encoded and / or decoded blocks created when a picture is split.

[0153] The unidirectional interpretation 700 predicts the current block 711 contained in the current frame 710 using a reference frame 730 having a reference block 731. The reference frame 730 may be placed temporally after the current frame 710 (e.g., as a successor reference frame), as shown, but in some examples it may be placed temporally before the current frame 710 (e.g., as a preceding reference frame). The current frame 710 is an example of a frame / picture being encoded / decoded at a particular time. The current frame 710 contains in the current block 711 an object that matches an object in the reference block 731 of the reference frame 730. The reference frame 730 is the frame used as a reference for encoding the current frame 710, and the reference block 731 is the block in the reference frame 730 that contains an object that is also contained in the current block 711 of the current frame 710.

[0154] The current block 711 is an arbitrary encoding unit being encoded / decoded at a specific stage of the encoding process. The current block 711 may be an entire segmented block, or a subblock if an affine interpretation mode is used. The current frame 710 is separated from the reference frame 730 by a certain time distance (TD) 733. TD 733 represents the time between the current frame 710 and the reference frame 730 in the video sequence and may be measured in units of frames. The prediction information of the current block 711 may refer to the reference frame 730 and / or reference block 731 by a reference index that indicates the direction and time distance between frames. During the period represented by TD 733, an object in the current block 711 moves from one position in the current frame 710 to another position in the reference frame 730 (e.g., the position of the reference block 731). For example, the object may move along a motion trajectory 713, which is the direction of the object's movement over time. The motion vector 735 represents the direction and magnitude of the object's movement along the motion trajectory 713 in TD 733. Therefore, the encoded motion vector 735, the reference block 731, and the residual including the difference between the current block 711 and the reference block 731 provide enough information to reconstruct the current block 711 and place the current block 711 in the current frame 710.

[0155] Figure 8 is a schematic diagram showing an example of a bidirectional interpretation 800. The bidirectional interpretation 800 can be used to determine the motion vectors of the encoded and / or decoded blocks created when a picture is split.

[0156] The bidirectional interpretation 800 is similar to the unidirectional interpretation 700, but uses a pair of reference frames to predict the current block 811 contained within the current frame 810. Thus, the current frame 810 and the current block 811 are substantially the same as the current frame 710 and the current block 711, respectively. The current frame 810 is temporally positioned between the preceding reference frame 820, which exists before the current frame 810 in the video sequence, and the succeeding reference frame 830, which exists after the current frame 810 in the video sequence. Apart from this, the preceding reference frame 820 and the succeeding reference frame 830 are substantially the same as the reference frame 730.

[0157] The current block 811 matches the preceding reference block 821 contained in the preceding reference frame 820 and the succeeding reference block 831 contained in the succeeding reference frame 830. Such matching indicates that, in the course of the video sequence, the object moves from the position of the preceding reference block 821 to the position of the succeeding reference block 831 along the motion trajectory 813, passing through the current block 811. The current frame 810 is a certain preceding time distance (TD0) 823 away from the preceding reference frame 820 and a certain succeeding time distance (TD1) 833 away from the succeeding reference frame 830. TD0 (823) indicates the time between the preceding reference frame 820 and the current frame 810 in the video sequence, in units of frames. TD1 (833) indicates the time between the current frame 810 and the succeeding reference frame 830 in the video sequence, in units of frames. Therefore, the object moves from the preceding reference block 821 to the current block 811 along the motion trajectory 813 over the period indicated by TD0 (823). The object moves from the current block 811 to the subsequent reference block 831 along the motion trajectory 813 over the period indicated by TD1(833). The prediction information for the current block 811 may refer to the preceding reference frame 820 and / or preceding reference block 821, and the subsequent reference frame 830 and / or subsequent reference block 831, by a pair of reference indices indicating the direction and time distance between frames.

[0158] The preceding motion vector (MV0) 825 represents the direction and magnitude of the object's motion over TD0(823) along the motion trajectory 813 (for example, between the preceding reference frame 820 and the current frame 810). The succeeding motion vector (MV1) 835 represents the direction and magnitude of the object's motion over TD1(833) along the motion trajectory 813 (for example, between the current frame 810 and the succeeding reference frame 830). Thus, the bidirectional interpretation 800 can encode and reconstruct the current block 811 using the preceding reference block 821 and / or succeeding reference block 831, MV0(825), and MV1(835).

[0159] In one embodiment, interpretation and / or bidirectional interpretation may be performed sample by sample (e.g., pixel by pixel) instead of block by block. That is, a motion vector pointing to each sample contained in the preceding reference block 821 and / or succeeding reference block 831 may be determined for each sample contained in the current block 811. In such an embodiment, the motion vectors 825 and 835 shown in Figure 8 represent multiple motion vectors corresponding to multiple samples contained in the current block 811, the preceding reference block 821, and the succeeding reference block 831.

[0160] In both merge mode and advanced motion vector prediction (AMVP) mode, the candidate list is generated by adding candidate motion vectors to the candidate list in an order defined by the candidate list determination pattern. Such candidate motion vectors may include motion vectors from unidirectional interpretation 700, bidirectional interpretation 800, or a combination thereof. Specifically, these motion vectors are generated for adjacent blocks when such blocks are encoded. Such motion vectors are added to the candidate list of the current block, and the motion vectors for the current block are selected from this candidate list. The motion vectors may then be signaled as indices of the selected motion vectors in the candidate list. The decoder can construct the candidate list using the same process as the encoder and determine the selected motion vectors from the candidate list based on the signaled indices. Thus, the candidate motion vectors include motion vectors generated according to unidirectional interpretation 700 and / or bidirectional interpretation 800, depending on which technique is used when encoding such adjacent blocks.

[0161] Figure 9 shows the video bitstream 900. As used herein, the video bitstream 900 may also be referred to as the encoded video bitstream, the bitstream, or a variation thereof. As shown in Figure 9, the bitstream 900 includes a sequence parameter set (SPS) 902, a picture parameter set (PPS) 904, a slice header 906, and image data 908.

[0162] SPS902 contains data common to all pictures included in a sequence of pictures (SOP). In contrast, PPS904 contains data common to all pictures. The slice header 906 contains information about the current slice, such as the slice type and which of the reference pictures is used. SPS902 and PPS904 are sometimes commonly referred to as parameter sets. SPS902, PPS904, and slice header 906 are types of network abstraction layer (NAL) units. A NAL unit is a syntactic structure that includes an indication of the type of data that follows (e.g., encoded video data). NAL units are classified into video coding layer (VCL) NAL units and non-VCL NAL units. A VCL NAL unit contains data representing the sample values ​​contained in the video picture, while a non-VCL NAL unit contains optional additional relevant information such as parameter sets (important header data that may apply to many VCL NAL units) and auxiliary extensions (timing information and other auxiliary data that may enhance the usefulness of the decoded video signal but are not necessary for decoding the sample values ​​contained in the video picture). A person skilled in the art will understand that bitstream 900 may also include other parameters and information when actually applied.

[0163] The image data 908 in Figure 9 includes data associated with an image or video being encoded or decoded. Image data 908 is sometimes simply referred to as the payload or data being carried in the bitstream 900. In one embodiment, image data 908 includes a CVS 914 (or CLVS) containing multiple pictures 910. The CVS 914 is the encoded video sequence of any encoded layer video sequence (CLVS) contained in the video bitstream 900. In particular, the CVS and CLVS are the same when the video bitstream 900 contains a single layer. The CVS and CLVS differ only when the video bitstream 900 contains multiple layers.

[0164] As shown in Figure 9, each slice of picture 910 may be contained in its own VCL's NAL unit 912. The set of VCL's NAL units 912 contained in CVS914 is sometimes called an access unit.

[0165] Figure 10 shows a partitioning method 1000 for picture 1010. Picture 1010 may be similar to any of the multiple pictures 910 shown in Figure 9. As shown, picture 1010 may be partitioned into multiple slices 1012. A slice is a spatially distinct region of a frame (e.g., a picture) that is encoded separately from any other region within the same frame. Three slices 1012 are shown in Figure 10, but more or fewer slices may be used in actual application. Each slice 1012 may be partitioned into multiple blocks 1014. The blocks 1014 in Figure 10 may be similar to the current block 811, preceding reference block 821, and succeeding reference block 831 in Figure 8. Blocks 1014 may represent CUs. Four blocks 1014 are shown in Figure 10, but more or fewer blocks may be used in actual application.

[0166] Each block 1014 may be divided into multiple samples 1016 (e.g., pixels). In one embodiment, the size of each block 1014 is measured in lumens. Figure 10 shows 16 samples 1016, but more or fewer samples may be used in actual applications.

[0167] In one embodiment, a conformance window 1060 is applied to picture 1010. As described above, the conformance window 1060 is used to crop, reduce, or otherwise modify the size of picture 1010 (e.g., a reconstructed / decoded picture) in the process of preparing the picture for output. For example, a decoder may apply the conformance window 1060 to picture 1010 to crop, trim, reduce, or otherwise modify the size of picture 1010 before the picture is output for display to the user. The size of the conformance window 1060 is determined by reducing the size of picture 1010 before output by applying conformance window upper offset 1062, conformance window lower offset 1064, conformance window left offset 1066, and conformance window right offset 1068 to picture 1010. In other words, only the portion of picture 1010 that exists within the conformance window 1060 is output. Thus, picture 1010 is cropped in size before output. In one embodiment, the first picture parameter set and the second picture parameter set each refer to the same sequence parameter set and have the same values ​​for picture width and picture height. Therefore, the first picture parameter set and the second picture parameter set also have the same values ​​for the conformance window.

[0168] Figure 11 shows one embodiment of a decoding method 1100 implemented by a video decoder (e.g., video decoder 400). Method 1100 may be performed after the bitstream to be decoded has been received directly or indirectly from a video encoder (e.g., video encoder 300). Method 1100 improves the decoding process by maintaining the same conformance window size for picture parameter sets having the same picture size. Thus, reference picture resampling (RPR) may remain enabled or turned on for the entire CVS. By maintaining a consistent conformance window size for picture parameter sets having the same picture size, encoding efficiency can be improved. Thus, in practice, the performance of the codec is improved, which leads to a more desirable user experience.

[0169] In block 1102, the video decoder receives a first picture parameter set (e.g., ppsA) and a second picture parameter set (e.g., ppsB), each referencing the same sequence parameter set. If the first and second picture parameter sets have the same values ​​with respect to picture width and picture height, then the first and second picture parameter sets have the same values ​​with respect to conformance window. In one embodiment, the picture width and picture height are measured in lumens.

[0170] In one embodiment, the picture width is specified as pic_width_in_luma_samples. In one embodiment, the picture height is specified as pic_height_in_luma_samples. In one embodiment, pic_width_in_luma_samples specifies the width of each decoded picture referencing the PPS in units of luma samples. In one embodiment, pic_height_in_luma_samples specifies the height of each decoded picture referencing the PPS in units of luma samples.

[0171] In one embodiment, the conformance window has a conformance window left offset, a conformance window right offset, a conformance window top offset, and a conformance window bottom offset, which together represent the conformance window size. In one embodiment, the conformance window left offset is specified as pps_conf_win_left_offset. In one embodiment, the conformance window right offset is specified as pps_conf_win_right_offset. In one embodiment, the conformance window top offset is specified as pps_conf_win_top_offset. In one embodiment, the conformance window bottom offset is specified as pps_conf_win_bottom_offset. In one embodiment, the conformance window size or conformance window value is signaled by PPS.

[0172] In block 1104, the video decoder applies a conformance window to the current picture corresponding to either the first or second picture parameter set. By doing so, the video encoder crops the current picture to the size of the conformance window.

[0173] In one embodiment, the method further comprises a step of decoding the current picture based on a resampled reference picture using interpretation. In one embodiment, the method further comprises a step of resampling a reference picture corresponding to the current picture using reference picture resampling (RPS). In one embodiment, resampling the reference picture changes the resolution of the reference picture.

[0174] In one embodiment, the method further includes a step of determining whether bidirectional optical flow (BDOF) is effective for decoding the picture based on the picture width, picture height, conformance window, and reference picture of the current picture. In one embodiment, the method further includes a step of determining whether decoder-side motion vector fine-tuning (DMVR) is effective for decoding the picture based on the picture width, picture height, conformance window, and reference picture of the current picture.

[0175] In one embodiment, the method further comprises the step of displaying an image generated using the current block on the display of an electronic device (e.g., a smartphone, tablet, laptop, personal computer, etc.).

[0176] Figure 12 shows one embodiment of method 1200 for encoding a video bitstream, performed by a video encoder (e.g., video encoder 300). Method 1200 may be performed when pictures (e.g., from video) are encoded into a video bitstream and then sent to a video decoder (e.g., video decoder 400). Method 1200 improves the encoding process by maintaining the same conformance window size for picture parameter sets having the same picture size. Thus, reference picture resampling (RPR) may remain enabled or turned on for the entire CVS. By maintaining a consistent conformance window size for picture parameter sets having the same picture size, encoding efficiency can be improved. Thus, in practice, codec performance improves, which leads to a more desirable user experience.

[0177] In block 1202, the video encoder generates a first picture parameter set and a second picture parameter set, each referencing the same sequence parameter set. If the first and second picture parameter sets have the same values ​​with respect to picture width and picture height, then the first and second picture parameter sets have the same values ​​with respect to conformance window. In one embodiment, the picture width and picture height are measured using lumens.

[0178] In one embodiment, the picture width is specified as pic_width_in_luma_samples. In one embodiment, the picture height is specified as pic_height_in_luma_samples. In one embodiment, pic_width_in_luma_samples defines the width of each decoded picture referencing the PPS in units of luma samples. In one embodiment, pic_height_in_luma_samples defines the height of each decoded picture referencing the PPS in units of luma samples.

[0179] In one embodiment, the conformance window has a conformance window left offset, a conformance window right offset, a conformance window top offset, and a conformance window bottom offset, which together represent the conformance window size. In one embodiment, the conformance window left offset is specified as pps_conf_win_left_offset. In one embodiment, the conformance window right offset is specified as pps_conf_win_right_offset. In one embodiment, the conformance window top offset is specified as pps_conf_win_top_offset. In one embodiment, the conformance window bottom offset is specified as pps_conf_win_bottom_offset. In one embodiment, the conformance window size or conformance window value is signaled by PPS.

[0180] In block 1204, the video encoder encodes the first picture parameter set and the second picture parameter set into a video bitstream. In block 1206, the video encoder stores the video bitstream for transmission to the video decoder. In one embodiment, the video encoder transmits the video bitstream containing the first picture parameter set and the second picture parameter set to the video decoder.

[0181] In one embodiment, a method for encoding a video bitstream is provided. The bitstream has a plurality of parameter sets and a plurality of pictures. Each of the plurality of pictures contains a plurality of slices. Each of the plurality of slices contains a plurality of encoding blocks. The method includes a step of generating a parameter set parameterSetA containing information including picture size picSizeA and conformance window confWinA, and writing it to the bitstream. This parameter may be a picture parameter set (PPS). The method further includes a step of generating another parameter set parameterSetB containing information including picture size picSizeB and conformance window confWinB, and writing it to the bitstream. This parameter may be a picture parameter set (PPS). The method further includes the steps of: if the values ​​of picSizeA in parameterSetA and picSizeB in parameterSetB are the same, then constraining the values ​​of the conformance window confWinA in parameterSetA and confWinB in parameterSetB to be the same; and if the values ​​of confWinA in parameterSetA and confWinB in parameterSetB are the same, then constraining the values ​​of the picture size picSizeA in parameterSetA and picSizeB in parameterSetB to be the same. The method further includes the step of encoding the bitstream.

[0182] In one embodiment, a method for decoding a video bitstream is provided. The bitstream has a plurality of parameter sets and a plurality of pictures. Each of the plurality of pictures contains a plurality of slices. Each of the plurality of slices contains a plurality of encoded blocks. The method comprises a step of analyzing the parameter set to obtain the picture size and conformance window size associated with the current picture currPic. The obtained information is used to derive the picture size and cropped size of the current picture. The method further comprises a step of analyzing another parameter set to obtain the picture size and conformance window size associated with a reference picture refPic. The obtained information is used to derive the picture size and cropped size of the reference picture. The method further includes the steps of: determining refPic as a reference picture for decoding the current block curBlock located within the current picture currPic; determining whether or not bidirectional optical flow (BDOF) is used or effective for decoding the current encoded block based on the picture size and conformance window of the current picture and the reference picture; and decoding the current block.

[0183] In one embodiment, if the picture size and conformance window of the current picture and the reference picture are different, the BDOF is not used or is invalid for decoding the current encoded block.

[0184] In one embodiment, a method for decoding a video bitstream is provided. The bitstream has a plurality of parameter sets and a plurality of pictures. Each of the plurality of pictures contains a plurality of slices. Each of the plurality of slices contains a plurality of encoded blocks. The method comprises a step of analyzing the parameter set to obtain the picture size and conformance window size associated with the current picture currPic. The obtained information is used to derive the picture size and cropped size of the current picture. The method further comprises a step of analyzing another parameter set to obtain the picture size and conformance window size associated with a reference picture refPic. The obtained information is used to derive the picture size and cropped size of the reference picture. This method further includes the steps of: determining refPic as a reference picture for decoding the current block curBlock located within the current picture currPic; determining whether decoder-side motion vector fine-tuning (DMVR) is used or effective for decoding the current encoded block based on the picture size and conformance window of the current picture and the reference picture; and decoding the current block.

[0185] In one embodiment, if the picture size and conformance window of the current picture and the reference picture are different, the DMVR is not used or is invalid for decoding the current encoded block.

[0186] In one embodiment, a method for encoding a video bitstream is provided. In one embodiment, the bitstream has a plurality of parameter sets and a plurality of pictures. Each of the plurality of pictures contains a plurality of slices. Each of the plurality of slices contains a plurality of encoding blocks. The method comprises the step of generating a parameter set that includes the picture size and conformance window size associated with the current picture currPic. This information is used to derive the picture size and cropped size of the current picture. The method further comprises the step of generating another parameter set that includes the picture size and conformance window size associated with a reference picture refPic. The obtained information is used to derive the picture size and cropped size of the reference picture. The method further comprises the step of restricting the time motion vector prediction (TMVP) of all slices belonging to the current picture currPic to the condition that the reference picture refPic must not be used as a reference picture at the same position if the picture sizes and conformance windows of the current picture and the reference picture are different. In other words, if the reference picture refPic is a reference picture at the same location for encoding a block contained in the current picture currPic for TMVP, then the picture size and conformance window of the current picture and the reference picture must be the same. This method further includes a step of decoding the bitstream.

[0187] In one embodiment, a method for decoding a video bitstream is provided. The bitstream has a plurality of parameter sets and a plurality of pictures. Each of the plurality of pictures contains a plurality of slices. Each of the plurality of slices contains a plurality of encoded blocks. The method comprises a step of analyzing the parameter set to obtain the picture size and conformance window size associated with the current picture currPic. The obtained information is used to derive the picture size and cropped size of the current picture. The method further comprises a step of analyzing another parameter set to obtain the picture size and conformance window size associated with a reference picture refPic. The obtained information is used to derive the picture size and cropped size of the reference picture. The method further includes the steps of determining refPic as the reference picture for decoding the current block curBlock located within the current picture currPic, and analyzing the syntax element (slice_DVMR_BDOF_enable_flag) to determine whether decoder-side motion vector fine-tuning (DMVR) and / or bidirectional optical flow (BDOF) are used or enabled for decoding the current encoded picture and slice. The method further includes the step of constraining the value of the syntax element (slice_DVMR_BDOF_enable_flag) to zero if the conformance window confWinA in parameterSetA and confWinB in parameterSetB are not the same, or if the values ​​of picSizeA in parameterSetA and picSizeB in parameterSetB are not the same.

[0188] The following explanation pertains to the base text and is a VVC working draft. That is, only increments are shown, and text included in the base text that is not described below remains unchanged. Deleted text is shown in italics, and added text is in bold.

[0189] The syntax and semantics of the sequence parameter set are provided. [Table 4]

[0190] The `max_width_in_luma_samples` setting specifies that the `pic_width_in_luma_samples` of any picture in which this SPS is active must be less than or equal to `max_width_in_luma_samples` as a requirement for bitstream conformance.

[0191] The `max_height_in_luma_samples` parameter specifies that the `pic_height_in_luma_samples` of any picture in which this SPS is active must be less than or equal to `max_height_in_luma_samples` as a requirement for bitstream conformance.

[0192] The syntax and semantics of the picture parameter set are provided. [Table 5]

[0193] `pic_width_in_luma_samples` specifies the width of each decoded picture referencing the PPS in units of luma samples. `pic_width_in_luma_samples` must not be equal to 0 and must be an integer multiple of `MinCbSizeY`.

[0194] `pic_height_in_luma_samples` specifies the height of each decoded picture referencing the PPS in units of luma samples. `pic_height_in_luma_samples` must not be equal to 0 and must be an integer multiple of `MinCbSizeY`.

[0195] For any active referenced picture whose width and height are reference_pic_width_in_luma_samples and reference_pic_height_in_luma_samples, all of the following conditions must be met for bitstream conformance to be valid.

[0196] ·2×pic_width_in_luma_samples≧reference_pic_width_in_luma_samples

[0197] ·2×pic_height_in_luma_samples≧reference_pic_height_in_luma_samples

[0198] ·pic_width_in_luma_samples≦8×reference_pic_width_in_luma_samples

[0199] ·pic_height_in_luma_samples≦8×reference_pic_height_in_luma_samples

[0200] The variables PicWidthInCtbsY, PicHeightInCtbsY, PicSizeInCtbsY, PicWidthInMinCbsY, PicHeightInMinCbsY, PicSizeInMinCbsY, PicSizeInSamplesY, PicWidthInSamplesC, and PicHeightInSamplesC are derived as follows:

[0201] ·PicWidthInCtbsY=Ceil(pic_width_in_luma_samples / CtbSizeY) (1)

[0202] ·PicHeightInCtbsY=Ceil(pic_height_in_luma_samples / CtbSizeY) (2)

[0203] ·PicSizeInCtbsY=PicWidthInCtbsY×PicHeightInCtbsY (3)

[0204] ·PicWidthInMinCbsY=pic_width_in_luma_samples / MinCbSizeY (4)

[0205] ·PicHeightInMinCbsY=pic_height_in_luma_samples / MinCbSizeY (5)

[0206] ·PicSizeInMinCbsY=PicWidthInMinCbsY×PicHeightInMinCbsY (6)

[0207] ·PicSizeInSamplesY=pic_width_in_luma_samples×pic_height_in_luma_samples (7)

[0208] ·PicWidthInSamplesC=pic_width_in_luma_samples / SubWidthC (8)

[0209] ·PicHeightInSamplesC=pic_height_in_luma_samples / SubHeightC (9)

[0210] A conformance_window_flag equal to 1 indicates that the conformance cropping window offset parameter is followed in PPS. A conformance_window_flag equal to 0 indicates that the conformance cropping window offset parameter does not exist.

[0211] conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset define the sample of the picture output from the decoding process, which references the PPS, by a rectangular area defined by picture coordinates for output. If conformance_window_flag is equal to 0, the values ​​of conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset are presumed to be equal to 0.

[0212] The conformance cropping window contains luma samples that have horizontal picture coordinates (including boundaries) from [SubWidthC×conf_win_left_offset] to [pic_width_in_luma_samples-(SubWidthC×conf_win_right_offset+1)] and vertical picture coordinates (including boundaries) from [SubHeightC×conf_win_top_offset] to [pic_height_in_luma_samples-(SubHeightC×conf_win_bottom_offset+1)].

[0213] The value of [SubWidthC × (conf_win_left_offset + conf_win_right_offset)] must be less than pic_width_in_luma_samples. Also, the value of [SubHeightC × (conf_win_top_offset + conf_win_bottom_offset)] must be less than pic_height_in_luma_samples.

[0214] The variables PicOutputWidthL and PicOutputHeightL are derived as follows:

[0215] ·PicOutputWidthL=pic_width_in_luma_samples-SubWidthC×(conf_win_right_offset+conf_win_left_offset) (10)

[0216] ·PicOutputHeightL=pic_height_in_pic_size_units-SubHeightC×(conf_win_bottom_offset+conf_win_top_offset) (11)

[0217] If ChromaArrayType is not equal to 0, the corresponding specified sample of the two chroma arrays is the sample with picture coordinates (x / SubWidthC, y / SubHeightC), where (x, y) are the picture coordinates of the specified chroma sample.

[0218] Note: The offset parameter of the conformance cropping window is only applied to the output. All internal decoding processes are applied to the uncropped picture size.

[0219] Assuming that PPS_A and PPS_B are picture parameter sets that refer to the same sequence parameter set, and that the values ​​of pic_width_in_luma_samples in PPS_A and PPS_B are the same, and that the values ​​of pic_height_in_luma_samples in PPS_A and PPS_B are the same, then all of the following conditions must be met for bitstream conformance to be true. The values ​​of conf_win_left_offset included in PPS_A and PPS_B are the same. The values ​​of conf_win_right_offset included in PPS_A and PPS_B are the same. The values ​​of conf_win_top_offset included in PPS_A and PPS_B are the same. The values ​​of conf_win_bottom_offset included in PPS_A and PPS_B are the same.

[0220] The following constraints are added to the semantics of collocated_ref_idx.

[0221] `collocated_ref_idx` defines the reference index of the same-location picture used for time motion vector prediction.

[0222] If slice_type is equal to P, or if slice_type is equal to B and collocated_from_l0_flag is equal to 1, collocated_ref_idx refers to the picture in list 0, and the value of collocated_ref_idx must be within the range of 0 to (NumRefIdxActive[0]-1) (including the boundary).

[0223] If slice_type is equal to B and collocated_from_l0_flag is equal to 0, collocated_ref_idx refers to the picture in List 1, and the value of collocated_ref_idx must be within the range of 0 to (NumRefIdxActive[1]-1) (including the boundary).

[0224] If collocated_ref_idx does not exist, its value is presumed to be equal to 0.

[0225] A requirement of bitstream conformance is that the picture referenced by collocated_ref_idx must be common to all slices of the encoded picture.

[0226] A requirement for bitstream conformance is that the resolution of the collocated_ref_idx and the referenced picture referenced by the current picture must be the same.

[0227] A requirement for bitstream conformance is that the picture size and conformance window of the referenced picture referenced by the collocated_ref_idx and the current picture must be the same.

[0228] The following conditions for setting dmvrFlag to 1 will be corrected.

[0229] • If all of the following conditions are met, dmvrFlag will be set to equal 1.

[0230] • sps_dmvr_enabled_flag is equal to 1.

[0231] • general_merge_flag[xCb][yCb] is equal to 1.

[0232] • Both predFlagL0[0][0] and predFlagL1[0][0] are equal to 1.

[0233] • mmvd_merge_flag[xCb][yCb] is equal to 0.

[0234] DiffPicOrderCnt(currPic,RefPicList[0][refIdxL0]) is equal to DiffPicOrderCnt(RefPicList[1][refIdxL1],currPic).

[0235] • BcwIdx[xCb][yCb] is equal to 0.

[0236] Both luma_weight_l0_flag[refIdxL0] and luma_weight_l1_flag[refIdxL1] are equal to 0.

[0237] • cbWidth is greater than or equal to 8.

[0238] • cbHeight is greater than or equal to 8.

[0239] • cbHeight × cbWidth is greater than or equal to 128.

[0240] If X is 0 and 1 respectively, the pic_width_in_luma_samples and pic_height_in_luma_samples of the reference picture refPicLX associated with refIdxLX are equal to the pic_width_in_luma_samples and pic_height_in_luma_samples of the current picture, respectively.

[0241] If X is 0 and 1 respectively, then pic_width_in_luma_samples, pic_height_in_luma_samples, conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset of the reference picture refPicLX associated with refIdxLX are equal to pic_width_in_luma_samples, pic_height_in_luma_samples, conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset of the current picture, respectively.

[0242] The following conditions for setting dmvrFlag to 1 will be corrected.

[0243] bdofFlag is set to TRUE if all of the following conditions are met.

[0244] • sps_bdof_enabled_flag is equal to 1.

[0245] • Both predFlagL0[xSbIdx][ySbIdx] and predFlagL1[xSbIdx][ySbIdx] are equal to 1.

[0246] • DiffPicOrderCnt(currPic,RefPicList[0][refIdxL0]) × DiffPicOrderCnt(currPic,RefPicList[1][refIdxL1]) is less than 0.

[0247] • MotionModelIdc[xCb][yCb] is equal to 0.

[0248] • merge_subblock_flag[xCb][yCb] is equal to 0.

[0249] • sym_mvd_flag[xCb][yCb] is equal to 0.

[0250] • BcwIdx[xCb][yCb] is equal to 0.

[0251] Both luma_weight_l0_flag[refIdxL0] and luma_weight_l1_flag[refIdxL1] are equal to 0.

[0252] • cbHeight is greater than or equal to 8.

[0253] If X is 0 and 1 respectively, the pic_width_in_luma_samples and pic_height_in_luma_samples of the reference picture refPicLX associated with refIdxLX are equal to the pic_width_in_luma_samples and pic_height_in_luma_samples of the current picture, respectively.

[0254] If X is 0 and 1 respectively, then pic_width_in_luma_samples, pic_height_in_luma_samples, conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset of the reference picture refPicLX associated with refIdxLX are equal to pic_width_in_luma_samples, pic_height_in_luma_samples, conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset of the current picture, respectively.

[0255] cIdx is equal to 0.

[0256] Figure 13 is a schematic diagram of a video encoding device 1300 (e.g., a video encoder 20 or a video decoder 30) according to one embodiment of the present disclosure. The video encoding device 1300 is suitable for implementing the disclosed embodiments as described herein. The video encoding device 1300 comprises an inlet port 1310 and a receiving unit (Rx) 1320 for receiving data, a processor, logic unit, or central processing unit (CPU) 1330 for processing the data, a transmitting unit (Tx) 1340 and an exit port 1350 for transmitting data, and a memory 1360 for storing data. The video encoding device 1300 may also comprise optical-to-electrical (OE) and electrical-to-optical (EO) conversion components for optical or electrical signal input or output, coupled to the inlet port 1310, the receiving unit 1320, the transmitting unit 1340, and the exit port 1350.

[0257] The processor 1330 is implemented by hardware and software. The processor 1330 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor 1330 communicates with the inlet port 1310, the receive unit 1320, the transmit unit 1340, the exit port 1350, and the memory 1360. The processor 1330 has an encoding module 1370. The encoding module 1370 implements the embodiments disclosed above. For example, the encoding module 1370 implements, processes, prepares, or provides various codec functions. Thus, the inclusion of the encoding module 1370 brings about a significant improvement in the functionality of the video encoding device 1300 and results in a change to another state of the video encoding device 1300. Alternatively, the encoding module 1370 may be implemented as an instruction stored in memory 1360 and executed by processor 1330.

[0258] The video encoding device 1300 may also include an input and / or output (I / O) device 1380 for exchanging data with the user. The I / O device 1380 may include output devices such as a display for displaying video data and a speaker for outputting audio data. The I / O device 1380 may also include input devices such as a keyboard, mouse, or trackball, and / or corresponding interfaces for interacting with such output devices.

[0259] The memory 1360 includes one or more disks, tape drives, and solid-state drives, and may be used as an overflow data storage device to store such programs when they are selected for execution, and to store instructions and data read during program execution. The memory 1360 may be volatile and / or non-volatile, and may be read-only memory (ROM), random-access memory (RAM), tri-level associative memory (TCAM), and / or static random-access memory (SRAM).

[0260] Figure 14 is a schematic diagram of one embodiment relating to the encoding means 1400. In one embodiment, the encoding means 1400 is implemented in a video encoding device 1402 (e.g., a video encoder 20 or a video decoder 30). The video encoding device 1402 includes a receiving means 1401. The receiving means 1401 is configured to receive a picture to encode or a bitstream to decode. The video encoding device 1402 includes a transmitting means 1407 coupled to the receiving means 1401. The transmitting means 1407 is configured to transmit a bitstream to a decoder or to transmit a decoded image to a display means (e.g., one of a plurality of I / O devices 1380).

[0261] The video encoding device 1402 includes a storage means 1403. The storage means 1403 is coupled to at least one of the receiving means 1401 or the transmitting means 1407. The storage means 1403 is configured to store instructions. The video encoding device 1402 also includes a processing means 1405. The processing means 1405 is coupled to the storage means 1403. The processing means 1405 is configured to execute instructions stored in the storage means 1403 to perform the method disclosed herein.

[0262] It should also be understood that the steps of the exemplary methods described herein do not necessarily have to be performed in the order described, and the order of the steps of such methods should be understood to be merely illustrative. Similarly, in methods consistent with various embodiments of this disclosure, additional steps may be included in such methods, and certain steps may be omitted or combined.

[0263] While this disclosure provides several embodiments, it should be understood that the disclosed systems and methods can be embodied in numerous other specific forms without departing from the spirit or scope of this disclosure. These embodiments should be considered illustrative and not limiting, and their purpose is not to limit them to the details shown herein. For example, various elements or components may be combined or integrated into other systems, or certain features may be omitted or not implemented.

[0264] Furthermore, techniques, systems, subsystems, and methods described and shown as separate or distinct in various embodiments may be combined with or integrated with other systems, modules, techniques, or methods without departing from the scope of this disclosure. Other things shown or described as being coupled to one another, directly coupled to one another, or communicating to one another may communicate, whether electrically, mechanically, or otherwise, by being indirectly coupled through some interface, device, or intermediate component. Other examples of modifications, substitutions, and alterations are readily apparent to those skilled in the art, and such examples may be made without departing from the spirit and scope disclosed herein. [Other possible items] (Item 1) A decoding method performed by a video decoder, A step in which the video decoder receives a first picture parameter set and a second picture parameter set that each refer to the same sequence parameter set, and when the first picture parameter set and the second picture parameter set have the same values with respect to picture width and picture height, the first picture parameter set and the second picture parameter set have the same values with respect to the compliance window; the receiving step, A step in which the video decoder applies the compliance window to the current picture corresponding to the first picture parameter set or the second picture parameter set; A method comprising: (Item 2) The method according to item 1, wherein the compliance window has a compliance window left offset, a compliance window right offset, a compliance window upper offset, and a compliance window lower offset. (Item 3) The method according to any one of items 1 to 2, further comprising a step of decoding the current picture corresponding to the first picture parameter set or the second picture parameter set using inter prediction after the compliance window is applied, wherein the inter prediction is based on a resampled reference picture. (Item 4) The method according to any one of items 1 to 3, further comprising a step of resampling a reference picture associated with the current picture corresponding to the first picture set or the second picture set using reference picture resampling (RPS). (Item 5) The method according to any one of items 1 to 4, wherein the resampling of the reference picture changes the resolution of the reference picture used to inter predict the current picture corresponding to the first picture set or the second picture set. (Item 6) The method according to any one of items 1 to 5, wherein the picture width and the picture height are measured with a lumen sample. (Item 7) The method further comprises a step of determining whether bidirectional optical flow (BDOF) is effective for decoding the picture based on the picture width, the picture height of the current picture, and the conformity window and the reference picture of the current picture, according to any one of items 1 to 6. (Item 8) The method further comprises a step of determining whether decoder-side motion vector fine-tuning (DMVR) is effective for decoding the picture based on the picture width, the picture height of the current picture, and the conformity window and the reference picture of the current picture, according to any one of items 1 to 6. (Item 9) The method further comprises a step of displaying an image generated using the current block on a display of an electronic device, according to any one of items 1 to 6. (Item 10) A method of encoding performed by a video encoder, a step of the video encoder generating a first picture parameter set and a second picture parameter set each referring to the same sequence parameter set, wherein when the first picture parameter set and the second picture parameter set have the same values with respect to picture width and picture height, the first picture parameter set and the second picture parameter set have the same values with respect to the conformity window; generating step; a step of the video encoder encoding the first picture parameter set and the second picture parameter set into a video bitstream; a step of the video encoder storing the video bitstream for transmission to a video decoder; A method comprising. (Item 11) The method according to item 10, wherein the conformance window has a conformance window left offset, a conformance window right offset, a conformance window upper offset, and a conformance window lower offset. (Item 12) The method according to any one of items 10 to 11, wherein the picture width and the picture height are measured in a lumen sample. (Item 13) The method according to any one of items 10 to 12, further comprising the step of transmitting the video bitstream, which includes the first picture parameter set and the second picture parameter set, to the video decoder. (Item 14) A receiver configured to receive an encoded video bitstream, A memory coupled to the receiver, wherein the memory stores instructions, A processor coupled to the memory, wherein the processor executes the instruction to the decoding device, Receiving a first picture parameter set and a second picture parameter set, each referencing the same sequence parameter set, wherein if the first picture parameter set and the second picture parameter set have the same values ​​with respect to picture width and picture height, then the first picture parameter set and the second picture parameter set have the same values ​​with respect to conformance window, Applying the conformance window to the current picture corresponding to the first picture parameter set or the second picture parameter set, A processor configured to perform the following: A decoding device equipped with the following features. (Item 15) The decoding device according to item 14, wherein the conformance window has a conformance window left offset, a conformance window right offset, a conformance window upper offset, and a conformance window lower offset. (Item 16) The decoding device further comprises decoding the current picture corresponding to the first picture parameter set or the second picture parameter set using interpretation after the conformance window has been applied, wherein the interpretation is based on a resampled reference picture, according to any one of items 14 to 15. (Item 17) The decoding device according to any one of items 15 to 16, further comprising a display configured to display an image generated based on the current picture. (Item 18) An encoding device, Memory containing instructions, A processor coupled to the memory, wherein the processor executes the instructions to the encoding device, To generate a first picture parameter set and a second picture parameter set, each referencing the same sequence parameter set, wherein if the first picture parameter set and the second picture parameter set have the same values ​​with respect to picture width and picture height, the first picture parameter set and the second picture parameter set have the same values ​​with respect to conformance window, Encoding the first picture parameter set and the second picture parameter set into a video bitstream, A processor configured to perform the following: A transmitter coupled to the processor, the transmitter being configured to transmit the video bitstream, which includes the first picture parameter set and the second picture parameter set, to a video decoder, An encoding device equipped with the following features. (Item 19) The encoding device according to item 18, wherein the conformance window has a conformance window left offset, a conformance window right offset, a conformance window upper offset, and a conformance window lower offset. (Item 20) An encoding device according to any one of items 18 to 19, wherein the picture width and picture height are measured by a lumen sample. (Item 21) A receiver configured to receive a picture to encode or a bitstream to decode, A transmitter coupled to the receiver, wherein the transmitter is configured to transmit the bitstream to a decoder or to transmit the decoded image to a display, A memory coupled to at least one of the receiver or the transmitter, wherein the memory is configured to store instructions, A processor coupled to the memory, wherein the processor is configured to execute the instructions stored in the memory to perform any of the methods in items 1 to 9 and any of items 10 to 13, An encoding device equipped with the following features. (Item 22) The encoding apparatus according to item 20, further comprising a display configured to display an image. (Item 23) Encoder and A decoder that communicates with the encoder, A system comprising the encoder or decoder, wherein the encoder or decoder includes a decoding device, encoding device, or coding device as described in any of items 15 to 22. (Item 24) A means for encoding, A receiving means configured to receive a picture to encode or a bitstream to decode, A transmitting means coupled to the receiving means, wherein the transmitting means is configured to transmit the bitstream to a decoding means or to transmit the decoded image to a display means, A storage means coupled to at least one of the receiving means or the transmitting means, wherein the storage means is configured to store instructions, A processing means coupled to the storage means, wherein the processing means is configured to execute the instructions stored in the storage means to perform any of items 1 to 9 and any of items 10 to 13, A means for encoding that includes the following:

Claims

1. A method for decoding a bitstream, A step of receiving a bitstream, wherein the bitstream includes a first picture parameter set (PPS) and a second PPS, each referencing the same sequence parameter set (SPS), the first PPS includes a first picture width parameter, a first picture height parameter, and a first conformance window flag, the second PPS includes a second picture width parameter, a second picture height parameter, and a second conformance window flag, the first conformance window flag being equal to 1 indicates that the first PPS further includes a first conformance window parameter, and the second conformance window flag being equal to 1 indicates that the second PPS further includes a second conformance window parameter, and The steps include analyzing the first picture width parameter and the first picture height parameter from the first PPS of the bitstream, The steps include analyzing the second picture width parameter and the second picture height parameter from the second PPS of the bitstream, If the first picture width parameter has the same value as the second picture width parameter and the first picture height parameter has the same value as the second picture height parameter, then it is determined that the first conformance window parameter has the same value as the second conformance window parameter, wherein the first conformance window parameter defines the samples output from the decoding process among the pictures referencing the first PPS by a rectangular region defined by picture coordinates for output, and the second conformance window parameter defines the samples output from the decoding process among the pictures referencing the second PPS by a rectangular region defined by picture coordinates for output. A method for providing this.

2. The first conformance window parameter or the second conformance window parameter includes a conformance window left offset, a conformance window right offset, a conformance window upper offset, and a conformance window lower offset. The method according to claim 1.

3. The first picture width parameter or the second picture width parameter is measured by a lumen sample, and the first picture height parameter or the second picture height parameter is measured by a lumen sample. The method according to claim 1 or 2.

4. A method for encoding a bitstream, wherein the method is The steps include encoding the sequence parameter set (SPS) into the bitstream, A step of encoding a first picture parameter set (PPS) and a second PPS into the bitstream, wherein the first PPS and the second PPS each refer to the same SPS, the first PPS includes a first picture width parameter, a first picture height parameter, and a first conformance window flag, the second PPS includes a second picture width parameter, a second picture height parameter, and a second conformance window flag, the first conformance window flag being equal to 1 indicates that the first PPS further includes a first conformance window parameter, and the second conformance window flag being equal to 1 indicates that the second PPS further includes a second conformance window parameter The steps include: if the first picture width parameter has the same value as the second picture width parameter and the first picture height parameter has the same value as the second picture height parameter, then the first conformance window parameter has the same value as the second conformance window parameter, the first conformance window parameter defines the samples output from the decoding process of the picture referencing the first PPS by a rectangular region defined by picture coordinates for output, and the second conformance window parameter defines the samples output from the decoding process of the picture referencing the second PPS by a rectangular region defined by picture coordinates for output. A method for providing this.

5. The first conformance window parameter or the second conformance window parameter includes a conformance window left offset, a conformance window right offset, a conformance window upper offset, and a conformance window lower offset. The method according to claim 4.

6. The first picture width parameter or the second picture width parameter is measured in a lumen sample. The first picture height parameter or the second picture height parameter is measured in a lumen sample. The method according to claim 4 or 5.

7. A decoding device comprising a processor and memory, wherein the memory is configured to store instructions, and the processor is configured to execute the instructions in the memory to perform the method according to any one of claims 1 to 3. Decoding device.

8. An encoding device comprising a processor and memory, wherein the memory is configured to store instructions, and the processor is configured to execute the instructions in the memory to perform the method according to any one of claims 4 to 6. Encoding device.

9. A device for storing a bitstream, the device comprising at least one storage medium and at least one communication interface, The at least one communication interface is configured to receive or transmit the bitstream, The at least one storage medium is configured to store the bitstream, The bitstream includes a first picture parameter set (PPS) and a second PPS, each referencing the same sequence parameter set (SPS), wherein the first PPS includes a first picture width parameter, a first picture height parameter, and a first conformance window flag, and the second PPS includes a second picture width parameter, a second picture height parameter, and a second conformance window flag, wherein the first conformance window flag equal to 1 indicates that the first PPS further includes a first conformance window parameter, and the second conformance window flag equal to 1 indicates that the second PPS further includes a second conformance window parameter. If the first picture width parameter has the same value as the second picture width parameter, and the first picture height parameter has the same value as the second picture height parameter, then the first conformance window parameter has the same value as the second conformance window parameter. The first conformance window parameter defines the samples output from the decoding process among the pictures that reference the first PPS, by a rectangular region defined by picture coordinates for output. The second conformance window parameter defines the samples output from the decoding process from the picture that references the second PPS, by a rectangular region defined by picture coordinates for output. device.

10. A method for storing a bitstream, The steps include receiving or transmitting the bitstream via a communication interface, A step of storing the bitstream in one or more storage media, wherein the bitstream includes a first picture parameter set (PPS) and a second PPS, each referring to the same sequence parameter set (SPS), the first PPS includes a first picture width parameter, a first picture height parameter, and a first conformance window flag, the second PPS includes a second picture width parameter, a second picture height parameter, and a second conformance window flag, the first conformance window flag being equal to 1 indicates that the first PPS further includes a first conformance window parameter, and the second conformance window flag being equal to 1 indicates that the second PPS further includes a second conformance window parameter. Equipped with, If the first picture width parameter has the same value as the second picture width parameter, and the first picture height parameter has the same value as the second picture height parameter, then the first conformance window parameter has the same value as the second conformance window parameter. The first conformance window parameter defines the samples output from the decoding process among the pictures that reference the first PPS, by a rectangular region defined by picture coordinates for output. The second conformance window parameter defines the samples output from the decoding process from the picture that references the second PPS, by a rectangular region defined by picture coordinates for output. method.

11. A device for transmitting a bitstream, wherein the device is At least one storage medium configured to store the bitstream, wherein the bitstream includes a first picture parameter set (PPS) and a second PPS, each referencing the same sequence parameter set (SPS), the first PPS includes a first picture width parameter, a first picture height parameter, and a first conformance window flag, the second PPS includes a second picture width parameter, a second picture height parameter, and a second conformance window flag, the first conformance window flag being equal to 1 indicates that the first PPS further includes a first conformance window parameter, and the second conformance window flag being equal to 1 indicates that the second PPS includes a second conformance window parameter The present invention further includes a meter, wherein if the first picture width parameter has the same value as the second picture width parameter and the first picture height parameter has the same value as the second picture height parameter, the first conformance window parameter has the same value as the second conformance window parameter, the first conformance window parameter defines a rectangular region defined by picture coordinates for output, which is a sample of the picture referencing the first PPS that is output from the decoding process, and the second conformance window parameter defines a rectangular region defined by picture coordinates for output, which is a sample of the picture referencing the second PPS that is output from the decoding process, and the present invention further includes at least one storage medium, A processor configured to acquire one or more bitstreams from one of the at least one storage mediums, A transmitter configured to transmit one or more bitstreams to a destination device A device equipped with the following features.

12. The device according to claim 11, wherein the first conformance window parameter or the second conformance window parameter includes a conformance window left offset, a conformance window right offset, a conformance window upper offset, and a conformance window lower offset.

13. A method for transmitting a bitstream, A step of storing the bitstream in at least one storage medium, wherein the bitstream includes a first picture parameter set (PPS) and a second PPS, each referring to the same sequence parameter set (SPS), the first PPS including a first picture width parameter, a first picture height parameter, and a first conformance window flag, and the second PPS including a second picture width parameter, a second picture height parameter, and a second conformance window flag, A first conformance window flag equal to 1 indicates that the first PPS further includes a first conformance window parameter, a second conformance window flag equal to 1 indicates that the second PPS further includes a second conformance window parameter, and if the first picture width parameter has the same value as the second picture width parameter and the first picture height parameter has the same value as the second picture height parameter, then the first conformance window parameter has the same value as the second conformance window parameter, the first conformance window parameter defines the samples of the picture referencing the first PPS that are output from the decoding process by a rectangular region defined by picture coordinates for output, and the second conformance window parameter defines the samples of the picture referencing the second PPS that are output from the decoding process by a rectangular region defined by picture coordinates for output, The steps include obtaining one or more bitstreams from one of the at least one storage mediums and transmitting the one or more bitstreams to a destination device. A method for providing this.

14. The first conformance window parameter or the second conformance window parameter includes a conformance window left offset, a conformance window right offset, a conformance window upper offset, and a conformance window lower offset. The method according to claim 13.

15. A program, which, when executed by a processor or computer, causes the processor or computer to perform the method described in any one of claims 1 to 3 or any one of claims 4 to 6. program.

16. A computer program that causes a video decoding device to perform the method according to any one of claims 1 to 3.

17. A computer program that causes a video encoding device to perform the method according to any one of claims 4 to 6.