Prediction weighted table-based image / video coding method and apparatus
The method optimizes image/video compression by parsing prediction weight tables from headers based on flags, addressing the need for efficient encoding/decoding of high-resolution and immersive media, thereby reducing signaling overhead and improving compression efficiency.
Patent Information
- Application Number
- JP2025151147
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-12-20
- Filing Date
- 2025-09-11
- Publication Date
- 2025-12-05
AI Technical Summary
The increasing demand for high-resolution and high-quality images/videos, such as 4K or 8K UHD, VR, AR, and holograms, necessitates highly efficient image/video compression technology to reduce transmission and storage costs while effectively handling diverse image characteristics.
A method and apparatus for encoding/decoding images/videos using a prediction weight table syntax, allowing the prediction weight table to be parsed from either a picture header or a slice header based on a flag, thereby optimizing signaling and reducing overhead.
Improves image/video compression efficiency by efficiently signaling prediction weight table syntax and reducing unnecessary signaling, leading to enhanced compression performance.
Smart Images

Figure 2025178278000001_ABST
Abstract
Description
[Technical Field]
[0001] The present technology relates to a method and apparatus for encoding / decoding image / video based on a prediction weighted table syntax. [Background technology]
[0002] Recently, demand for high-resolution, high-quality images / videos of 4K or 8K or higher, such as UHD (Ultra High Definition) images / videos, is increasing in various fields. As the resolution and quality of image / video data increases, the amount of information or bits to be transmitted increases relatively compared to existing image / video data, which increases transmission and storage costs when transmitting image data using existing media such as wired or wireless broadband lines or storing image / video data using existing storage media.
[0003] In addition, interest in and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have been increasing in recent years, and the broadcast of images / videos with different image characteristics from real images, such as game images, has been increasing.
[0004] Therefore, there is a need for highly efficient image / video compression technology to effectively compress and transmit, store, and play back high-resolution, high-quality image / video information having the above-mentioned various characteristics. Summary of the Invention [Problem to be solved by the invention]
[0005] The technical problem of this document is to provide a method and apparatus for improving image / video coding efficiency.
[0006] Another technical problem of this document is to provide a method and apparatus for efficiently signaling prediction weight table syntax.
[0007] Another technical problem of this document is to provide a method and apparatus for reducing signaling overhead related to weighted prediction. [Means for solving the problem]
[0008] According to one embodiment of the present document, a video decoding method performed by a video decoding device includes the steps of parsing a flag related to weighted prediction from a bitstream, parsing a prediction weight table syntax from the bitstream based on the flag, and performing weighted prediction on a current block in a current picture based on the prediction weight table syntax to reconstruct the current picture, wherein based on the value of the flag being 1, the prediction weight table syntax can be parsed from a picture header of the bitstream, and based on the value of the flag being 0, the prediction weight table syntax can be parsed from a slice header of the bitstream.
[0009] According to another embodiment of the present document, a video encoding method performed by a video encoding device includes the steps of deriving motion information of a current block, performing weighted prediction on the current block based on the motion information, and encoding video information including a flag related to the weighted prediction and a prediction weight table syntax, wherein the value of the flag can be set to 1 based on the prediction weight table syntax being included in a picture header of the video information, and the value of the flag can be set to 0 based on the prediction weight table syntax being included in a slice header of the video information.
[0010] According to yet another embodiment of the present document, there is provided a computer-readable digital storage medium, the digital storage medium including information for causing a video decoding device to perform a video decoding method, the video decoding method including the steps of parsing a flag related to weighted prediction from video information, parsing a prediction weight table syntax from the video information based on the flag, and performing weighted prediction on a current block in a current picture based on the prediction weight table syntax to reconstruct the current picture, wherein based on the value of the flag being 1, the prediction weight table syntax can be parsed from a picture header of the bitstream, and based on the value of the flag being 0, the prediction weight table syntax can be parsed from a slice header of the bitstream. [Effects of the Invention]
[0011] According to one embodiment of this document, the overall image / video compression efficiency can be improved.
[0012] According to one embodiment of this document, prediction weight table syntax can be signaled efficiently.
[0013] According to one embodiment of the present document, unnecessary signaling can be reduced when transmitting information related to weighted prediction. [Brief explanation of the drawings]
[0014] [Figure 1] 1 illustrates schematically an example of a video / image coding system to which embodiments of the present document may be applied. [Figure 2] 1 is a diagram illustrating the configuration of a video / image encoding device to which an embodiment of this document can be applied. [Figure 3] 1 is a diagram illustrating the configuration of a video / image decoding device to which an embodiment of the present document can be applied. [Figure 4] 1 illustrates an example of a video / image encoding method based on inter prediction. [Figure 5] 1 illustrates an example of a video / image decoding method based on inter prediction. [Figure 6] 1 illustrates an example of a video / image encoding method and associated components according to embodiments of the present document. [Figure 7] 1 illustrates an example of a video / image encoding method and associated components according to embodiments of the present document. [Figure 8] 1 illustrates an example of a video / image decoding method and related components according to an embodiment of the present document. [Figure 9] 1 illustrates an example of a video / image decoding method and related components according to an embodiment of the present document. [Figure 10] 1 illustrates an example of a content streaming system to which the embodiments disclosed herein are applicable. DETAILED DESCRIPTION OF THE INVENTION
[0015] Because the disclosure of this document can be modified in various ways and can have various embodiments, specific embodiments will be illustrated in the drawings and described in detail. The terms used in this document are used merely to describe specific embodiments and are not intended to limit the technical ideas of this document. The singular expressions "a," "an," "an," "the," and the like include the expression "at least one" unless the context clearly indicates otherwise. In this document, the terms "comprise," "have," and the like are intended to specify the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, and should be understood not to preclude the presence or possibility of addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0016] Meanwhile, each component in the drawings described in this document is illustrated independently for the convenience of describing different characteristic functions, and does not mean that each component is realized by separate hardware or software. For example, two or more components may be combined to form a single component, or a single component may be divided into multiple components. Embodiments in which each component is integrated and / or separated are also included within the scope of this document as long as they do not deviate from the essence of the method disclosed herein.
[0017] In this document, " / " and "," should be interpreted to indicate "and / or." For example, "A / B" is interpreted as "A and / or B," and "A, B" is interpreted as "A and / or B." Additionally, "A / B / C" means "at least one of A, B, and / or C." Also, "A, B, C" means "at least one of A, B, and / or C." (In this document, the term " / " and "," should be interpreted to indicate "and / or." For instance, the expression "A / B" may mean "A and / or B." Further, "A, B" may mean "A and / or B." Further, "A / B / C" may mean "at least one of A, B, and / or C." Also, "A / B / C" may mean "at least one of A, B, and / or C.")
[0018] Additionally, in this document, "or" should be interpreted as "and / or." For example, "A or B" may mean 1) only "A," 2) only "B," or 3) "A and B." In other words, the term "or" in this document should be interpreted to indicate "and / or." For instance, the expression "A or B" may comprise 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted to indicate "additionally or alternatively."
[0019] Furthermore, parentheses used herein may mean "for example." Specifically, when "prediction (intra prediction)" is displayed, "intra prediction" may be suggested as an example of "prediction." Alternatively, "prediction" in this specification is not limited to "intra prediction," and "intra prediction" may be suggested as an example of "prediction." Furthermore, when "prediction (i.e., intra prediction)" is displayed, "intra prediction" may be suggested as an example of "prediction."
[0020] In this document, technical features individually described in one drawing may be embodied individually or simultaneously.
[0021] Hereinafter, the embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Hereinafter, the same reference numerals will be used to refer to the same components in the drawings, and redundant description of the same components will be omitted.
[0022] FIG. 1 illustrates a schematic diagram of an example video / image coding system to which embodiments of this document may be applied.
[0023] As shown in Figure 1, a video / image coding system includes a first device (source device) and a second device (receiving device). The source device can transmit encoded video / image information or data to the receiving device in file or streaming form via a digital storage medium or a network.
[0024] The source device may include a video source, an encoding device, and a transmitting unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / video encoding device, and the decoding device may be referred to as a video / video decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, which may be a separate device or an external component.
[0025] A video source can acquire video / images through a video / image capture, synthesis, or generation process. A video source can include a video / image capture device and / or a video / image generation device. A video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, a virtual video / image can be generated via a computer, etc., in which case the video / image capture process can be replaced with a process in which related data is generated.
[0026] An encoding device can encode input video / images. The encoding device can perform a series of procedures such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0027] The transmitter can transmit the encoded video / image information or data output in the form of a bitstream to a receiver of a receiving device via a digital storage medium or a network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include elements for generating a media file in a predetermined file format and elements for transmission via a broadcasting / communication network. The receiver can receive / extract the bitstream and transmit it to a decoding device.
[0028] The decoding device can decode the video / image by performing a series of steps such as inverse quantization, inverse transform, and prediction, which correspond to the operations of the encoding device.
[0029] The renderer can render the decoded video / image, and the rendered video / image can be displayed via a display unit.
[0030] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document may be applied to methods disclosed in the VVC (versatile video coding) standard. The methods / embodiments disclosed in this document may also be applied to methods disclosed in the EVC (essential video coding) standard, the AV1 (AOMedia Video 1) standard, the AVS2 (2nd generation audio video coding standard), or next-generation video / image coding standards (e.g., H.267, H.268, etc.).
[0031] Various embodiments relating to video / image coding are presented in this document, and unless otherwise stated, the embodiments may be performed in combination with each other.
[0032] In this document, video can refer to a collection of a series of images over time. A picture generally refers to a unit that shows one image at a specific time, and a slice / tile is a unit that constitutes part of a picture in coding. A slice / tile contains one or more CTUs (coding tree units). A picture is made up of one or more slices / tiles. A picture is made up of one or more tile groups. A tile group contains one or more tiles. A brick refers to a rectangular area of a CTU row within a tile in a picture. A tile is partitioned into multiple bricks, and each brick is made up of one or more CTU rows within the tile. A tile that is not partitioned into multiple bricks may also be called a brick. Brick scan refers to a specific sequential ordering of CTUs that partition a picture, where the CTUs are aligned within a brick by a CTU raster scan, the bricks within a tile are aligned consecutively by a raster scan of the bricks in the tile, and the tiles within a picture are aligned consecutively by a raster scan of the tiles in the picture. A tile is a specific tile column and a rectangular region of a CTU within the specific tile column. The tile column is a rectangular region of a CTU, where the rectangular region has a height equal to the height of the picture and a width specified by a syntax element in a picture parameter set. The tile row is a rectangular region of a CTU, where the rectangular region has a width specified by a syntax element in a picture parameter set and a height that may be equal to the height of the picture. Tile scan refers to a specific sequential ordering of CTUs that partition a picture, where the CTUs are aligned consecutively by a CTU raster scan within a tile, and the tiles within a picture are aligned consecutively by a raster scan of the tiles in the picture. A slice contains an integer number of bricks of a picture, and the integer number of bricks is contained in one NAL unit. A slice can consist of multiple complete tiles, or it can be a contiguous sequence of complete bricks of one tile. In this document, the terms tile group and slice may be used interchangeably.For example, in this document, a tile group / tile group header may be referred to as a slice / slice header.
[0033] A pixel or a pel may refer to the smallest unit constituting one picture (or image). A "sample" may also be used as a term corresponding to a pixel. A sample may generally refer to a pixel or a pixel value, may refer to only a pixel / pixel value of a luma component, or may refer to only a pixel / pixel value of a chroma component. Alternatively, a sample may refer to a pixel value in the spatial domain, or may refer to a transform coefficient in the frequency domain when such a pixel value is transformed into the frequency domain.
[0034] A unit may refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. The term unit may be used interchangeably with terms such as block or area. In general, an M×N block may include samples (or a sample array) consisting of M columns and N rows, or a set (or an array) of transform coefficients.
[0035] The term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block can refer to a set of samples or transform coefficients consisting of M columns and N rows. A sample generally refers to a pixel or pixel value, and can refer to only a pixel / pixel value of a luma component or only a pixel / pixel value of a chroma component. A sample can also be used as a term corresponding to one pixel or pel of a picture (or image).
[0036] 2 is a diagram illustrating the configuration of a video / image encoding device to which the embodiments of this document can be applied. Hereinafter, the term "video encoding device" includes the image encoding device.
[0037] As shown in FIG. 2, the encoding apparatus 200 may include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter predictor 221 and an intra predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. The image dividing unit 210, the predicting unit 220, the residual processing unit 230, the entropy encoding unit 240, the adding unit 250, and the filtering unit 260 may be configured by one or more hardware components (e.g., an encoder chipset or a processor) depending on the embodiment. Also, the memory 270 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.
[0038] The image division unit 210 may divide an input image (or picture, frame) input to the encoding device 200 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, the coding units may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) using a quad-tree, binary-tree, ternary-tree (QTBTTT) structure. For example, one coding unit may be divided into multiple coding units of deeper depths based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and then the binary tree structure and / or the ternary structure may be applied. Alternatively, the binary tree structure may be applied first. The coding procedure according to the present disclosure may be performed based on a final coding unit that is not further divided. In this case, the largest coding unit may be used as the final coding unit based on coding efficiency according to image characteristics, or the coding unit may be recursively divided into coding units of lower depths as needed, and a coding unit of an optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may each be divided or partitioned from the final coding unit.The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0039] The encoding apparatus 200 subtracts a prediction signal (predicted block, prediction sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from an input video signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, a unit in the encoder 200 that subtracts the prediction signal (predicted block, prediction sample array) from the input video signal (original block, original sample array) is called a subtraction unit 231. The prediction unit 220 may perform prediction on a current block (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit 220 determines whether intra prediction or inter prediction is to be applied for the current block or for each CU. The prediction unit 220 may generate various information related to prediction, such as prediction mode information, as will be described later in the description of each prediction mode, and transmit the information to the entropy encoding unit 240. The prediction information can be encoded in the entropy encoding unit 240 and output in the form of a bitstream.
[0040] The intra prediction unit 222 may predict the current block by referring to samples in the current picture. The referenced samples may be located adjacent to or distant from the current block depending on the prediction mode. Prediction modes in intra prediction may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, DC mode and planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the granularity of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the settings. The intra prediction unit 222 may also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.
[0041] The inter prediction unit 221 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction direction (such as L0 prediction, L1 prediction, or Bi prediction). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, a collocated CU (col CU), or the like, and the reference picture including the temporal neighboring block may be called a collocated picture (colPic). For example, the inter predictor 221 may construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive a motion vector and / or a reference picture index for the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of a skip mode or a merge mode, the inter predictor 221 may use motion information of neighboring blocks as motion information of the current block. In the case of the skip mode, unlike the merge mode, a residual signal may not be transmitted.In the case of motion vector prediction (MVP) mode, the motion vector of the neighboring block can be used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0042] The prediction unit 220 generates a prediction signal based on various prediction methods, which will be described later. For example, the prediction unit 200 can apply intra prediction or inter prediction for predicting a block, or can simultaneously apply intra prediction and inter prediction. This is called combined inter and intra prediction (CIIP). The prediction unit can also use intra block copy (IBC) prediction mode or palette mode for predicting a block. The IBC prediction mode or palette mode can be used for content image / video coding, such as games, as in screen content coding (SCC). IBC basically performs prediction within a current picture, but is similar to inter prediction in that it derives a reference block within the current picture. That is, IBC can use at least one of the inter prediction techniques described in this document. Palette mode can be considered an example of intra coding or intra prediction. When palette mode is applied, sample values within a picture can be signaled based on information about a palette table and a palette index.
[0043] The prediction signal generated via the prediction unit (including the inter prediction unit 221 and / or the intra prediction unit 222) can be used to generate a reconstructed signal or a residual signal.
[0044] The transform unit 232 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a graph-based transform (GBT), or a conditionally non-linear transform (CNT). Here, GBT refers to a transform obtained from a graph representing inter-pixel relationship information. CNT refers to a transform obtained based on a predicted signal generated using all previously reconstructed pixels. The transform process may be applied to pixel blocks having the same square size or to non-square blocks of variable size.
[0045] The quantization unit 233 quantizes the transform coefficients and transmits the quantized signal to the entropy encoding unit 240. The entropy encoding unit 240 encodes the quantized signal (information about the quantized transform coefficients) and outputs it as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantization unit 233 may rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on a coefficient scan order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0046] The entropy encoding unit 240 can perform various encoding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoding unit 240 can encode information required for video / image restoration (e.g., values of syntax elements) together with or separately from the quantized transform coefficients. The encoded information (e.g., encoded video / video information) can be transmitted or stored in the form of a bitstream in network abstraction layer (NAL) units. The video / video information can further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / video information can also include general constraint information. In this document, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device are included in video / image information. The video / image information is encoded through the encoding procedure described above and included in the bitstream. The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmitter (not shown) for transmitting and / or a storage unit (not shown) for storing the signal output from the entropy encoding unit 240 may be configured as an internal / external element of the encoding device 200, or the transmitter may be included in the entropy encoding unit 240.
[0047] The quantized transform coefficients output from the quantization unit 233 may be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) is reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transform unit 235. The adder 250 adds the reconstructed residual signal to the prediction signal output from the inter prediction unit 221 or the intra prediction unit 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). When there is no residual for the current block, such as when skip mode is applied, a predicted block may be used as the reconstructed block. The adder 250 may also be referred to as a reconstruction unit or a reconstructed block generator. The generated reconstructed signal is used for intra prediction of the next block to be processed in the current picture, and may also be used for inter prediction of the next picture after filtering, as described below.
[0048] Meanwhile, LMCS (luma mapping with chroma scaling) can be applied in the picture encoding and / or restoration process.
[0049] The filtering unit 260 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 260 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and store the modified reconstructed picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, and a bilateral filter. The filtering unit 260 generates various information related to filtering and transmits it to the entropy encoding unit 240, as will be described later in the description of each filtering method. The entropy encoding unit 240 encodes the filtering information and outputs it in the form of a bitstream.
[0050] The modified reconstructed picture sent to the memory 270 can be used as a reference picture in the inter prediction unit 221. This allows the encoding apparatus to avoid prediction mismatch between the encoding apparatus 100 and the decoding apparatus when inter prediction is applied, and also improves coding efficiency.
[0051] The DPB of the memory 270 may store a modified reconstructed picture to be used as a reference picture in the inter predictor 221. The memory 270 may store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of a block in an already reconstructed picture. The stored motion information may be transmitted to the inter predictor 221 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 270 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 222.
[0052] FIG. 3 is a diagram illustrating the configuration of a video / image decoding device to which the embodiments of this document can be applied.
[0053] As shown in FIG. 3, the decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter predictor 331 and an intra predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. Depending on the embodiment, the entropy decoding unit 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be implemented as a single hardware component (e.g., a decoder chipset or processor). The memory 360 may include a decoded picture buffer (DPB) or may be implemented as a digital storage medium. The hardware components may further include a memory 360 as an internal / external component.
[0054] When a bitstream including video / image information is input, the decoding apparatus 300 can reconstruct an image corresponding to the process in which the video / image information was processed by the encoding apparatus of FIG. 3. For example, the decoding apparatus 300 can derive units / blocks based on block division-related information obtained from the bitstream. The decoding apparatus 300 can perform decoding using a processing unit applied by the encoding apparatus. Therefore, the processing unit for decoding can be, for example, a coding unit, and the coding unit can be divided from a coding tree unit or a maximal coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the coding unit. The reconstructed image signal decoded and output by the decoding apparatus 300 can be reproduced via a reproduction device.
[0055] The decoding apparatus 300 receives a signal output from the encoding apparatus of FIG. 2 in the form of a bitstream, and the received signal is decoded by the entropy decoding unit 310. For example, the entropy decoding unit 310 may parse the bitstream to derive information (e.g., video / video information) necessary for image restoration (or picture restoration). The video / video information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / video information may also include general constraint information. The decoding apparatus may decode pictures based on the information on the parameter sets and / or the general constraint information. Signaled / received information and / or syntax elements, which will be described later in this document, may be decoded through the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 decodes information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC (context-adaptive variable length coding), or CABAC (context-adaptive arithmetic coding), and outputs values of syntax elements required for image restoration and quantized values of transform coefficients related to the residual.More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in a bitstream, determines a context model using information about the syntax element to be decoded and decoding information about neighboring and current blocks or information about symbols / bins decoded in a previous step, predicts the occurrence probability of bins according to the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values of each syntax element. After determining the context model, the CABAC entropy decoding method may update the context model using information about the decoded symbols / bins for the context model of the next symbol / bin. Prediction information from the information decoded by the entropy decoding unit 310 is provided to a prediction unit (inter prediction unit 332 and intra prediction unit 331), and residual values entropy decoded by the entropy decoding unit 310, i.e., quantized transform coefficients and related parameter information, may be input to the residual processing unit 320.
[0056] The residual processing unit 320 may derive a residual signal (residual block, residual sample, residual sample array). Information related to filtering among the information decoded by the entropy decoding unit 310 is provided to the filtering unit 350. A receiving unit (not shown) for receiving a signal output from the encoding device may be further configured as an internal / external element of the decoding device 300, and the receiving unit may be a component of the entropy decoding unit 310. The decoding device according to this document may be referred to as a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder includes the entropy decoding unit 310, and the sample decoder includes at least one of the inverse quantization unit 321, the inverse transform unit 322, the adder 340, the filtering unit 350, the memory 360, the inter prediction unit 332, and the intra prediction unit 331.
[0057] The inverse quantization unit 321 may inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit 321 may rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding apparatus. The inverse quantization unit 321 may inverse quantize the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.
[0058] The inverse transform unit 322 performs inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0059] In this document, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When the quantization / dequantization is omitted, the quantized transform coefficients may be referred to as transform coefficients. When the transform / inverse transform is omitted, the transform coefficients may be referred to as coefficients or residual coefficients, or may still be referred to as transform coefficients for uniformity of expression.
[0060] In this document, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about transform coefficients, and the information about the transform coefficients may be signaled via residual coding syntax. Transform coefficients may be derived based on the residual information (or information about the transform coefficients), and scaled transform coefficients may be derived through an inverse transform (scaling) of the transform coefficients. Residual samples may be derived based on an inverse transform (transform) of the scaled transform coefficients. This may be similarly applied / expressed in other parts of this document.
[0061] The prediction unit 330 performs prediction on the current block and generates a predicted block including prediction samples for the current block. The prediction unit 330 may determine whether intra prediction or inter prediction is applied to the current block based on information about the prediction output from the entropy decoding unit 310, and may determine a specific intra / inter prediction mode.
[0062] The prediction unit 330 may generate a prediction signal based on various prediction methods, which will be described later. For example, the prediction unit may apply intra prediction or inter prediction for predicting a block, or may simultaneously apply intra prediction and inter prediction. This may be referred to as combined inter and intra prediction (CIIP). The prediction unit may also use an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode can be used for content video / movie coding, such as games, as in screen content coding (SCC). IBC basically performs prediction within a current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, information regarding a palette table and a palette index is included in the video / picture information and signaled.
[0063] The intra prediction unit 331 can predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located far away from the current block depending on the prediction mode. Prediction modes in intra prediction include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 331 can also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.
[0064] The inter prediction unit 332 may derive a predicted block for the current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information includes a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter prediction unit 332 may construct a motion information candidate list based on the neighboring blocks and derive a motion vector and / or a reference picture index for the current block based on received candidate selection information. Inter prediction may be performed based on various prediction modes, and the prediction information may include information indicating the inter prediction mode for the current block.
[0065] The adder 340 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the acquired residual signal to a predicted signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 332 and / or the intra prediction unit 331). When there is no residual for the current block, such as when a skip mode is applied, the predicted block may be used as a reconstructed block.
[0066] The adder 340 may be referred to as a reconstruction unit or a reconstruction block generator. The generated reconstruction signal may be used for intra prediction of a next block to be processed in the current picture, may be output after filtering as described below, or may be used for inter prediction of a next picture.
[0067] Meanwhile, LMCS (luma mapping with chroma scaling) can be applied during picture decoding.
[0068] The filtering unit 350 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 350 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may transmit the modified reconstructed picture to the memory 60, specifically, to the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.
[0069] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter predictor 332. The memory 360 can store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in an already reconstructed picture. The stored motion information is transmitted to the inter predictor 332 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 360 can store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 331.
[0070] In this document, the embodiments described for the filtering unit 260, inter prediction unit 221, and intra prediction unit 222 of the encoding device 200 can also be applied identically or correspondingly to the filtering unit 350, inter prediction unit 332, and intra prediction unit 331 of the decoding device 300, respectively.
[0071] Meanwhile, the video / picture coding method according to the present document may be performed based on the following partitioning structure. Specifically, the aforementioned procedures, such as prediction, residual processing (e.g., inverse transform, inverse quantization), syntax element coding, and filtering, may be performed based on the CTUs and CUs (and / or TUs and PUs) derived based on the partitioning structure. The block partitioning procedure is performed by the image partitioning unit 210 of the encoding device, and partition-related information may be encoded by the entropy encoding unit 240 and transmitted to the decoding device in the form of a bitstream. The entropy decoding unit 310 of the decoding device may derive a block partitioning structure for the current picture based on the partitioning-related information acquired from the bitstream, and perform a series of procedures for video decoding (e.g., prediction, residual processing, block / picture reconstruction, in-loop filtering, etc.) based on the block partitioning structure. The CU size and the TU size may be the same, or multiple TUs may exist within a CU region. Meanwhile, the CU size may generally indicate the size of a luma component (sample) coding block (CB). The TU size may generally indicate the size of a luma component (sample) transform block (TB). The chroma component (sample) CB or TB size may be derived based on the luma component (sample) CB or TB size according to a component ratio according to the color format of a picture / video (chroma format, e.g., 4:4:4, 4:2:2, 4:2:0, etc.). The TU size may be derived based on maxTbSize. For example, if the CU size is larger than maxTbSize, a plurality of TUs (TBs) of the maxTbSize may be derived from the CU, and transform / inverse transform may be performed in units of the TUs (TBs). Furthermore, for example, if intra prediction is applied, the intra prediction mode / type may be derived in units of the CU (or CB), and neighboring reference sample derivation and predicted sample generation procedures may be performed in units of the TU (or TB).In this case, one or more TUs (or TBs) may exist within one CU (or CB) region, and in this case, the multiple TUs (or TBs) may share the same intra prediction mode / type.
[0072] Furthermore, in the video / image coding according to this document, video processing units may have a hierarchical structure. A picture may be divided into one or more tiles, bricks, slices, and / or tile groups. A slice may include one or more bricks. A brick may include one or more CTU rows within the tile. A slice may include an integer number of bricks in the picture. A tile group may include one or more tiles. A tile may include one or more CTUs. The CTUs may be divided into one or more CUs. A tile is a rectangular region including CTUs in a specific tile row and a specific tile column within a picture. A tile group may include an integer number of tiles according to tile raster scanning within a picture. A slice header may carry information / parameters applicable to the slice (block within the slice). If an encoding / decoding device has a multi-core processor, the encoding / decoding procedures for the tiles, slices, bricks, and / or tile groups may be processed in parallel. In this document, the terms slice and tile group may be used interchangeably. That is, the tile group header may be referred to as a slice header. Here, a slice may have one of slice types including an intra (I) slice, a predictive (P) slice, and a bi-predictive (B) slice. For blocks in an I slice, inter prediction is not used for prediction, and only intra prediction can be used. Of course, even in this case, original sample values can be coded and signaled without prediction. For blocks in a P slice, intra prediction or inter prediction can be used, and if inter prediction is used, only uni prediction can be used. On the other hand, for blocks in a B slice, intra prediction or inter prediction can be used, and if inter prediction is used, up to bi prediction can be used.
[0073] In an encoding device, the tile / tile group, brick, slice, maximum and minimum coding unit sizes can be determined based on the characteristics of the video image (e.g., resolution) or taking into consideration coding efficiency or parallel processing, and information regarding this or information that can guide this can be included in the bitstream.
[0074] The decoding device can acquire information indicating whether a tile / tile group, brick, slice, or CTU within a tile of the current picture is divided into multiple coding units, etc. If such information is acquired (transmitted) only under specific conditions, efficiency can be improved.
[0075] Meanwhile, as described above, one picture may include multiple slices, and one slice may include a slice header and slice data. In this case, one picture header may be added to multiple slices (slice header and slice data set) in one picture. The picture header (picture header syntax) may include information / parameters commonly applicable to the picture. The slice header (slice header syntax) may include information / parameters commonly applicable to the slices. An adaptation parameter set (APS) or a picture parameter set (PPS) may include information / parameters commonly applicable to one or more pictures. A sequence parameter set (SPS) may include information / parameters commonly applicable to one or more sequences. A video parameter set (VPS) may include information / parameters commonly applicable to multiple layers. A decoding parameter set (DPS) may include information / parameters commonly applicable to all videos. The DPS may include information / parameters related to concatenation of a coded video sequence (CVS).
[0076] In this document, the higher level syntax may include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, picture header syntax, and slice header syntax.
[0077] Also, for example, information regarding the division and configuration of tiles / tile groups / bricks / slice can be configured in the encoding device based on the higher level syntax and transmitted to the decoding device in the form of a bitstream.
[0078] Meanwhile, a video / image encoding procedure based on inter prediction may generally include, for example:
[0079] FIG. 4 shows an example of a video / image encoding method based on inter prediction.
[0080] Referring to FIG. 4, an encoding apparatus performs inter prediction on a current block (S400). The encoding apparatus may derive an inter prediction mode and motion information for the current block and generate a predicted sample for the current block. Here, the inter prediction mode determination, motion information derivation, and predicted sample generation procedures may be performed simultaneously, or one procedure may be performed before the other procedures. For example, an inter prediction unit of the encoding apparatus may include a prediction mode determination unit, a motion information derivation unit, and a predicted sample derivation unit, in which the prediction mode determination unit may determine a prediction mode for the current block, the motion information derivation unit may derive motion information for the current block, and the predicted sample derivation unit may derive a predicted sample for the current block. For example, the inter prediction unit of the encoding apparatus may search for a block similar to the current block within a predetermined region (search region) of a reference picture through motion estimation and derive a reference block whose difference from the current block is minimal or equal to or less than a predetermined standard. Based on this, a reference picture index indicating a reference picture in which the reference block is located can be derived, and a motion vector can be derived based on a position difference between the reference block and the current block. The encoding apparatus can determine a mode to be applied to the current block from various prediction modes. The encoding apparatus can compare rate-distortion (RD) costs for the various prediction modes to determine an optimal prediction mode for the current block.
[0081] For example, when a skip mode or a merge mode is applied to the current block, the encoding apparatus may construct a merge candidate list and derive a reference block, among reference blocks indicated by merge candidates included in the merge candidate list, whose difference between the current block and the current block is minimum or equal to or less than a predetermined criterion. In this case, a merge candidate associated with the derived reference block may be selected, and merge index information indicating the selected merge candidate may be generated and signaled to a decoding apparatus. Motion information of the current block may be derived using motion information of the selected merge candidate.
[0082] As another example, when the (A)MVP mode is applied to the current block, the encoding apparatus may construct an (A)MVP candidate list and use a motion vector of an MVP (motion vector predictor) candidate selected from the MVP candidates included in the (A)MVP candidate list as the MVP of the current block. In this case, for example, a motion vector pointing to a reference block derived by the motion estimation may be used as the motion vector of the current block, and the MVP candidate having the smallest difference from the motion vector of the current block among the MVP candidates may be the selected MVP candidate. A motion vector difference (MVD), which is the difference obtained by subtracting the MVP from the motion vector of the current block, may be derived. In this case, information regarding the MVD may be signaled to the decoding apparatus. Also, when the (A)MVP mode is applied, the value of the reference picture index may be configured as reference picture index information and separately signaled to the decoding apparatus.
[0083] The encoding apparatus may derive residual samples based on the predicted samples (S410) by comparing the original samples of the current block with the predicted samples.
[0084] The encoding apparatus encodes video information including prediction information and residual information (S420). The encoding apparatus may output the encoded video information in the form of a bitstream. The prediction information may include prediction mode information (e.g., skip flag, merge flag, or mode index) and information about motion information, which are information related to the prediction procedure. The information about the motion information may include candidate selection information (e.g., merge index, MVP flag, or MVP index), which is information for deriving a motion vector. The information about the motion information may also include information about the MVD and / or reference picture index information. The information about the motion information may also include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. The residual information is information about the residual sample. The residual information may also include information about quantized transform coefficients for the residual sample.
[0085] The output bitstream can be stored in a (digital) storage medium and then transmitted to the decoding device, or can be transmitted to the decoding device via a network.
[0086] Meanwhile, as described above, the encoding apparatus can generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on the reference samples and the residual samples. This is because the encoding apparatus derives the same prediction result as that performed by the decoding apparatus, thereby improving coding efficiency. Therefore, the encoding apparatus can store the reconstructed picture (or reconstructed samples, reconstructed blocks) in a memory and use it as a reference picture for inter prediction. As described above, an in-loop filtering procedure can be further applied to the reconstructed picture.
[0087] A video / picture decoding procedure based on inter prediction may generally include, for example:
[0088] FIG. 5 illustrates an example of a video / picture decoding method based on inter prediction.
[0089] The decoding device may perform operations corresponding to those performed by the encoding device, and may perform prediction on the current block based on the received prediction information to derive predicted samples.
[0090] 5, a decoding apparatus may determine a prediction mode for the current block based on prediction information received from a bitstream (S500). The decoding apparatus may determine which inter-prediction mode is applied to the current block based on prediction mode information in the prediction information.
[0091] For example, it may determine whether a merge mode or an (A)MVP mode is applied to the current block based on a merge flag, or may select one of various inter prediction mode candidates based on the merge index. The inter prediction mode candidates may include various inter prediction modes such as skip mode, merge mode, and / or (A)MVP mode.
[0092] The decoding apparatus derives motion information of the current block based on the determined inter prediction mode (S510). For example, when a skip mode or a merge mode is applied to the current block, the decoding apparatus may construct a merge candidate list (described below) and select one merge candidate from among a plurality of merge candidates included in the merge candidate list. The selection may be performed based on the selection information (merge index) described above. Motion information of the selected merge candidate may be used to derive motion information of the current block. The motion information of the selected merge candidate may be used as motion information of the current block.
[0093] As another example, when the (A)MVP mode is applied to the current block, the decoding apparatus may construct an (A)MVP candidate list and use a motion vector predictor (MVP) selected from among the MVP candidates included in the (A)MVP candidate list as the MVP of the current block. The selection may be performed based on the selection information (MVP flag or MVP index). In this case, the MVD of the current block may be derived based on information related to the MVD, and the motion vector of the current block may be derived based on the MVP of the current block and the MVD. Also, the decoding apparatus may derive a reference picture index of the current block based on the reference picture index information. A picture pointed to by the reference picture index in the reference picture list for the current block may be derived as a reference picture referenced for inter-prediction of the current block.
[0094] Alternatively, the motion information of the current block may be derived without constructing a candidate list, in which case the candidate list construction described above may be omitted.
[0095] The decoding apparatus may generate predictive samples for the current block based on the motion information of the current block (S520). In this case, the reference picture may be derived based on a reference picture index of the current block, and the predictive samples of the current block may be derived using samples of a reference block to which the motion vector of the current block points on the reference picture. In this case, as described below, a predictive sample filtering procedure may be further performed on all or some of the predictive samples of the current block, depending on the circumstances.
[0096] For example, the inter-prediction unit of the decoding device may include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit, and may determine a prediction mode for the current block based on prediction mode information received by the prediction mode determination unit, derive motion information (such as a motion vector and / or a reference picture index) of the current block based on information regarding the motion information received by the motion information derivation unit, and derive a prediction sample of the current block by the prediction sample derivation unit.
[0097] The decoding apparatus generates residual samples for the current block based on the received residual information (S530). The decoding apparatus generates reconstructed samples for the current block based on the predicted samples and the residual samples, and can generate a reconstructed picture based on the reconstructed samples (S540). Thereafter, as described above, an in-loop filtering procedure can be further applied to the reconstructed picture.
[0098] Meanwhile, a predicted block for the current block may be derived based on motion information derived according to the prediction mode of the current block. The predicted block may include predicted samples (prediction sample array) of the current block. If the motion vector of the current block points to a fractional sample unit, an interpolation procedure may be performed, through which predicted samples of the current block may be derived based on reference samples in fractional sample units within a reference picture. If affine inter-prediction is applied to the current block, predicted samples may be generated based on a sample / sub-block unit MV (motion vector). If bi-prediction is applied, predicted samples derived through a weighted sum or weighted average (according to phase) of predicted samples derived based on L0 prediction (i.e., prediction using a reference picture in a reference picture list L0 and MVL0) and predicted samples derived based on L1 prediction (i.e., prediction using a reference picture in a reference picture list L1 and MVL1) may be used as predicted samples of the current block. When bi-prediction is applied, if the reference picture used for L0 prediction and the reference picture used for L1 prediction are located in different temporal directions relative to the current picture (i.e., if bi-prediction is possible but supports bidirectional prediction), this can be called true bi-prediction.
[0099] As described above, reconstructed samples and reconstructed pictures can be generated based on the derived predicted samples, and then procedures such as in-loop filtering can be performed.
[0100] Meanwhile, weighted sample prediction may be used in inter prediction. Weighted sample prediction may be referred to as weighted prediction. Weighted prediction may be applied when the slice type of a current slice in which a current block (e.g., CU) is located is a P slice or a B slice. That is, weighted prediction may be used not only when bi-prediction is applied but also when uni-prediction is applied. For example, as described below, weighted prediction may be determined based on weightedPredFlag, and the value of weightedPredFlag may be determined based on signaled pps_weighted_pred_flag (for a P slice) or pps_weighted_bipred_flag (for a B slice). For example, if slice_type is P, weightedPredFlag may be set as pps_weighted_pred_flag. Otherwise (if slice_type is B), weightedPredFlag may be set as pps_weighted_bipred_flag.
[0101] The predicted samples or values of the predicted samples that are the output of the weighted prediction may be referred to as pbSamples.
[0102] Weighted prediction procedures can be broadly divided into default weighted (sample) prediction procedures and explicit weighted (sample) prediction procedures. The weighted (sample) prediction procedure may also refer to only the explicit weighted (sample) prediction procedure. For example, when the value of weightedPredFlag is 0, the predicted sample values (pbSamples) can be derived based on the default weighted (sample) prediction procedure. When the value of weightedPredFlag is 1, the predicted sample values (pbSamples) can be derived based on the explicit weighted (sample) prediction procedure.
[0103] On the other hand, when bi-prediction is applied to the current block, predicted samples can be derived based on a weighted average. Conventionally, a bi-predictive signal (i.e., a bi-predictive sample) could be derived through a simple average of an L0 predicted signal (L0 predicted sample) and an L1 predicted signal (L1 predicted sample). That is, a bi-predictive sample was derived by averaging an L0 predicted sample based on an L0 reference picture and MVL0 and an L1 predicted sample based on an L1 reference picture and MVL1. However, according to this document, when bi-prediction is applied, a bi-predictive signal (bi-predictive sample) can be derived through a weighted average of an L0 predicted signal and an L1 predicted signal.
[0104] Bi-directional optical flow (BDOF) can be used to refine a bi-prediction signal. BDOF is used to calculate improved motion information and generate predicted samples when bi-prediction is applied to a current block (e.g., CU). The process of calculating the improved motion information can be included in the motion information derivation step.
[0105] For example, BDOF can be applied at a 4x4 sub-block level, i.e., BDOF can be performed in units of 4x4 sub-blocks within a current block. BDOF can be applied only to the luma component. Alternatively, BDOF can be applied only to the chroma component, or to the luma and chroma components.
[0106] Meanwhile, as mentioned above, a high level syntax (HLS) can be coded / signaled for video / image coding, and video / image information can be included in the HLS.
[0107] A coded picture can consist of one or more slices. Parameters describing a coded picture are signaled in a picture header, and parameters describing a slice are signaled in a slice header. The picture header itself is carried in the form of an NAL unit. The slice header is located at the beginning of the NAL unit that contains the payload of the slice (i.e., slice data).
[0108] Each picture is associated with a picture header. A picture can be composed of different types of slices (intra-coded slices (i.e., I-slices) and inter-coded slices (i.e., P-slices and B-slices)). Thus, the picture header can include syntax elements necessary for intra-slices and inter-slices of a picture.
[0109] A picture can be divided into sub-pictures, tiles, and / or slices. Sub-picture signaling can be present in the sequence parameter set (SPS), and tile and square slice signaling can be present in the picture parameter set (PPS). Raster-scan slice signaling can be present in the slice header.
[0110] When weighted prediction is applied for inter prediction of the current block, the weighted prediction may be performed based on information related to weighted prediction.
[0111] The weighted prediction procedure can be initiated based on two flags in the SPS.
[0112] As an example, for weighted prediction, the SPS syntax may include syntax elements such as those shown in Table 1 below.
[0113] [Table 1]
[0114] In Table 1, if the value of sps_weighted_pred_flag is 1, this may indicate that weighted prediction is applied to the P slice that references the SPS.
[0115] When the value of sps_weighted_bipred_flag is 1, this may indicate that weighted prediction is applied to a B slice that references the SPS. When the value of sps_weighted_bipred_flag is 0, this may indicate that weighted prediction is not applied to a B slice that references the SPS.
[0116] The two flags signaled in the SPS indicate whether weighted prediction is applied to P and B slices in a coded video sequence (CVS).
[0117] On the other hand, for weighted prediction, the PPS syntax can include syntax elements such as those shown in Table 2 below.
[0118] [Table 2]
[0119] In Table 2, if the value of pps_weighted_pred_flag is 0, this may indicate that weighted prediction is not applied to the P slice referring to the PPS. If the value of pps_weighted_pred_flag is 1, this may indicate that weighted prediction is applied to the P slice referring to the PPS. If the value of pps_weighted_pred_flag is 0, the value of pps_weighted_pred_flag is 0.
[0120] If the value of pps_weighted_bipred_flag is 0, this may indicate that weighted prediction is not applied to the B slice that references the PPS. If the value of pps_weighted_bipred_flag is 1, this may indicate that explicit weighted prediction is applied to the B slice that references the PPS. If the value of sps_weighted_bipred_flag is 0, the value of pps_weighted_bipred_flag is 0.
[0121] Additionally, the slice header syntax may include syntax elements such as those shown in Table 3 below.
[0122] [Table 3-1]
[0123] [Table 3-2]
[0124] In Table 3, slice_pic_parameter_set_id indicates the value of pps_pic_parameter_set_id for the PPS in use. The value of slice_pic_parameter_set_id is in the range of 0 to 63.
[0125] The value of the temporary ID (TempralID) of the current picture must be greater than or equal to the tempralID value of the PPS with pps_pic_parameter_set_id such as slice_pic_parameter_set_id.
[0126] Meanwhile, the prediction weight table syntax can include information about weighted prediction as shown in Table 4 below.
[0127] [Table 4-1]
[0128] [Table 4-2]
[0129] In Table 4, luma_log2_weight_denom is the base 2 logarithm of the denominator for all luma weighting factors. The value of luma_log2_weight_denom is in the range 0 to 7.
[0130] delta_chroma_log2_weight_denom is the base 2 log difference in the denominator for all chroma weight values. If delta_chroma_log2_weight_denom is not present, it is inferred to 0.
[0131] ChromaLog2WeightDenom is derived as luma_log2_weight_denom+delta_chroma_log2_weight_denom, and its value is in the range 0 to 7.
[0132] A value of 1 for luma_weight_l0_flag[i] indicates the presence of a weighting factor for the luma component of (reference picture) list 0 (L0) prediction using RefPicList[0][i]. A value of 0 for luma_weight_l0_flag[i] indicates the absence of such a weighting factor.
[0133] When the value of chroma_weight_l0_flag[i] is 1, this indicates that there is a weighting coefficient for the chroma prediction value of L0 prediction using RefPicList[0][i]. When the value of chroma_weight_l0_flag[i] is 0, this indicates that such a weighting coefficient does not exist. When chroma_weight_l0_flag[i] does not exist, it is inferred to be 0.
[0134] delta_luma_weight_l0[i] is the difference of the weighting coefficients applied to the luma prediction value for L0 prediction using RefPicList[0][i].
[0135] LumaWeightL0[i] is derived to be (1<<luma_log2_weight_denom)+delta_luma_weight_l0[i]. When luma_weight_l0_flag[i] is 1, the value of delta_luma_weight_l0[i] is included in the range from -128 to 127. When luma_weight_l0_flag[i] is 0, LumaWeightL0[i] is inferred to be 2 luma_log2_weight_denom is inferred.
[0136] luma_offset_l0[i] is the additive offset applied to the luma prediction value for L0 prediction using RefPicList[0][i]. The value of luma_offset_l0[i] is included in the range from -128 to 127. When the value of luma_weight_l0_flag[i] is 0, the value of luma_offset_l0[i] is inferred to be 0.
[0137] delta_chroma_weight_l0[i][j] is the difference of the weighting coefficients applied to the chroma prediction value for L0 prediction using RefPicList[0][i] where j is 0 for Cb and j is 1 for Cr.
[0138] ChromaWeightL0[i][j] is derived to (1<<ChromaLog2WeightDenom)+delta_chroma_weight_l0[i][j]. When chroma_weight_l0_flag[i] is 1, the value of delta_chroma_weight_l0[i][j] is included in the range from -128 to 127. When chroma_weight_l0_flag[i] is 0, ChromaWeightL0[i][j] is inferred to be 2 ChromaLog2WeightDenom as inferred.
[0139] delta_chroma_offset_l0[i][j] is the difference of the addition offset applied to the chroma prediction value for L0 prediction, using RefPicList[0][i] where j is 0 for Cb and j is 1 for Cr.
[0140] The value of delta_chroma_offset_l0[i][j] is included in the range from -4×128 to 4×127. When the value of chroma_weight_l0_flag[i] is 0, the value of ChromaOffsetL0[i][j] is inferred to be 0.
[0141] The above prediction weight table syntax is often used to modify the sequence when there is a scene change. The existing prediction weight table syntax is signaled in the slice header when the PPS flag for weighted prediction is enabled and the slice type is P, or when the PPS flag for weighted bi-prediction is enabled and the slice type is B. However, when the scene changes, it often occurs that one or several frames need to adjust the prediction weight table. Usually, when the PPS is shared among multiple frames, it may be unnecessary to signal the information regarding weighted prediction for all frames referring to the PPS.
[0142] The following drawings are created to explain a specific example of the present document. The names of specific devices and names of specific signals / information shown in the drawings are presented for illustrative purposes only, and the technical features of the present specification are not limited to the specific names used in the following drawings.
[0143] This document provides the following methods to solve the above problems, each of which can be applied independently or in combination with each other.
[0144] 1. Tools for weighted prediction (information about weighted prediction) can be applied at the picture level rather than the slice level: weighting values are applied to a specific reference picture of a picture and are used for all slices of that picture.
[0145] 2. The prediction weight table syntax can be signaled at the picture level rather than the slice level. For this purpose, the prediction weight table syntax can be signaled in the picture header (PH) or picture parameter set (PPS).
[0146] 3. When weighted prediction is applied to a picture, all slices within the picture may have the same active reference pictures, which includes the order of the active reference pictures in the reference picture list (RPL) (i.e., L0 for P slices, L0 and L1 for B slices).
[0147] 4. Alternatively, if the above does not apply, the following may apply:
[0148] a. The signaling of weighted prediction is independent of the reference picture list signaling, i.e., there is no assumption in the prediction weight table signaling about the order of reference pictures in the reference picture list.
[0149] b. There is no signaling of weighted prediction values for reference pictures in L0 and L1. Weight values are provided directly for reference pictures.
[0150] c. Instead of two loops for signaling weight values for reference pictures, only one loop can be used, where in each loop, the reference picture associated with the weight value to be signaled is first identified.
[0151] d. Reference picture identification is based on the picture order count (POC) value.
[0152] e. For bit saving, instead of signaling the POC value of the reference picture, the delta POC value between the reference picture and the current picture can be signaled.
[0153] 5. In addition to item 4 above, the following can be applied for signaling the delta POC value between the reference picture and the current picture, so that the absolute delta POC value can be signaled as follows:
[0154] a. For the first signaled delta POC, this is the delta between the POC of the reference picture and the current picture.
[0155] b. For the signaled delta POC rest (ie, when i starts from 1), this is the delta between the POC of the ith reference picture and the (i-1)th reference picture.
[0156] 6. The two flags in the PPS can be combined into a single control flag (e.g., pps_weighted_pred_flag), which can be used to indicate the presence of additional flags in the picture header.
[0157] a. The flag in the picture header may be changed depending on the PPS flag (The flag in the PH may be conditioned on the PPS flag), and if the NAL unit type is not IDR (instantaneous decoding refresh), may further indicate the presence of pred_weighted_table() data (prediction weight table syntax).
[0158] 7. The two flags signaled in PPS (pps_weighted_pred_flag and pps_weighted_bipred_flag) can be combined into one flag, which can use the existing name pps_weighted_pred_flag.
[0159] 8. A flag can be signaled in a picture header to indicate whether weighted prediction is applied to the picture associated with the picture header. The flag can be referred to as pic_weighted_pred_flag.
[0160] The presence of a.pic_weighted_pred_flag may depend on the value of pps_weighted_pred_flag. If the value of pps_weighted_pred_flag is 0, pic_weighted_pred_flag is not present and its value may be inferred to be 0.
[0161] b. If the value of pic_weighted_pred_flag is 1, signaling of pred_weighted_table() can be present in the picture header.
[0162] 9. Alternatively, if weighted prediction is enabled (i.e., pps_weighted_pred_flag has a value of 1 or pps_weighted_bipred_flag has a value of 1), information regarding weighted prediction may still be present in the slice header and the following may apply.
[0163] a. A new flag can be signaled to indicate whether information about weighted prediction is present in the slice header, which can be called slice_weighted_pred_present_flag.
[0164] The presence of b.slice_weighted_pred_present_flag can be determined by the type of slice, the values of pps_weighted_pred_flag and pps_weighted_bipred_flag.
[0165] In this document, information about weighted prediction may include information / syntax elements about weighted prediction described in Tables 1 to 4. Video / image information may include various information for inter prediction, such as information about weighted prediction, residual information, and inter prediction mode information. The inter prediction mode information may include information / syntax elements, such as information indicating whether merge mode or MVP mode can be applied to the current block, and selection information for selecting one of the motion candidates in a motion candidate list. For example, when merge mode is applied to the current block, a merge candidate list may be constructed based on neighboring blocks of the current block, and one candidate may be selected / used (based on a merge index) from the merge candidate list to derive motion information of the current block. As another example, when MVP mode is applied to the current block, an MVP candidate list may be constructed based on neighboring blocks of the current block, and one candidate may be selected / used (based on an MVP flag) from the MVP candidate list to derive motion information of the current block.
[0166] In one embodiment, for weighted prediction during inter prediction, the PPS may include syntax elements such as those shown in Table 5 below, and the semantics thereof are as shown in Table 6 below.
[0167] [Table 5]
[0168] [Table 6]
[0169] Referring to Tables 5 and 6, if the value of pps_weighted_pred_flag is 0, this may indicate that weighted prediction is not applied to the P or B slice referring to the PPS, and if the value of pps_weighted_pred_flag is 1, this may indicate that weighted prediction is applied to the P or B slice referring to the PPS.
[0170] The picture header can also include syntax elements as shown in Table 7 below, with semantics as shown in Table 8 below.
[0171] [Table 7]
[0172] [Table 8]
[0173] Referring to Tables 7 and 8, if the value of pic_weighted_pred_flag is 0, this may indicate that weighted prediction is not applied to the P or B slice that references the picture header. If the value of pic_weighted_pred_flag is 1, this may indicate that weighted prediction is applied to the P or B slice that references the picture header.
[0174] If the value of pic_weighted_pred_flag is 1, all slices in the picture associated with the picture header may have the same reference picture list. Otherwise, if the value of pic_weighted_pred_flag is 1, the value of pic_rpl_present_flag may be 1.
[0175] In the absence of the above conditions, pic_weighted_pred_flag can be signaled as shown in Table 9 below.
[0176] [Table 9]
[0177] On the other hand, the slice header can include syntax elements such as those shown in Table 10 below.
[0178] [Table 10]
[0179] In addition, the prediction weight table syntax may include syntax elements such as those shown in Table 11 below, and the semantics for these elements may be as shown in Table 12 below.
[0180] [Table 11-1]
[0181] [Table 11-2]
[0182] [Table 12]
[0183] Referring to Table 11 and Table 12, mum_l0_weighted_ref_pics may indicate the number of weighted reference pictures in reference picture list 0. The value of num_l0_weighted_ref_pics is in the range from 0 to MaxDecPicBuffMinus1+14.
[0184] mum_l1_weighted_ref_pics may indicate the number of weighted reference pictures in reference picture list 1. The value of num_l1_weighted_ref_pics is in the range from 0 to MaxDecPicBuffMinus1+14.
[0185] If the value of luma_weight_l0_flag[i] is 1, it indicates that there is a weighting factor for the luma component of list 0 (L0) prediction using RefPicList[0][i].
[0186] A value of 1 for chroma_weight_l0_flag[i] indicates the presence of a weighting factor for the chroma prediction value for L0 prediction using RefPicList[0][i]. A value of 0 for chroma_weight_l0_flag[i] indicates the absence of such a weighting factor.
[0187] If the value of luma_weight_l1_flag[i] is 1, it indicates that there is a weighting factor for the luma component of list 1 (L1) prediction using RefPicList[0][i].
[0188] chroma_weight_l1_flag[i] indicates the presence of a weighting factor for the chroma prediction value for L1 prediction using RefPicList[0][i]. If chroma_weight_l0_flag[i] has a value of 0, it indicates that no such weighting factor exists.
[0189] For example, when weighted prediction is applied to a current block, the encoding apparatus may generate count information for weighted reference pictures in a reference picture list of the current block based on the weighted prediction. The count information may refer to count information for weight values signaled for items (reference pictures) in an L0 reference picture list and / or an L1 reference picture list. That is, the value of the count information may be the same as the number of weighted reference pictures in the corresponding reference picture list (L0 and / or L1). Therefore, if the value of the count information is n, the prediction weight table syntax may include n weighting factor-related flags for the reference picture lists. The weighting factor-related flags may correspond to luma_weight_l0_flag, luma_weight_l1_flag, chroma_weight_l0_flag, and / or chroma_weight_l0_flag in Table 11. A weighting value for the current picture may be derived based on the weighting factor-related flags.
[0190] When weighted bi-prediction is applied to the current block, the prediction weight table syntax may include number information for weighted reference pictures in an L1 reference picture list and number information for weighted reference pictures in the L0 reference picture list independently, as shown in Table 11. The weighting coefficient-related flag may be included independently for each of the number information for weighted reference pictures in the L1 reference picture list and number information for weighted reference pictures in the L0 reference picture list. That is, the prediction weight table syntax may include the same number of luma_weight_l0_flag and / or chroma_weight_l0_flag as the number of weighted reference pictures in the L0 reference picture list, and may include the same number of luma_weight_l1_flag and / or chroma_weight_l1_flag as the number of weighted reference pictures in the L1 reference picture list.
[0191] The encoding apparatus may encode video information including the number information, weighting coefficient-related flags, etc., and output the encoded video information in the form of a bitstream. Here, the number information and weighting coefficient-related flags may be included in a prediction weight table syntax in the video information as shown in Table 11. The prediction weight table syntax may be included in a picture header or a slice header in the video information. A weighted prediction-related flag may be included in a picture parameter set and / or a picture header to indicate whether the prediction weight table syntax is included in the picture header, i.e., to indicate whether information related to weighted prediction is present in the picture header. If the weighted prediction-related flag is included in the picture parameter set, it may correspond to pps_weighted_pred_flag in Table 5. If the weighted prediction-related flag is included in the picture header, it may correspond to pic_weighted_pred_flag in Table 7. Alternatively, both pps_weighted_pred_flag and pic_weighted_pred_flag may be included in the video information to indicate whether the prediction weight table syntax is included in the picture header.
[0192] When a flag related to weighted prediction is parsed from a bitstream, the decoding apparatus can parse a prediction weight table syntax from the bitstream based on the parsed flag. The flag related to weighted prediction can be parsed from a picture parameter set and / or the picture header of the bitstream. In other words, the flag related to weighted prediction can include pps_weighted_pred_flag and / or pic_weighted_pred_flag. If the value of pps_weighted_pred_flag and / or pic_weighted_pred_flag is 1, the decoding apparatus can parse the prediction weight table syntax from the picture header of the bitstream.
[0193] When the prediction weight table syntax is parsed from the picture header (when the value of pps_weighted_pred_flag and / or pic_weighted_pred_flag is 1), the decoding device may apply information about weighted prediction included in the prediction weight table syntax to all slices in the current picture. In other words, when the prediction weight table syntax is parsed from the picture header, all slices in the picture associated with the picture header may have the same reference picture list.
[0194] The decoding apparatus may parse number information for weighted reference pictures in a reference picture list of the current block based on the prediction weight table syntax. The value of the number information may be the same as the number of weighted reference pictures in the reference picture list. When weighted bi-prediction is applied to the current block, the decoding apparatus may parse the number information for weighted reference pictures in an L1 reference picture list and the number information for weighted reference pictures in an L0 reference picture list independently from the prediction weight table syntax.
[0195] The decoding apparatus may parse weighting factor-related flags for the reference picture list from the prediction weight table syntax based on the number information. The weighting factor-related flags may correspond to luma_weight_l0_flag, luma_weight_l1_flag, chroma_weight_l0_flag, and / or chroma_weight_l0_flag in Table 11. For example, if the value of the number information is n, the decoding apparatus may parse n weighting factor-related flags from the prediction weight table syntax. Then, the decoding apparatus may derive weighting values for the reference pictures of the current block based on the weighting factor-related flags, and perform weighted prediction on the current block based on the weighting factor-related flags to generate or derive predicted samples. Thereafter, the decoding apparatus may generate or derive reconstructed samples for the current block based on the predicted samples, and reconstruct the current picture based on the reconstructed samples.
[0196] As another embodiment, for weighted prediction during inter prediction, the picture header may include syntax elements such as those shown in Table 13 below, and the semantics thereof may be as shown in Table 14 below.
[0197] [Table 13]
[0198] [Table 14]
[0199] Referring to Tables 13 and 14, if the value of pic_weighted_pred_flag is 0, this may indicate that weighted prediction is not applied to a P or B slice that references the picture header. If the value of pic_weighted_pred_flag is 1, this may indicate that weighted prediction is applied to a P or B slice that references the picture header. If the value of sps_weighted_pred_flag is 0, the value of pic_weighted_pred_flag is 0.
[0200] On the other hand, the slice header can include syntax elements such as those shown in Table 15 below.
[0201] [Table 15-1]
[0202] [Table 15-2]
[0203] Referring to Table 15, a flag for weighted prediction (pic_weighted_pred_flag) may indicate whether the prediction weight table syntax (information about weighted prediction) is present in the picture header or the slice header. When the value of pic_weighted_pred_flag is 1, it may indicate that the prediction weight table syntax (information about weighted prediction) is not present in the slice header but may be present in the picture header. When the value of pic_weighted_pred_flag is 0, it may indicate that the prediction weight table syntax (information about weighted prediction) is not present in the picture header but may be present in the slice header. Although Tables 13 and 14 show that the flag for weighted prediction is signaled in the picture header, the flag for weighted prediction may also be signaled in the picture parameter set.
[0204] For example, when weighted prediction is applied to a current block, the encoding apparatus may perform the weighted prediction and, based on the weighted prediction, encode video information including a flag and prediction weight table syntax related to the weighted prediction. In this case, when the prediction weight table syntax is included in a picture header of the video information, the encoding apparatus may set the flag to 1, and when the prediction weight table syntax is included in a slice header of the video information, the encoding apparatus may set the flag to 0. When the flag is set to 1, information related to the weighted prediction included in the prediction weight table syntax may be applied to all slices in the current picture. When the flag is set to 0, information related to the weighted prediction included in the prediction weight table syntax may be applied to a slice associated with a slice header among slices in the current picture. Therefore, when the prediction weight table syntax is included in the picture header, all slices in the picture associated with the picture header may have the same reference picture list, and when the prediction weight table syntax is included in the slice header, slices associated with the slice header may have the same reference picture list.
[0205] Meanwhile, the prediction weight table syntax may include number information for weighted reference pictures in a reference picture list of the current block, weighting factor-related flags, etc. As described above, the number information may mean number information for weighting values signaled for items (reference pictures) in the L0 reference picture list and / or the L1 reference picture list, and the value of the number information may be the same as the number of weighted reference pictures in the corresponding reference picture list (L0 and / or L1). Therefore, if the value of the number information is n, the prediction weight table syntax may include n weighting factor-related flags for reference picture lists. The weighting factor-related flags may correspond to luma_weight_l0_flag, luma_weight_l1_flag, chroma_weight_l0_flag, and / or chroma_weight_l0_flag in Table 11.
[0206] When weighted bi-prediction is applied to the current block, the encoding apparatus may generate a prediction weight table syntax including number information for weighted reference pictures in an L1 reference picture list and number information for weighted reference pictures in the L0 reference picture list. The prediction weight table syntax may include the weighting factor-related flags independently for the number information for weighted reference pictures in the L1 reference picture list and the number information for weighted reference pictures in the L0 reference picture list. That is, the prediction weight table syntax may include the same number of luma_weight_l0_flag and / or chroma_weight_l0_flag as the number of weighted reference pictures in the L0 reference picture list, and may include the same number of luma_weight_l1_flag and / or chroma_weight_l1_flag as the number of weighted reference pictures in the L1 reference picture list.
[0207] When a flag related to weighted prediction is parsed from a bitstream, the decoding apparatus may parse a prediction weight table syntax from the bitstream based on the parsed flag. The flag related to weighted prediction may be parsed from a picture parameter set and / or the picture header of the bitstream. In other words, the flag related to weighted prediction may correspond to pps_weighted_pred_flag and / or pic_weighted_pred_flag. When the flag related to weighted prediction has a value of 1, the decoding apparatus may parse the prediction weight table syntax from a picture header of the bitstream. When the flag related to weighted prediction has a value of 0, the decoding apparatus may parse the prediction weight table syntax from a slice header of the bitstream.
[0208] When a prediction weight table syntax is parsed from a picture header, the decoding apparatus can apply information about weighted prediction included in the prediction weight table syntax to all slices in a current picture. In other words, when a prediction weight table syntax is parsed from a picture header, all slices in a picture associated with the picture header can have the same reference picture list. When the prediction weight table syntax is parsed from a slice header, the decoding apparatus can apply information about weighted prediction included in the prediction weight table syntax to slices in a current picture associated with the slice header. In other words, when a prediction weight table syntax is parsed from a picture header, slices associated with the slice header can have the same reference picture list.
[0209] The decoding apparatus may parse number information for weighted reference pictures in a reference picture list of the current block based on the prediction weight table syntax. The value of the number information may be the same as the number of weighted reference pictures in the reference picture list. When weighted bi-prediction is applied to the current block, the decoding apparatus may parse the number information for weighted reference pictures in an L1 reference picture list and the number information for weighted reference pictures in an L0 reference picture list independently from the prediction weight table syntax.
[0210] The decoding apparatus may parse weighting factor-related flags for the reference picture list from the prediction weight table syntax based on the number information. The weighting factor-related flags may correspond to the luma_weight_l0_flag, luma_weight_l1_flag, chroma_weight_l0_flag, and / or chroma_weight_l0_flag. For example, if the value of the number information is n, the decoding apparatus may parse n weighting factor-related flags from the prediction weight table syntax. The decoding apparatus may then derive weight values for the reference pictures of the current block based on the weighting factor-related flags, and perform inter prediction on the current block based on the weighting factor-related flags to generate or derive predicted samples. The decoding apparatus may then generate or derive reconstructed samples for the current block based on the predicted samples, and generate a reconstructed picture for the current picture based on the reconstructed samples.
[0211] In another embodiment, the prediction weight table syntax may include syntax elements such as those shown in Table 16 below, and the semantics thereof may be as shown in Table 17 below.
[0212] [Table 16]
[0213]
Table 17
[0214] In Tables 16 and 17, if pic_poc_delta_sign[i] does not exist, it is inferred to be 0. DeltaPocWeightedRefPic[i] for i within the range from 0 to num_weighted_ref_pics_minus1 can be derived as follows.
[0215]
Number
[0216] Also, ChromaWeight[i][j] can be derived as (1<<ChromaLog2WeightDenom)+delta_chroma_weight[i][j]. When the value of chroma_weight_flag[i] is 1, the value of delta_chroma_weight[i][j] is within the range from -128 to 127. When the value of chroma_weight_flag[i] is 0, ChromaWeight[i][j] can be derived as 2ChromaLog2WeightDenom.
[0217] Also, ChromaOffset[i][j] can be derived as follows.
[0218]
Number
[0219] The value of delta_chroma_offset[i][j] is within the range from -4*128 to 4*127. When the value of chroma_weight_flag[i] is 0, the value of ChromaOffset[i][j] is inferred to be 9.
[0220] sumWeightFlags may be derived as the sum of luma_weight_flag[i]+2*chroma_weight_flag[i], where i is in the range of 0 to num_weighted_ref_pics_minus1. If slice_type is P, sumWeightL0Flags is less than or equal to 24.
[0221] If the current slice is a P slice or a B slice and the value of pic_weighted_pred_flag is 1, L0ToWeightedRefIdx[i] may indicate a mapping between an index in the list of weighted reference pictures and the i-th reference picture L0, where i is in the range from 0 to NumRefIdxActive[0]-1 and may be derived as follows:
[0222]
number
[0223] If the current slice is a B slice and the value of pic_weighted_pred_flag is 1, L1ToWeightedRefIdx[i] may indicate a mapping between an index in the list of weighted reference pictures and the i-th active reference picture L1, where i is in the range from 0 to NumRefIdxActive[1]-1 and may be derived as follows:
[0224]
number
[0225] If luma_weight_l0_flag[i] occurs, it is replaced with luma_weight_flag[L0ToWeightedRefIdx[i]], and if luma_weight_l1_flag[i] occurs, it is replaced with luma_weight_flag[L1ToWeightedRefIdx[i]].
[0226] When LumaWeightL0[i] occurs, it is replaced with LumaWeight[L0ToWeightedRefIdx[i]], and when LumaWeightL1[i] occurs, it is replaced with LumaWeight[L1ToWeightedRefIdx[i]].
[0227] When luma_offset_l0[i] occurs, it is replaced with luma_offset[L0ToWeightedRefIdx[i]], and when luma_offset_l1[i] occurs, it is replaced with luma_offset[L1ToWeightedRefIdx[i]].
[0228] When ChromaWeightL0[i] occurs, it is replaced with ChromaWeight[L0ToWeightedRefIdx[i]], and when ChromaWeightL1[i] occurs, it is replaced with ChromaWeight[L1ToWeightedRefIdx[i]].
[0229] In another embodiment, the slice header syntax may include syntax elements such as those shown in Table 18 below, and the semantics thereof may be as shown in Table 19 below.
[0230] [Table 18]
[0231] [Table 19]
[0232] Referring to Tables 18 and 19, a flag indicating whether or not the prediction weight table syntax is present in the slice header can be signaled. The flag can be signaled in the slice header and can be referred to as slice_weight_pred_present_flag.
[0233] If the value of slice_weight_pred_present_flag is 1, this may indicate that the prediction weight table syntax is present in the slice header. If the value of slice_weight_pred_present_flag is 0, this may indicate that the prediction weight table syntax is not present in the slice header, i.e., that the prediction weight table syntax is present in the picture header.
[0234] In another embodiment, the prediction weight table syntax can be parsed in the slice header, and an adaptation parameter set including syntax elements such as those shown in Table 20 below can be signaled.
[0235] [Table 20]
[0236] Each APS RBSP must be available to the decoding process before it can be referenced by being included in at least one access unit with a TemporalId that is smaller than or equal to the TemporalId of the coded slice NAL unit that references it or is provided through external means.
[0237] aspLayerId refers to the nuh_layer_id of the APS NAL unit. If the layer whose nuh_layer_id is aspLayerId is an independent layer (i.e., vps_independent_layer_flag[GeneralLayerIdx[aspLayerId]] is 1), the APS NAL unit containing the APS RBSP has the same nuh_layer_id as the nuh_layer_id of the coded slice NAL unit that references it. Otherwise, the APS NAL unit containing the APS RBSP has the same nuh_layer_id as the coded slice NAL unit that references it, or the same nuh_layer_id as the nuh_layer_id of a direct dependent layer of the layer containing the coded slice NAL unit that references it.
[0238] All APS NAL units with a particular value of adaptation_parameter_set_id and a particular value of aps_params_type within an access unit have the same content.
[0239] The adaptation_parameter_set_id provides an identifier for the APS so that it can be referenced by other syntax elements.
[0240] If aps_params_type is ALF_APS, SCALING_APS, or PRED_WEIGHT_APS, the value of adaptation_parameter_set_id is in the range from 0 to 7.
[0241] If aps_params_type is LMCS_APS, the value of adaptation_parameter_set_id is in the range from 0 to 3.
[0242] aps_params_type indicates the type of APS parameters included in the APS as shown in Table 21 below. If the value of aps_params_type is 1 (LMCS_APS), the value of adaptation_parameter_set_id is in the range from 0 to 3.
[0243] [Table 21]
[0244] Each type of APS uses a separate value space for adaptation_parameter_set_id.
[0245] APS NAL units (with specific values of adaptation_parameter_set_id and aps_params_type) can be shared between pictures, and other slices within a picture can reference other ALF APS units.
[0246] If the value of aps_extension_flag is 0, this indicates that the aps_extension_data_flag syntax element is not present in the APS RBSP syntax structure. If the value of aps_extension_flag is 1, this indicates that the aps_extension_data_flag syntax element is present in the APS RBSP syntax structure.
[0247] aps_extension_data_flag can have any value.
[0248] As mentioned above, a new aps_params_type (PRED_WEIGHT_APS) can be added to the existing types, and the slice header can be modified to signal the APS ID instead of pred_weight_table() as shown in Table 22 below.
[0249] [Table 22]
[0250] In Table 22, slice_pred_weight_aps_id indicates the adaptation_parameter_set_id of the prediction weight table APS. The TemporalId of an APS NAL unit having the same aps_params_type as PRED_WEIGHT_APS and the same adaptation_parameter_set_id as slice_pred_weight_aps_id is less than or equal to the TemporalId of the coded slice NAL unit.
[0251] If the slice_pred_weight_aps_id syntax element is present in a slice header, the value of slice_pred_weight_aps_id is the same for all slices of a picture.
[0252] In this case, a prediction weight table syntax such as the following Table 23 can be signaled:
[0253] [Table 23-1]
[0254] [Table 23-2]
[0255] In Table 23, if the value of num_lists_active_flag is 1, this may indicate that prediction weight table information is signaled for one reference picture list. If the value of num_lists_active_flag is 0, this may indicate that prediction weight table information for two reference picture lists (L0 and L1) is not signaled.
[0256] numRefIdxActive[i] can be used to indicate the number of active reference indexes. The value of numRefIdxActive[i] is in the range of 0 to 14.
[0257] The syntax in Table 23 indicates whether information for one or two lists is parsed in the APS when num_lists_active_flag is parsed.
[0258] Instead of Table 23, a prediction weight table syntax such as Table 24 below can also be used.
[0259] [Table 24-1]
[0260] [Table 24-2]
[0261] In Table 24, if the value of num_lists_active_flag is 1, this may indicate that prediction weight table information is signaled for one reference picture list. If the value of num_lists_active_flag is 0, this may indicate that prediction weight table information for two reference picture lists is not signaled.
[0262] 6 and 7 show a schematic diagram of an example video / image encoding method and associated components according to embodiments of the present document.
[0263] The video / image encoding method disclosed in Figure 6 may be performed by the (video / image) encoding apparatus 200 disclosed in Figures 2 and 7. Specifically, for example, S600 and S610 in Figure 6 may be performed by the prediction unit 220 of the encoding apparatus 200, and S620 may be performed by the entropy encoding unit 240 of the encoding apparatus 200. The video / image encoding method disclosed in Figure 6 may include the embodiments described above in this document.
[0264] Specifically, referring to FIGS. 6 and 7, the prediction unit 220 of the encoding apparatus may derive motion information of a current block in a current picture based on motion estimation (S600). For example, the encoding apparatus may use an original block in an original picture with respect to the current block to search for a similar reference block with high correlation within a predetermined search range in the reference picture in fractional pixel units, thereby deriving motion information. Block similarity may be derived based on a difference in sample values based on phase. For example, block similarity may be calculated based on the sum of absolute differences (SAD) between the current block (or a template of the current block) and the reference block (or a template of the reference block). In this case, motion information may be derived based on the reference block with the smallest SAD within the search range. The derived motion information may be signaled to the decoding apparatus in various ways based on the inter-prediction mode.
[0265] The prediction unit 220 of the encoding apparatus may perform weighted (sample) prediction on a current block based on motion information of the current block, thereby generating prediction samples (prediction block) and prediction-related information for the current block (S610). The prediction-related information may include prediction mode information (e.g., merge mode, skip mode, etc.), information about motion information, etc. The information about the motion information may include candidate selection information (e.g., merge index, MVP flag, or MVP index) for deriving a motion vector. The information about the motion information may also include the above-mentioned information about MVD and / or reference picture index information. The information about the motion information may also include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. For example, when the slice type of the current slice is a P slice or a B slice, the prediction unit 220 may perform weighted prediction on a current block in the current slice. The weighted prediction may be used when not only bi-prediction but also uni-prediction is applied to the current block.
[0266] The residual processor 230 of the encoding device may generate residual samples and residual information based on the prediction samples generated by the prediction unit 220 and the original picture (original block, original sample). Here, the residual information may include information on the (quantized) transform coefficients for the residual samples as information on the residual samples.
[0267] The adder (or reconstruction unit) of the encoding device can generate reconstructed samples (reconstructed pictures, reconstruction blocks, reconstructed sample arrays) by adding the residual samples generated by the residual processing unit 230 and the predicted samples generated by the prediction unit 220.
[0268] The entropy encoding unit 240 of the encoding device can encode video information including prediction-related information generated by the prediction unit 220, residual information generated by the residual processing unit 230, flags related to weighted prediction, prediction weight table syntax, etc. (S620).
[0269] For example, the entropy encoding unit 240 of the encoding device may encode video information based on at least one of Tables 5 to 23 and output the encoded video information in the form of a bitstream. Specifically, the entropy encoding unit 240 of the encoding device may set a value of a flag related to weighted prediction to 1 based on whether the prediction weight table syntax of this document is included in the picture header of the video information, and may set a value of the flag related to weighted prediction to 0 based on whether the prediction weight table syntax is included in the slice header of the video information. Alternatively, the entropy encoding unit 240 of the encoding device may set a value of the flag related to weighted prediction to 1 based on whether information related to weighted prediction included in the prediction weight table syntax is applied to all slices in the current picture including the current block, and may set a value of the flag related to weighted prediction to 0 based on whether information related to weighted prediction included in the prediction weight table syntax is applied to a slice associated with the slice header among slices in the current picture. When the prediction weight table syntax is included in the picture header, all slices in the picture associated with the picture header may have the same reference picture list, and when the prediction weight table syntax is included in the slice header, slices associated with the slice header may have the same reference picture list. The flag for weighted prediction may be included in a picture parameter set or a picture header of video information and transmitted to a decoding device. The flag for weighted prediction may be information indicating whether information on weighted prediction exists in the picture header.
[0270] Meanwhile, the prediction unit 220 of the encoding apparatus may generate number information for weighted reference pictures in a reference picture list for the weighted prediction based on weighted prediction based on motion information. In this case, the entropy encoding unit 240 of the encoding apparatus may encode video information including the number information. The number information may be included in a prediction weight table syntax in the video information, and in this case, the prediction weight table syntax may also be included in a picture header in the video information. Here, the value of the number information may be the same as the number of the weighted reference pictures in the reference picture list. The prediction weight table syntax may include the same number of weighting coefficient-related flags as the number of the number information. For example, if the value of the number information is n, the prediction weight table syntax may include n weighting coefficient-related flags. When weighted bi-prediction is applied, the prediction weight table syntax may include the number information and / or the weighting coefficient-related flags independently for L0 and L1. In other words, the number information for weighted reference pictures in L0 and the number information for weighted reference pictures in L1 can be signaled independently within the prediction weight table syntax without any dependency on each other (independent of the number of active reference pictures for each list).
[0271] 8 and 9 illustrate an example of a video / image decoding method and related components according to embodiments of the present document.
[0272] The video / image decoding method disclosed in Figure 8 may be performed by the (video / image) decoding apparatus 300 disclosed in Figures 3 and 9. Specifically, for example, S800 and S810 of Figure 8 may be performed by the entropy decoding unit 310 of the decoding apparatus. S820 of Figure 8 may be performed by the residual processing unit 320, the prediction unit 330, and the addition unit 340 of the decoding apparatus. The video / image decoding method disclosed in Figure 8 may include the embodiments described above in this document.
[0273] 8 and 9, the entropy decoding unit 310 of the decoding device may parse a flag related to weighted prediction from a bitstream (S800) and may parse a prediction weight table syntax from the bitstream based on the flag related to weighted prediction (S810). The flag related to weighted prediction may be parsed from a picture parameter set or a picture header of the bitstream and may indicate whether information related to weighted prediction (prediction weight table syntax) is present in the picture header. For example, if the flag related to weighted prediction has a value of 1, the entropy decoding unit 310 of the decoding device may parse the prediction weight table syntax from the picture header of the bitstream, and if the flag related to weighted prediction has a value of 0, the entropy decoding unit 310 may parse the prediction weight table syntax from the slice header of the bitstream. When the value of the flag for weighted prediction is 1, information on weighted prediction included in the prediction weight table syntax may be applied to all slices in the current picture, and when the value of the flag for weighted prediction is 0, information on weighted prediction included in the prediction weight table may be applied to a slice associated with the slice header among the slices in the current picture. When the prediction weight table syntax is parsed from the picture header, all slices in the picture associated with the picture header may have the same reference picture list, and when the prediction weight table syntax is parsed from the slice header, slices associated with the slice header may have the same reference picture list.
[0274] Meanwhile, the entropy decoding unit 310 of the decoding apparatus may parse number information from the prediction weight table syntax. The value of the number information may be the same as the number of weighted reference pictures in a reference picture list. The entropy decoding unit 310 of the decoding apparatus may parse weighting factor-related flags from the prediction weight table syntax based on the number information, the number of which is equal to the number of the number information. For example, if the value of the number information is n, the prediction weight table syntax may include n weighting factor-related flags. When weighted bi-prediction is applied, the prediction weight table syntax may include the number information and / or the weighting factor-related flags independently for each of L0 and L1. For example, the number information for weighted reference pictures in L0 and the number information for weighted reference pictures in L1 may be parsed independently in the prediction weight table syntax without depending on each other (without depending on the number of active reference pictures for each list).
[0275] The decoding apparatus may reconstruct the current picture by performing weighted prediction on a current block in the current picture based on prediction-related information (e.g., inter / intra prediction classification information, intra prediction mode information, inter prediction mode information, and information on weighted prediction) acquired from the bitstream (S820). Here, the information on weighted prediction may include a prediction weight table syntax. For example, the prediction unit 330 of the decoding apparatus may derive weight values for weighted prediction based on parsed weighting coefficient-related flags based on number information in the prediction weight table syntax. Specifically, for example, if the value of number information in the prediction weight table syntax is n, the prediction unit 330 of the decoding apparatus may parse n weighting coefficient-related flags from the prediction weight table syntax. Then, the prediction unit 330 of the decoding apparatus may perform weighted prediction on the current block based on the weight values, and derive predicted samples of the current block.
[0276] Meanwhile, the residual processing unit 320 of the decoding device can generate residual samples based on residual information acquired from a bitstream. The adder 340 of the decoding device can generate reconstructed samples based on the predicted samples generated by the prediction unit 330 and the residual samples generated by the residual processing unit 320. The adder 340 of the decoding device can then generate a reconstructed picture (reconstructed block) based on the reconstructed samples.
[0277] Thereafter, if necessary, in-loop filtering procedures such as deblocking filtering, SAO and / or ALF procedures can be applied to the reconstructed picture to improve the subjective / objective image quality.
[0278] In the above embodiments, the method is described based on a flowchart with a series of steps or blocks, but the embodiment is not limited to the order of the steps, and any step may occur in a different order or simultaneously with other steps described above. Furthermore, those skilled in the art should understand that the steps shown in the flowchart are not exclusive, and other steps may be included, or one or more steps in the flowchart may be deleted without affecting the scope of the embodiments herein.
[0279] The methods according to the embodiments of the present document described above can be implemented in software form, and the encoding device and / or decoding device according to the present document can be included in devices that perform video processing, such as TVs, computers, smartphones, set-top boxes, and display devices.
[0280] In this document, when an embodiment is implemented in software, the method described above may be implemented with modules (processes, functions, etc.) that perform the functions described above. The modules may be stored in memory and executed by a processor. The memory may be internal or external to the processor and may be coupled to the processor in various well-known ways. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described herein may be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in each drawing may be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information (e.g., information on instructions) or algorithms for implementation may be stored on a digital storage medium.
[0281] In addition, the decoding device and encoding device to which the embodiment(s) of this document are applied may be included in a multimedia broadcast transmitting / receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video interaction device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, a custom video (VoD) service providing device, an over-the-top (OTT) video (over-the-top) device, an internet streaming service providing device, a three-dimensional (3D) video device, a virtual reality (VR) device, an augmented reality (AR) device, an image telephone video device, a vehicle terminal (e.g., a vehicle terminal (including an autonomous vehicle), an airplane terminal, a ship terminal, etc.), a medical video device, etc., and may be used to process a video signal or a data signal. For example, over-the-top (OTT) video (over-the-top) device may include a game console, a Blu-ray player, an internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.
[0282] In addition, a processing method to which the embodiment(s) of this document is applied may be produced in the form of a computer-executable program and stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiment(s) of this document may also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. Examples of the computer-readable recording medium include Blu-ray Discs (BDs), Universal Serial Buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable recording medium also includes media embodied in the form of carrier waves (e.g., transmission via the Internet). A bitstream generated by the encoding method may be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.
[0283] Furthermore, the embodiment(s) of this document may be embodied in a computer program product by program code, which may be executed by a computer in accordance with the embodiment(s) of this document, and which may be stored on a computer-readable carrier.
[0284] FIG. 10 illustrates an example of a content streaming system to which the embodiments disclosed herein may be applied.
[0285] Referring to FIG. 10, a content streaming system to which the embodiments of this document are applied can broadly include an encoding server, a streaming server, a web server, a media repository, a user device, and a multimedia input device.
[0286] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server may be omitted.
[0287] The bitstream may be generated by an encoding method or a bitstream generation method applied to an embodiment of this document, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0288] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, which controls commands and responses between devices in the content streaming system.
[0289] The streaming server can receive content from a media repository and / or an encoding server. For example, if content is received from the encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a certain period of time to provide a smooth streaming service.
[0290] Examples of the user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays (HMDs)), digital TVs, desktop computers, and digital signs.
[0291] Each server in the content streaming system can be operated as a distributed server, in which case data received by each server can be processed in a distributed manner.
Claims
1. 1. A video decoding method performed by a video decoding device, comprising: obtaining a bitstream; parsing flags associated with weighted prediction from the bitstream; parsing a prediction weight table syntax from the bitstream based on the flag; deriving prediction samples based on the prediction weight table syntax; the prediction weight table syntax relates luma weighting factors and chroma weighting factors for the weighted prediction; based on the flag having a value of 1, the prediction weight table syntax is parsed from a picture header of the bitstream; The method of claim 1, wherein the prediction weight table syntax is parsed from a slice header of the bitstream based on the value of the flag being 0.
2. 1. A video encoding method performed by a video encoding device, comprising: deriving motion information of a current block; generating a prediction sample by performing weighted prediction on the current block based on the motion information; encoding video information including a flag associated with the weighted prediction and a prediction weight table syntax; encoding the video information to generate a bitstream; the prediction weight table syntax relates luma weighting factors and chroma weighting factors for the weighted prediction; the value of the flag is determined as 1 based on the prediction weight table syntax being included in a picture header of the video information; The method, wherein the value of the flag is determined as 0 based on the prediction weight table syntax being included in a slice header of the video information.
3. 1. A method for transmitting data relating to a video, comprising: obtaining a bitstream for the video, the bitstream comprising: deriving motion information of a current block; generating a prediction sample by performing weighted prediction on the current block based on the motion information; encoding video information including a flag associated with the weighted prediction and a prediction weight table syntax; generating the bitstream by encoding the video information; transmitting the data including the bitstream; the prediction weight table syntax relates luma weighting factors and chroma weighting factors for the weighted prediction; the value of the flag is determined as 1 based on the prediction weight table syntax being included in a picture header of the video information; The method, wherein the value of the flag is determined as 0 based on the prediction weight table syntax being included in a slice header of the video information.