Decoding device, encoding device, and image data transmission device

By parsing and encoding the prediction weight table syntax and deriving the weighted prediction weights, the problem of large amount of information in high-resolution image/video encoding is solved, and more efficient image/video compression is achieved and signaling overhead is reduced.

CN120825575APending Publication Date: 2025-10-21LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511265573.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-12-20
Filing Date
2020-12-11
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing technologies have problems with high-resolution, high-quality image/video encoding, such as large amounts of information and high transmission and storage costs. Especially in image/video broadcasting for virtual reality and immersive media, efficient compression technology is needed to reduce the number of bits and signaling overhead.

Method used

By parsing the prediction weighting table syntax in a video decoding device, deriving the weight of the weighted prediction, and reconstructing the predicted samples of the current block based on the weight, and at the same time generating and encoding the number information about the weighted reference picture list in the video encoding device, redundant signaling and the number of bits are reduced.

Benefits of technology

The video/image compression efficiency is improved, the number of bits and signaling overhead of weighted prediction are reduced, and more efficient image/video encoding is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120825575A_ABST
    Figure CN120825575A_ABST
Patent Text Reader

Abstract

The invention provides a decoding apparatus, an encoding apparatus, and an image data transmission apparatus. A video decoding method performed by a video decoding device according to the present document comprises the steps of: parsing a prediction weighting table syntax from a bitstream; parsing information about the number of reference pictures in the reference picture list from the prediction weighting table syntax; deriving a weight of the weighted prediction based on information on the number of reference pictures; deriving a prediction sample of the current block by performing weighted prediction on the current block based on the weight; and reconstructing the current picture based on the prediction sample, where the prediction weighting table syntax can be parsed from a picture header of the bitstream.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the original invention patent application number 202080096547.X (international application number: PCT / KR2020 / 018134, application date: December 11, 2020, invention name: Image / video coding method and device based on weighted prediction). Technical Field

[0002] The present disclosure relates to a method and apparatus for encoding / decoding an image / video by performing inter-frame prediction based on weighted prediction. Background Art

[0003] Recently, the demand for high-resolution, high-quality images / videos, such as 4K or 8K or higher ultra-high-definition (UHD) images / videos, has been increasing in various fields. As the image / video resolution or quality becomes higher, a relatively larger amount of information or bits is transmitted compared to conventional image / video data. Therefore, if the image / video data is transmitted via a medium such as an existing wired / wireless broadband line or stored in a traditional storage medium, the cost for transmission and storage is likely to increase.

[0004] In addition, there is growing interest and demand for virtual reality (VR) and artificial reality (AR) content and immersive media such as holograms; and there is also growing broadcasting of images / videos that exhibit image / video characteristics that are different from actual images / videos (e.g., game images / videos).

[0005] Therefore, highly efficient image / video compression technology is required to effectively compress and transmit, store, or play high-resolution, high-quality images / videos exhibiting various characteristics as described above. Summary of the Invention

[0006] Technical issues

[0007] The technical subject of this document is to provide a method and device for improving the efficiency of video / image encoding.

[0008] Another technical subject of this document is to provide methods and apparatus for efficiently signaling prediction weighting table syntax.

[0009] Yet another technical subject of this document is to provide methods and devices for reducing signaling overhead regarding weighted prediction.

[0010] Yet another technical subject matter of this document is to provide methods and apparatus for reducing the number of bits used for image information.

[0011] Means of solving the problem

[0012] According to an embodiment of the present document, a video decoding method performed by a video decoding device may include the following steps: parsing a prediction weighting table syntax from a bitstream; parsing information about the number of weighted reference pictures in a reference picture list from the prediction weighting table syntax; deriving weights for weighted prediction based on the quantity information; deriving prediction samples of the current block by performing the weighted prediction on the current block based on the weights; and reconstructing the current picture based on the prediction samples, wherein the prediction weighting table syntax can be parsed from a picture header of the bitstream.

[0013] According to another embodiment of the present document, a video encoding method performed by a video encoding device may include the following steps: deriving motion information about a current block; performing weighted prediction on the current block based on the motion information; generating quantity information about weighted reference pictures in a reference picture list for the weighted prediction; and encoding image information including the quantity information, wherein the quantity information may be included in a prediction weighting table syntax in the image information, and the prediction weighting table syntax may be included in a picture header of the image information.

[0014] According to another embodiment of the present document, a computer-readable digital storage medium may include information that enables a video decoding device to perform a video decoding method, and the video decoding method may include the following steps: parsing a prediction weighting table syntax from image information; parsing information about the number of weighted reference pictures in a reference picture list from the prediction weighting table syntax; deriving weights for weighted prediction based on the quantity information; deriving prediction samples of the current block by performing the weighted prediction on the current block based on the weights; and reconstructing the current picture based on the prediction samples, wherein the prediction weighting table syntax can be parsed from a picture header of the image information.

[0015] Effects of the present invention

[0016] According to the implementation of this document, the overall video / image compression efficiency can be improved.

[0017] According to the embodiments of this document, the prediction weighting table syntax can be efficiently signaled.

[0018] According to the embodiments of this document, redundant signaling for transmitting information about weighted prediction can be reduced.

[0019] According to the embodiments of this document, the number of bits used for weighted prediction can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 An example of a video / image coding system to which embodiments of this document may be applied is schematically illustrated.

[0021] Figure 2 is a diagram schematically illustrating a configuration of a video / image encoding device to which an embodiment of this document can be applied.

[0022] Figure 3 is a diagram schematically illustrating a configuration of a video / image decoding device to which embodiments of this document can be applied.

[0023] Figure 4 An example of a video / image encoding method based on inter-frame prediction is illustrated.

[0024] Figure 5 An inter-frame predictor in an encoding device is schematically illustrated.

[0025] Figure 6 An example of a video / image decoding method based on inter-frame prediction is illustrated.

[0026] Figure 7 An inter-frame predictor in a decoding device is schematically illustrated.

[0027] Figure 8 and Figure 9 An example of a video / image encoding method and related components according to an embodiment of this document is schematically illustrated.

[0028] Figure 10 and Figure 11 An example of a video / image decoding method and related components according to an embodiment of this document is schematically illustrated.

[0029] Figure 12 An example of a content streaming system to which the embodiments disclosed in this document can be applied is illustrated. DETAILED DESCRIPTION

[0030] This document relates to video / image coding. For example, the methods / implementations disclosed in this document can be applied to methods disclosed in the Versatile Video Coding (VVC) standard. In addition, the methods / implementations disclosed in this document can be applied to methods disclosed in the Essential Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the 2nd generation Audio Video Coding standard (AVS2), or next-generation video / image coding standards (e.g., H.267, H.268, etc.).

[0031] Various embodiments related to video / image coding are presented in this document, and unless otherwise specified, the embodiments may also be performed in combination with each other.

[0032] In this document, video may refer to a series of images over time. A picture generally refers to a unit representing an image at a specific time frame, and a slice / tile refers to a unit that constitutes a part of a picture in terms of coding. A slice / tile may include one or more coding tree units (CTUs). A picture may consist of one or more slices / tiles. A picture may consist of one or more tile groups. A tile group may include one or more tiles. A tile may represent a rectangular area of ​​a CTU row within a tile in a picture. A tile may be divided into multiple tiles, each of which may consist of one or more CTU rows within a tile. A tile that is not divided into multiple tiles may also be referred to as a tile. Tile scanning is a specific sequential ordering of the CTUs of a partitioned picture, where CTUs are sequentially ordered within a tile in a CTU raster scan, tiles within a tile are sequentially ordered in a raster scan of the tiles of the tile, and tiles in a picture are sequentially ordered in a raster scan of the tiles of the picture. A tile is a rectangular area of ​​a CTU within a specific tile column and a specific tile row in a picture. A tile column is a rectangular area of ​​a CTU with a height equal to the height of the picture and a width specified by a syntax element in the picture parameter set. A tile row is a rectangular area of ​​a CTU with a height specified by a syntax element in the picture parameter set and a width equal to the width of the picture. Patch scan is a specific sequential ordering of the CTUs that partition a picture, where the CTUs are ordered consecutively in a CTU raster scan within the tile and the tiles in the picture are ordered consecutively in a raster scan of the tiles of the picture. A slice comprises an integer number of tiles of a picture that can be contained in only a single NAL unit. A slice can consist of multiple complete tiles, or just a sequence of consecutive complete tiles of a tile. In this document, tile group and slice can be used interchangeably. For example, in this document, a tile group / tile group header can be referred to as a slice / slice header.

[0033] A pixel or picture element (pel) may refer to the smallest unit constituting a picture (or image). In addition, "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component.

[0034] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of ​​a picture and information related to the area. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, a unit may be used interchangeably with terms such as block or area. In general, an M×N block may include M columns and N rows of samples (or sample arrays) or a set (or array) of transform coefficients. Alternatively, a sample may refer to a pixel value in the spatial domain, and when such a pixel value is transformed into the frequency domain, it may refer to a transform coefficient in the frequency domain.

[0035] In some cases, the term "unit" may be used interchangeably with terms such as "block" or "region." In general, an M×N block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample may generally represent a pixel or pixel value, and may represent only a pixel / pixel value of a luminance component or only a pixel / pixel value of a chrominance component. A sample may be used as a term corresponding to a pixel or picture element configuring a picture (or image).

[0036] The disclosure of this document can be modified in various forms, and specific embodiments thereof will be described and illustrated in the accompanying drawings. The terms used in this document are only used to describe specific embodiments and are not intended to limit the methods disclosed in this document. Singular expressions include the expression "at least one" as long as it is clearly interpreted differently. Terms such as "including" and "having" are intended to indicate the presence of the features, quantities, steps, operations, elements, components, or combinations thereof used in the document, and therefore it should be understood that the possibility of the presence or addition of one or more different features, quantities, steps, operations, elements, components, or combinations thereof is not excluded.

[0037] In addition, the various configurations of the drawings described in this document are independent illustrations for explaining the functions of the features that are different from each other, and do not mean that the various configurations are implemented by different hardware or different software. For example, two or more configurations can be combined to form a single configuration, and a single configuration can also be divided into multiple configurations. Without departing from the gist of the disclosed method of this document, implementations of combined and / or separate configurations are included within the scope of the disclosure of this document.

[0038] In this document, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" may mean "A and / or B". Furthermore, "A, B" may mean "A and / or B". Furthermore, "A / B / C" may mean "at least one of A, B, and / or C". Furthermore, "A / B / C" may mean "at least one of A, B, and / or C".

[0039] Furthermore, in this document, the term "or" should be interpreted as meaning "and / or." For example, the expression "A or B" may include 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted as meaning "additionally or alternatively."

[0040] Furthermore, brackets used in this document may mean "for example." Specifically, when "prediction (intra-frame prediction)" is expressed, it may indicate that "intra-frame prediction" is presented as an example of "prediction." In other words, the term "prediction" in this document is not limited to "intra-frame prediction" and may indicate that "intra-frame prediction" is presented as an example of "prediction." Furthermore, even when "prediction (i.e., intra-frame prediction)" is expressed, it may indicate that "intra-frame prediction" is presented as an example of "prediction."

[0041] In this document, technical features described separately in one drawing may be implemented separately or may be implemented simultaneously.

[0042] Hereinafter, the embodiments of the present invention will be described in detail with reference to the accompanying drawings. In addition, in all drawings, the same reference numerals may be used to indicate the same elements, and the same description of the same elements will be omitted.

[0043] Figure 1 An example of a video / image encoding system to which embodiments of this document can be applied is illustrated.

[0044] Reference Figure 1 The video / image coding system may include a first device (source device) and a second device (receiving device). The source device may send the coded video / image information or data in the form of a file or stream to the receiving device via a digital storage medium or a network.

[0045] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.

[0046] A video source may acquire a video / image through a process of capturing, synthesizing, or generating a video / image. A video source may include a video / image capture device and / or a video / image generation device. For example, a video / image capture device may include one or more cameras, a video / image archive including previously captured videos / images, etc. For example, a video / image generation device may include a computer, a tablet computer, and a smartphone, and may (electronically) generate a video / image. For example, a virtual video / image may be generated by a computer, etc. In this case, the video / image capture process may be replaced by a process of generating relevant data.

[0047] An encoding device can encode input video / images. For compression and coding efficiency, the encoding device may perform a series of processes such as prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output as a bitstream.

[0048] The transmitter can transmit the encoded image / image information or data, output as a bitstream, in the form of a file or stream to a receiver of a receiving device via a digital storage medium or network. Digital storage media may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter may include components for generating a media file in a predetermined file format and may also include components for transmission via a broadcast / communication network. The receiver may receive / extract the bitstream and transmit the received bitstream to a decoding device.

[0049] The decoding device may decode a video / image by performing a series of processes such as dequantization, inverse transformation, and prediction corresponding to the operations of the encoding device.

[0050] The renderer may render the decoded video / image, and the rendered video / image may be displayed on a display.

[0051] Figure 2 is a diagram schematically illustrating a configuration of a video / image encoding device to which embodiments of this document can be applied. Hereinafter, a device referred to as a video encoding device may include an image encoding device.

[0052] Reference Figure 2 , the encoding device 200 includes and is configured with an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the image partitioner 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 described above may be configured by one or more hardware components (e.g., an encoder chipset or processor). In addition, the memory 270 may include a decoded picture buffer (DPB) or may also be configured by a digital storage medium. The hardware components may also include the memory 270 as an internal / external component.

[0053] The image partitioner 210 can partition the input image (or picture, frame) input to the encoding device 200 into one or more processing units. For example, a processing unit can be referred to as a coding unit (CU). In this case, the coding unit can be recursively partitioned from a coding tree unit (CTU) or a maximum coding unit (LCU) based on a quadtree, binary tree, and ternary tree (QTBTTT) structure. For example, a coding unit can be partitioned into multiple coding units of greater depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, the quadtree structure can be applied first, and the binary tree structure and / or ternary tree structure can be applied later. Alternatively, the binary tree structure can be applied first. The encoding process according to this document can be performed based on the final coding unit that is no longer partitioned. In this case, based on image characteristics, coding efficiency, etc., the maximum coding unit can be directly used as the final coding unit, or, if necessary, the coding unit can be recursively partitioned into coding units of greater depth so that a coding unit of optimal size can be used as the final coding unit. Here, the encoding process may include processes such as prediction, transformation, and reconstruction (to be described later). In another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, each of the prediction unit and the transform unit may be split or partitioned from the above-mentioned final coding unit. The prediction unit may be a unit for sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0054] The encoding device 200 can subtract the prediction signal (prediction block, prediction sample array) output from the inter predictor 221 or the intra predictor 222 from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to the transformer 232. In this case, as shown, the unit in the encoding device 200 that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be referred to as the subtractor 231. The predictor 220 can perform prediction on a processing target block (hereinafter referred to as the current block) and generate a prediction block including prediction samples of the current block. The predictor 220 can determine whether to apply intra prediction or inter prediction in units of the current block or CU. As described later in the description of each prediction mode, the predictor 220 can generate various types of information regarding prediction (such as prediction mode information) and send the generated information to the entropy encoder 240, as described below in the description of each prediction mode. Information about the prediction may be encoded by the entropy encoder 240 and output in the form of a bitstream.

[0055] The intra-frame predictor 222 can predict the current block with reference to samples in the current picture. Depending on the prediction mode, the referenced samples can be located near the current block or can be spaced apart. In intra-frame prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. For example, the non-directional mode can include a DC mode and a planar mode. For example, depending on the level of detail of the prediction direction, the directional mode can include 33 directional prediction modes or 65 directional prediction modes. However, this is only an example, and more or fewer directional prediction modes than the above number can be used depending on the settings. The intra-frame predictor 222 can also use the prediction mode applied to the neighboring block to determine the prediction mode applied to the current block.

[0056] The inter-frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. Motion information can include a motion vector and a reference picture index. It can also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, bi-prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in the reference picture. The reference picture including the reference block and the reference picture including the temporally neighboring block can be the same or different. Temporally neighboring blocks can be referred to as collocated reference blocks, collocated CUs (colCUs), etc., and the reference picture including temporally neighboring blocks can be referred to as collocated pictures (colPics). For example, the inter-frame predictor 221 can configure a motion information candidate list based on the neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter-frame predictor 221 can use the motion information of the neighboring blocks as the motion information of the current block. In the case of skip mode, unlike merge mode, the residual signal may not be transmitted. The motion vector prediction (MVP) mode indicates the motion vector of the current block by using the motion vector of the neighboring block as the motion vector predictor and signaling the motion vector difference.

[0057] The predictor 220 can generate a prediction signal based on various prediction methods described later. For example, the predictor 220 can apply intra prediction or inter prediction to predict a block, and can apply intra prediction and inter prediction at the same time. This can be referred to as combined inter and intra prediction (CIIP). In addition, the predictor can be based on an intra block copy (IBC) prediction mode or a palette mode for predicting blocks. The IBC prediction mode or palette mode can be used for image / video encoding of content such as games, for example, screen content coding (SCC). IBC basically performs prediction in the current picture, but it can be performed similarly to inter prediction in that a reference block is derived in the current picture. That is, IBC can use at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample values ​​in the picture can be signaled based on information about the palette table and the palette index.

[0058] The prediction signal generated by the predictor (including the inter-frame predictor 221 and / or the intra-frame predictor 222 ) may be used to generate a reconstructed signal or may be used to generate a residual signal.

[0059] The transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of the following: discrete cosine transform (DCT), discrete sine transform (DST), graph-based transform (GBT), or conditional nonlinear transform (CNT). Here, when the relationship information between pixels is illustrated as a graph, GBT means a transform obtained from the graph. CNT means a transform obtained based on a prediction signal generated by using all previously reconstructed pixels. In addition, the transform process can also be applied to pixel blocks having the same square size, or can also be applied to variable-sized blocks that are not square.

[0060] The quantizer 233 quantizes the transform coefficients and transmits the quantized transform coefficients to the entropy encoder 240, and the entropy encoder 240 encodes the quantized signal (information about the quantized transform coefficients) and outputs the encoded signal as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantizer 233 may rearrange the quantized transform coefficients in the block form in a one-dimensional vector form based on the coefficient scanning order, and may generate information about the transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.

[0061] The entropy encoder 240 can implement various encoding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoder 240 can also encode information necessary for video / image reconstruction (e.g., syntax element values, etc.), in addition to quantized transform coefficients, either together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in units of network abstraction layer (NAL) units. The video / image information can also include information about various parameter sets, such as the adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). Furthermore, the video / image information can also include general constraint information. In this document, the video / image information can include information and / or syntax elements signaled / transmitted from the encoding device to the decoding device. The video / image information can be encoded through the aforementioned encoding process and thus included in the bitstream. The bitstream can be transmitted over a network or stored on a digital storage medium. Here, the network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmitting unit (not shown) for transmitting a signal output from the entropy encoder 240 and / or a storage unit (not shown) for storing the signal may be configured as an internal / external element of the encoding device 200, or the transmitting unit may also be included in the entropy encoder 240.

[0062] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, dequantization and inverse transform can be applied to the quantized transform coefficients by the dequantizer 234 and the inverse transformer 235 to reconstruct the residual signal (residual block or residual sample). The adder 250 can add the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). When there is no residual for the processing target block, such as when the skip mode is applied, the prediction block can be used as a reconstructed block. The adder 250 can be referred to as a restorer or a restored block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the current picture, or can be used for inter-frame prediction of the next picture after filtering, as described below.

[0063] Furthermore, luma mapping and chroma scaling (LMCS) may also be applied during the picture encoding and / or reconstruction process.

[0064] The filter 260 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 270 (specifically, the DPB of the memory 270). For example, the various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 260 can generate various types of information related to filtering and transmit the generated information to the entropy encoder 240, as described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoder 240 and output in the form of a bitstream.

[0065] The modified reconstructed picture transmitted to the memory 270 may be used as a reference picture in the inter predictor 221. When inter prediction is applied by the encoding apparatus, prediction mismatch between the encoding apparatus 200 and the decoding apparatus may be avoided and encoding efficiency may be improved.

[0066] The DPB of the memory 270 can store the corrected reconstructed picture for use as a reference picture in the inter-frame predictor 221. The memory 270 can store the motion information of the block from which the motion information in the current picture is derived (or encoded) and / or the motion information of the block in the reconstructed picture. The stored motion information can be transmitted to the inter-frame predictor 221 to be used as the motion information of the spatially adjacent block or the motion information of the temporally adjacent block. The memory 270 can store the reconstructed samples of the reconstructed block in the current picture and can transmit the reconstructed samples to the intra-frame predictor 222.

[0067] Figure 3 is a diagram for schematically explaining a configuration of a video / image decoding device to which an embodiment of this document can be applied.

[0068] Reference Figure 3 , the decoding device 300 may include and be configured with an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an intra-frame predictor 331 and an inter-frame predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 described above may be configured by one or more hardware components (e.g., a decoder chipset or processor). In addition, the memory 360 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware component may also include the memory 360 as an internal / external component.

[0069] When a bit stream including video / image information is input, the decoding device 300 may generate a Figure 2 The processing of video / image information in the illustrated encoding device reconstructs the image. For example, the decoding device 300 can derive a unit / block based on block partition related information obtained from the bitstream. The decoding device 300 can perform decoding using a processing unit applied to the encoding device. Therefore, for example, the processing unit of decoding can be a coding unit, and the coding unit can be divided from a coding tree unit or a maximum coding unit according to a quadtree structure, a binary tree structure and / or a ternary tree structure. One or more transform units can be derived from the coding unit. In addition, the reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a reproduction device.

[0070] The decoding device 300 can receive the data in the form of a bit stream from Figure 2The received signal is output by the encoding device, and the entropy decoder 310 can decode the received signal. For example, the entropy decoder 310 can parse the bitstream to derive information required for image reconstruction (or picture reconstruction) (e.g., video / image information). The video / image information may also include information about various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may also include general constraint information. The decoding device may also decode the picture based on the information about the parameter set and / or general constraint information. The signaled / received information and / or syntax elements described later in this document may be decoded and obtained from the bitstream through a decoding process. For example, the entropy decoder 310 may decode the information within the bitstream based on a coding method such as exponential Golomb coding, context-adaptive variable length coding (CAVLC), or context-adaptive binary arithmetic coding (CABAC), and output syntax elements required for image reconstruction and quantized values ​​of transform coefficients for the residual. More specifically, the CABAC entropy decoding method can receive bins corresponding to various syntax elements in the bitstream, determine a context model by using the decoding target syntax element information, the decoding information of the decoding target block, or the information of the symbol / bin decoded in the previous stage, and perform arithmetic decoding on the bin by predicting the probability of the bin appearing according to the determined context model, and generate a symbol corresponding to the value of each syntax element. In this case, the CABAC entropy decoding method can update the context model by using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. The information related to the prediction among the information decoded by the entropy decoder 310 can be provided to the predictor (inter-frame predictor 332 and intra-frame predictor 331), and the residual value (i.e., quantized transform coefficients and related parameter information) that has been entropy decoded in the entropy decoder 310 can be input to the residual processor 320.

[0071] The residual processor 320 can derive a residual signal (residual block, residual sample, or residual sample array). In addition, information about filtering among the information decoded by the entropy decoder 310 can be provided to the filter 350. At the same time, a receiver (not shown) for receiving a signal output from the encoding device can also be configured as an internal / external element of the decoding device 300, or the receiver can be a component of the entropy decoder 310. At the same time, the decoding device according to this document can be referred to as a video / image / picture decoding device, and the decoding device can be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of the following: a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.

[0072] The dequantizer 321 may dequantize the quantized transform coefficients to output the transform coefficients. The dequantizer 321 may rearrange the quantized transform coefficients in a two-dimensional block form. In this case, the rearrangement may be performed based on the coefficient scanning order performed by the encoding device. The dequantizer 321 may dequantize the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.

[0073] The inverse transformer 322 performs an inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).

[0074] In this document, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantized transform coefficient may be referred to as a transform coefficient. When transform / inverse transform is omitted, the transform coefficient may be referred to as a coefficient or a residual coefficient, or may still be referred to as a transform coefficient for consistency of expression.

[0075] In this document, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about the transform coefficients, and the information about the transform coefficients may be signaled using residual coding syntax. The transform coefficients may be derived based on the residual information (or information about the transform coefficients), and the scaled transform coefficients may be derived by inverse transforming (scaling) the transform coefficients. Residual samples may be derived based on the inverse transform (transform) of the scaled transform coefficients. This may also be applied / expressed in other parts of this document.

[0076] The predictor 330 may perform prediction on the current block and generate a prediction block including prediction samples of the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block and determine a specific intra / inter prediction mode based on the information on prediction output from the entropy decoder 310.

[0077] The predictor 330 can generate a prediction signal based on various prediction methods described below. For example, the predictor can apply intra prediction or inter prediction to predict a block, and can apply intra prediction and inter prediction simultaneously. This can be referred to as combined inter and intra prediction (CIIP). In addition, the predictor can predict a block based on an intra block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for image / video encoding of content such as games, such as screen content coding (SCC). IBC can basically perform prediction in the current picture, but can be performed similarly to inter prediction so that a reference block is derived within the current picture. That is, IBC can use at least one inter prediction technique described in this document. The palette mode can be considered an example of intra coding or intra prediction. When the palette mode is applied, information about the palette table and palette index can be included in the video / image information and notified with a signal.

[0078] The intra-frame predictor 331 can predict the current block by referencing samples in the current picture. Depending on the prediction mode, the referenced samples may be located near the current block or may be spaced apart from the current block. In intra-frame prediction, the prediction mode may include multiple non-directional modes and multiple directional modes. The intra-frame predictor 331 may determine the prediction mode to be applied to the current block by using the prediction modes applied to neighboring blocks.

[0079] The inter-frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. Motion information can include a motion vector and a reference picture index. Motion information can also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, bi-prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter-frame predictor 332 can construct a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and information about the prediction can include information indicating the inter-frame prediction mode used for the current block.

[0080] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block or prediction sample array) output from the predictor (including the inter-frame predictor 332 and / or the intra-frame predictor 331). If there is no residual for the processing target block, for example, when skip mode is applied, the prediction block can be used as the reconstructed block.

[0081] The adder 340 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra-frame prediction of the next block to be processed in the current picture, and as described later, may also be output through filtering or may also be used for inter-frame prediction of the next picture.

[0082] In addition, Luma Mapping and Chroma Scaling (LMCS) can also be applied to the picture decoding process.

[0083] The filter 350 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and store the modified reconstructed image in the memory 360, specifically, in the DPB of the memory 360. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0084] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter-frame predictor 332. The memory 360 can store the motion information of the block from which the motion information in the current picture is derived (decoded) and / or the motion information of the block in the reconstructed picture. The stored motion information can be transmitted to the inter-frame predictor 332 to be used as the motion information of the spatially adjacent block or the motion information of the temporally adjacent block. The memory 360 can store the reconstructed samples of the reconstructed block in the current picture and transmit the reconstructed samples to the intra-frame predictor 331.

[0085] In this document, the embodiments described in the filter 260 , the inter predictor 221 , and the intra predictor 222 of the encoding apparatus 200 may be equally applied to or correspond to the filter 350 , the inter predictor 332 , and the intra predictor 331 .

[0086] In addition, the video / image encoding method according to the present document can be performed based on the following partition structure. Specifically, the above-mentioned prediction, residual processing ((inverse) transform and (de)quantization), syntax element encoding and filtering processes can be performed based on the CTU and CU (and / or TU and PU) derived based on the partition structure. The block partitioning process can be performed by the image partitioner 210 of the above-mentioned encoding device, and the partition-related information can be processed by the entropy encoder 240 (encoding) and can be transmitted to the decoding device in the form of a bitstream. The entropy decoder 310 of the decoding device can derive the block partition structure of the current picture based on the partition-related information obtained from the bitstream, and based on this, a series of processes for image decoding (e.g., prediction, residual processing, block / picture reconstruction, in-loop filtering, etc.) can be performed. The CU size and the TU size can be equal to each other, or multiple TUs can exist within the CU area. In addition, the CU size can generally represent the luminance component (sample) coding block (CB) size. The TU size can generally represent the luminance component (sample) transform block (TB) size. The chroma component (sample) CB or TB size can be derived based on the luma component (sample) CB or TB size according to the component ratio according to the color format (chroma format, for example, 4:4:4, 4:2:2, 4:2:0, etc.) of the picture / image. The TU size can be derived based on maxTbSize. For example, if the CU size is larger than maxTbSize, multiple TUs (TBs) of maxTbSize can be derived from the CU, and transformation / inverse transformation can be performed in units of TU (TB). In addition, for example, in the case of applying intra-frame prediction, the intra-frame prediction mode / type can be derived in units of CU (or CB), and the neighboring reference sample derivation and prediction sample generation process can be performed in units of TU (or TB). In this case, one or more TUs (or TBs) may exist in one CU (or CB) area, and in this case, multiple TUs (or TBs) may share the same intra-frame prediction mode / type.

[0087] In addition, in the video / image coding according to this document, the image processing unit may have a hierarchical structure. A picture may be partitioned into one or more tiles, tiles, slices, and / or tile groups. A slice may include one or more tiles. A tile may include one or more CTU rows within a tile. A slice may include an integer number of tiles of a picture. A tile group may include one or more tiles. A tile may include one or more CTUs. A CTU may be partitioned into one or more CUs. A tile represents a rectangular area of ​​a CTU within a specific tile column and a specific tile row in a picture. A tile group may include an integer number of tiles according to a raster scan of tiles in a picture. A slice header may carry information / parameters that can be applied to the corresponding slice (block in a slice). If the encoding / decoding device has a multi-core processor, the encoding / decoding processes for tiles, slices, tiles, and / or tile groups may be processed in parallel. In this document, slices or tile groups may be used interchangeably. That is, the tile group header may be referred to as a slice header. Here, the slice may have one of the slice types including intra (I) slice, predicted (P) slice, and bi-predicted (B) slice. When predicting a block in an I slice, inter prediction may not be used, and only intra prediction may be used. Of course, even in this case, signaling may be performed by encoding the original sample values ​​without prediction. With respect to a block in a P slice, intra prediction or inter prediction may be used, and when inter prediction is used, only unidirectional prediction may be used. In addition, with respect to a block in a B slice, intra prediction or inter prediction may be used, and when inter prediction is used, up to bi-prediction may be used.

[0088] Taking into account coding efficiency or parallel processing or based on the characteristics of the video image (e.g., resolution), the encoding device can determine the block / block group, tile, slice, and maximum and minimum coding unit sizes, and their information or information that can derive them can be included in the bitstream.

[0089] The decoding device can obtain information indicating the tiles / tile groups, tiles, and slices of the current picture, as well as whether the CTU in the tile has been partitioned into multiple coding units. Efficiency can be improved by only obtaining (sending) this information under specific conditions.

[0090] As described above, a picture may include multiple slices, and a slice may include a slice header and slice data. In this case, a picture header may be further added for multiple slices in a picture (a collection of slice headers and slice data). The picture header (picture header syntax) may include information / parameters that are commonly applied to the picture. The slice header (slice header syntax) may include information / parameters that are commonly applied to the slices. The adaptation parameter set (APS) or the picture parameter set (PPS) may include information / parameters that are commonly applied to one or more pictures. The sequence parameter set (SPS) may include information / parameters that are commonly applied to one or more sequences. The video parameter set (VPS) may include information / parameters that are commonly applied to multiple layers. The decoding parameter set (DPS) may include information / parameters that are commonly applied to the entire video. The DPS may include information / parameters related to the concatenation of the coded video sequence (CVS).

[0091] In this document, the high-level syntax may include at least one of an APS syntax, a PPS syntax, an SPS syntax, a VPS syntax, a DPS syntax, a picture header syntax, and a slice header syntax.

[0092] In addition, for example, information on partitioning and configuration of tiles / tile groups / tiles / slices may be configured in an encoding device based on a high-level syntax and may be transmitted to a decoding device in the form of a bitstream.

[0093] The video / image encoding process based on inter-frame prediction may schematically include the following contents.

[0094] Figure 4 An example of a video / image encoding method based on inter-frame prediction is illustrated, and Figure 5 An inter-frame predictor in an encoding device is schematically illustrated.

[0095] Reference Figure 4 and Figure 5, the encoding device performs inter prediction on the current block (S400). The encoding device may derive an inter prediction mode and motion information for the current block, and may generate a prediction sample of the current block. Here, the processes for determining the inter prediction mode, deriving motion information, and generating prediction samples may be performed simultaneously, or one process may be performed before the other process. For example, the inter predictor 221 of the encoding device may include a prediction mode determiner 221_1, a motion information deriver 221_2, and a prediction sample deriver 221_3, wherein the prediction mode determiner 221_1 may determine the prediction mode of the current block, the motion information deriver 221_2 may derive motion information about the current block, and the prediction sample deriver 221_3 may derive the prediction sample of the current block. For example, the inter predictor of the encoding device may search for a block similar to the current block within a predetermined area (search area) of a reference picture through motion estimation, and may derive a reference block whose difference with the current block is a minimum value or a predetermined reference level or less. The inter-frame predictor can derive a reference picture index based on the reference block, which indicates the reference picture where the reference block is located, and can derive a motion vector based on the position difference between the reference block and the current block. The encoding device can determine the mode to be applied to the current block among various prediction modes. The encoding device can compare the rate-distortion (RD) cost of various prediction modes and determine the optimal prediction mode for the current block.

[0096] For example, when skip mode or merge mode is applied to the current block, the encoding device may construct a merge candidate list and derive a reference block from the reference blocks indicated by the merge candidates included in the merge candidate list, the reference block having a difference with the current block that is a minimum value or a predetermined reference level or less. In this case, a merge candidate associated with the derived reference block may be selected, and merge index information indicating the selected merge candidate may be generated and signaled to the decoding device. Motion information about the selected merge candidate may be used to derive motion information about the current block.

[0097] In another example, when the (A)MVP mode is applied to the current block, the encoding device can construct an (A)MVP candidate list and can use the motion vector of the MVP candidate selected from the motion vector predictor (MVP) candidates included in the (A)MVP candidate list as the MVP of the current block. For example, in this case, the motion vector of the reference block derived by motion estimation can be used as the motion vector of the current block, and the MVP candidate with the smallest difference from the motion vector of the current block among the MVP candidates can be the selected MVP candidate. A motion vector difference (MVD) can be derived, which is the difference obtained by subtracting the MVP from the motion vector of the current block. In this case, information about the MVD can be notified to the decoding device with a signal. When the (A)MVP mode is applied, the value of the reference picture index can be configured as reference picture index information and can be separately notified to the decoding device with a signal.

[0098] The encoding apparatus may induce residual samples based on the prediction samples (S410). The encoding apparatus may induce residual samples by comparing original samples of the current block with the prediction samples.

[0099] The encoding device encodes the image information including prediction information and residual information (S420). The encoding device may output the encoded image information in the form of a bitstream. The prediction information is information related to the prediction process and may include prediction mode information (e.g., a skip flag, a merge flag, or a mode index) and information about motion information. The information about the motion information may include candidate selection information (e.g., a merge index, an mvp flag, or an mvp index), which is information used to derive a motion vector. In addition, the information about the motion information may include information about MVD and / or reference picture index information. In addition, the information about the motion information may include information indicating whether L0 prediction, L1 prediction, or dual prediction is applied. The residual information is information about the residual sample. The residual information may include information about the quantized transform coefficients of the residual sample.

[0100] The output bitstream may be stored in a (digital) storage medium and sent to the decoding device, or may be sent to the decoding device over a network.

[0101] As described above, the encoding device can generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on the reference samples and the residual samples. The reconstructed picture is used by the encoding device to derive the same result as the prediction result derived by the decoding device and is used to improve coding efficiency. Therefore, the encoding device can store the reconstructed picture (or reconstructed samples and reconstructed blocks) in a memory and use it as a reference picture for inter-frame prediction. As described above, the in-loop filtering process can be further applied to the reconstructed picture.

[0102] For example, the video / image decoding process based on inter-frame prediction may schematically include the following contents.

[0103] Figure 6 An example of a video / image decoding method based on inter-frame prediction is illustrated, and Figure 7 An inter-frame predictor in a decoding device is schematically illustrated.

[0104] The decoding device may perform operations corresponding to the above operations performed by the encoding device. The decoding device may predict the current block based on the received prediction information and may derive a prediction sample.

[0105] Specifically, refer to Figure 6 and Figure 7 , the decoding device may determine a prediction mode of the current block based on the prediction information received from the bitstream (S500). The decoding device may determine which inter-frame prediction mode to apply to the current block based on the prediction mode information in the prediction information.

[0106] For example, whether to apply merge mode to the current block or whether to determine (A)MVP mode can be determined based on the merge flag. Alternatively, one of various inter-frame prediction mode candidates can be selected based on the merge index. Inter-frame prediction mode candidates can include various inter-frame prediction modes such as skip mode, merge mode and / or (A)MVP mode.

[0107] The decoding device derives motion information about the current block based on the determined inter-frame prediction mode (S510). For example, when the skip mode or merge mode is applied to the current block, the decoding device can construct a merge candidate list described later and select a merge candidate from the merge candidates included in the merge candidate list. The selection can be performed based on the above-mentioned selection information (merge index). The motion information about the selected merge candidate can be used to derive the motion information about the current block. The motion information about the selected merge candidate can be used as the motion information about the current block.

[0108] In another example, when the (A)MVP mode is applied to the current block, the decoding device may construct an (A)MVP candidate list and may use the motion vector of a motion vector predictor (MVP) candidate selected from the MVP candidates included in the (A)MVP candidate list as the MVP of the current block. The selection may be performed based on the above-mentioned selection information (MVP flag or MVP index). In this case, the decoding device may derive the MVD of the current block based on the information about the MVD, and may derive the motion vector of the current block based on the MVP and MVD of the current block. In addition, the decoding device may derive the reference picture index of the current block based on the reference picture index information. The picture indicated by the reference picture index in the reference picture list of the current block may be derived as the reference picture referenced by the inter-frame prediction of the current block.

[0109] The motion information about the current block may be derived without constructing a candidate list, in which case the construction of the candidate list may be omitted.

[0110] The decoding device may generate prediction samples for the current block based on the motion information about the current block (S520). In this case, a reference picture may be derived based on a reference picture index of the current block, and the prediction samples of the current block may be derived using samples of the reference block indicated by the motion vector of the current block in the reference picture. In this case, as described later, a prediction sample filtering process may be further performed on all or some of the prediction samples of the current block, depending on the situation.

[0111] For example, the inter-frame predictor 332 of the decoding device may include a prediction mode determiner 332_1, a motion information deriver 332_2 and a prediction sample deriver 332_3, wherein the prediction mode determiner 332_1 can determine the prediction mode of the current block based on the received prediction mode information, the motion information deriver 332_2 can derive motion information (motion vector and / or reference picture index) about the current block based on the received information about the motion information, and the prediction sample deriver 332_3 can derive the prediction sample of the current block.

[0112] The decoding device generates residual samples of the current block based on the received residual information (S530). The decoding device may generate reconstructed samples of the current block based on the predicted samples and the residual samples, and may generate a reconstructed picture based on the reconstructed samples (S540). Subsequently, as described above, the in-loop filtering process may be further applied to the reconstructed picture.

[0113] The prediction block for the current block can be derived based on motion information derived according to the prediction mode of the current block. The prediction block can include prediction samples (prediction sample array) for the current block. When the motion vector of the current block indicates a partial sample unit, an interpolation process can be performed to derive prediction samples for the current block based on reference samples in the partial sample unit in a reference picture. When affine inter prediction is applied to the current block, prediction samples can be generated based on motion vectors (MVs) in sample / sub-block units. When bi-prediction is applied, prediction samples derived by a weighted sum or weighted average of prediction samples derived based on L0 prediction (i.e., prediction using reference pictures in reference picture list L0 and MVL0) and prediction samples derived based on L1 prediction (i.e., prediction using reference pictures in reference picture list L1 and MVL1 (depending on the phase)) can be used as prediction samples for the current block. Applying bi-prediction where the reference pictures used for L0 prediction and L1 prediction are located in different temporal directions relative to the current picture (i.e., corresponding to bi-prediction and bi-directional prediction) is referred to as true bi-prediction.

[0114] As described above, reconstructed samples and a reconstructed picture may be generated based on the derived prediction samples, and then an in-loop filtering process may be performed.

[0115] In inter-frame prediction, weighted sample prediction can be used. Weighted sample prediction can be called weighted prediction. When the slice type of the current slice where the current block (e.g., CU) is located is a P slice or a B slice, weighted prediction can be applied. That is, weighted prediction can be used not only when bi-prediction is applied, but also when uni-prediction is applied. For example, as described below, weighted prediction can be determined based on weightedPredFlag, and the value of weightedPredFlag can be determined based on the signaled pps_weighted_pred_flag (in the case of a P slice) or pps_weighted_bipred_flag (in the case of a B slice). For example, when slice_type is p, weightedPredFlag can be set to pps_weighted_pred_flag. Otherwise (when slice_type is B), weightedPredFlag can be set to pps_weighted_bipred_flag.

[0116] The prediction samples or the values ​​of the prediction samples as the output of the weighted prediction may be referred to as pbSamples.

[0117] The weighted prediction process can be largely divided into the default weighted (sample) prediction process and the explicit weighted (sample) prediction process. The weighted (sample) prediction process can refer only to the explicit weighted (sample) prediction process. For example, when the value of weightedPredFlag is 0, the value of the prediction sample (pbSamples) can be derived based on the default weighted (sample) prediction process. When the value of weightedPredFlag is 1, the value of the prediction sample (pbSamples) can be derived based on the explicit weighted (sample) prediction process.

[0118] When bi-prediction is applied to the current block, a prediction sample can be derived based on a weighted average. Conventionally, a bi-prediction signal (i.e., a bi-prediction sample) can be derived by simply averaging an L0 prediction signal (L0 prediction sample) and an L1 prediction signal (L1 prediction sample). That is, the bi-prediction sample is derived as an average of an L0 prediction sample based on an L0 reference picture and MVL0 and an L1 prediction sample based on an L1 reference picture and MVL1. However, according to this document, when bi-prediction is applied, a bi-prediction signal (bi-prediction sample) can be derived by weighted averaging an L0 prediction signal and an L1 prediction signal.

[0119] Bidirectional optical flow (BDOF) can be used to refine the bi-prediction signal. BDOF is used to generate prediction samples by calculating improved motion information when bi-prediction is applied to the current block (e.g., CU), and the process of calculating the improved motion information can be included in the motion information derivation operation.

[0120] For example, BDOF can be applied horizontally in 4x4 sub-blocks. That is, BDOF can be performed in units of 4×4 sub-blocks within the current block. BDOF can only be applied to the luma component. Alternatively, BDOF can be applied only to the chroma components, or to both the luma and chroma components.

[0121] As described above, High Level Syntax (HLS) can be encoded for video / image encoding / signaling. Video / image information can be included in HLS.

[0122] A coded picture may include one or more slices. Parameters describing the coded picture are signaled in the picture header, and parameters describing the slice are signaled in the slice header. The picture header is carried in the form of a separate NAL unit. The slice header is present at the beginning of the NAL unit, which includes the slice payload (i.e., slice data).

[0123] Each picture is associated with a picture header. A picture may include different types of slices (intra-coded slices (i.e., I slices) and inter-coded slices (i.e., P slices and B slices). Therefore, the picture header may include syntax elements required for intra-frame slices and inter-frame slices of the picture.

[0124] A picture can be divided into sub-pictures, tiles, and / or slices. Sub-picture signaling can be present in the sequence parameter set (SPS), and tile and square slice signaling can be present in the picture parameter set (PPS). Raster scan slice signaling can be present in the slice header.

[0125] When weighted prediction is applied to inter prediction of a current block, weighted prediction may be performed based on information about weighted prediction.

[0126] The weighted prediction process can be started based on two flags in SPS.

[0127] For example, the syntax elements shown in Table 1 below may be included in the SPS syntax regarding weighted prediction.

[0128] [Table 1]

[0129]

[0130] In Table 1, a value of sps_weighted_pred_flag equal to 1 may indicate that weighted prediction is applied to the P slice referencing the SPS.

[0131] A value of sps_weighted_bipred_flag equal to 1 may indicate that weighted prediction is applied to B slices of the reference SPS. A value of sps_weighted_bipred_flag equal to 0 may indicate that weighted prediction is not applied to B slices of the reference SPS.

[0132] Two flags signaled in the SPS indicate whether weighted prediction is applied to P slices and B slices in a coded video sequence (CVS).

[0133] The syntax elements shown in Table 2 below may be included in the PPS syntax regarding weighted prediction.

[0134] [Table 2]

[0135]

[0136] In Table 2, a value of pps_weighted_pred_flag equal to 0 may indicate that weighted prediction is not applied to the P slice referencing the PPS. A value of pps_weighted_pred_flag equal to 1 may indicate that weighted prediction is applied to the P slice referencing the PPS. When the value of sps_weighted_pred_flag is 0, the value of pps_weighted_pred_flag is 0.

[0137] A value of pps_weighted_bipred_flag equal to 0 may indicate that weighted prediction is not applied to B slices referencing the PPS. A value of pps_weighted_bipred_flag equal to 1 may indicate that explicit weighted prediction is applied to B slices referencing the PPS. When the value of sps_weighted_bipred_flag is 0, the value of pps_weighted_bipred_flag is 0.

[0138] In addition, the syntax elements shown in Table 3 below may be included in the slice header syntax.

[0139] [Table 3]

[0140]

[0141] In Table 3, slice_pic_parameter_set_id indicates the value of pps_pic_parameter_set_id of the used PPS. The value of slice_pic_parameter_set_id ranges from 0 to 63, inclusive.

[0142] The value of the temporary ID (TempralID) of the current picture needs to be greater than or equal to the TempralID value of the PPS having the same pps_pic_parameter_set_id as slice_pic_parameter_set_id.

[0143] The prediction weight table syntax may include information about the weighted predictions shown in Table 4 below.

[0144] [Table 4]

[0145]

[0146] In Table 4, luma_log2_weight_denom is the base-2 logarithm of the denominator of all luma weighting factors. The value of luma_log2_weight_denom is within the range from 0 to 7, inclusive.

[0147] delta_chroma_log2_weight_denom is the difference in the base-2 logarithms of the denominators of all chroma weighting factors. When delta_chroma_log2_weight_denom is not present, delta_chroma_log2_weight_denom is inferred to be 0.

[0148] ChromaLog2WeightDenom is derived as luma_log2_weight_denom+

[0149] delta_chroma_log2_weight_denom, and its value is within the range from 0 to 7 inclusive.

[0150] A value of luma_weight_10_flag[i] equal to 1 indicates that weighting factors for the luma component predicted using list 0 (L0) of (reference picture) RefPicList[0][i] are present. A value of luma_weight_10_flag[i] equal to 0 indicates that these weighting factors are not present.

[0151] A value of chroma_weight_10_flag[i] equal to 1 indicates that weighting factors for the chroma prediction values ​​of the L0 prediction of RefPicList[0][i] are used. A value of chroma_weight_10_flag[i] equal to 0 indicates that these weighting factors are not present. When chroma_weight_10_flag[i] is not present, chroma_weight_10_flag[i] is inferred to be 0.

[0152] delta_luma_weight_10[i] is the difference of the weighting factors for the luma prediction values used for L0 prediction using RefPicList[0][i].

[0153] LumaWeightL0[i] is inferred to be (1<<luma_log2_weight_denom)+delta_luma_weight_l0[i]. When luma_weight_10_flag[i] is 1, the value of delta_luma_weight_10[i] is included in the range from -128 to 127. When luma_weight_10_flag[i] is 0, LumaWeightL0[i] is inferred to be 2 luma _log2_weight_denom .

[0154] luma_offset_10[i] is the cumulative offset for the luma prediction values used for L0 prediction using RefPicList[0][i]. The value of luma_offset_10[i] is included in the range from -128 to 127. When the value of luma_weight_10_flag[i] is 0, the value of luma_offset_10[i] is inferred to be 0.

[0155] delta_chroma_weight_l0[i][j] is the difference of the weighting factors for the chroma prediction values used for L0 prediction using RefPicList[0][i], where for Cb, j is 0, and for Cr, j is 1.

[0156] ChromaWeightL0[i][j] is derived as (1<<Chromalog2WeightDenom)+delta_chroma_weight_l0[i][j]. When chroma_weight_10_flag[i] is 1, the value of delta_chroma_weight_10[i][j] is included in the range from -128 to 127. When chroma_weight_l0_flag[i] is 0, ChromaWeightL0[i][j] is inferred to be 2 ChromaLog2WeightDenom .

[0157] delta_chroma_offset_l0[i][j] is the cumulative offset for the chroma prediction values used for L0 prediction using RefPicList[0][i], where for Cb, j is 0, and for Cr, j is 1.

[0158] The value of delta_chroma_offset_10[i][j] is included in the range from -4x128 to 4x127. When the value of chroma_weight_10_flag[i] is 0, the value of ChromaOffsetL0[i][j] is inferred to be 0.

[0159] The prediction weighting table syntax is often used to modify sequences during scene changes. When the PPS flag for weighted prediction is enabled and the slice type is P, or when the PPS flag for weighted bi-prediction is enabled and the slice type is B, the existing prediction weighting table syntax is signaled in the slice header. However, it is often the case that when a scene changes, the prediction weighting table needs to be adjusted for one or more frames. In general, when multiple frames share a PPS, it may not be necessary to signal information about weighted prediction for all frames that reference the PPS.

[0160] The following figures are provided to describe specific examples of this document. Since specific terms of the devices illustrated in the figures or specific signal / message terms are provided for illustration, the technical features of the present disclosure are not limited to the specific terms used in the following figures.

[0161] This document provides the following methods to address the above issues. These methods can be used independently or in combination with each other.

[0162] 1. Tools that can apply weighted prediction at the picture level instead of the slice level (information about weighted prediction). The weighting values ​​are applied to a specific reference picture of the picture and are used for all slices of the picture.

[0163] 2. The prediction weighting table syntax can be signaled at the picture level instead of the slice level. To this end, the prediction weighting table syntax can be signaled in the picture header (PH) or the picture parameter set (PPS).

[0164] 3. When weighted prediction is applied to a picture, all slices in the picture can have the same active reference picture. This includes the order of active reference pictures in the reference picture list (RPL) (i.e., L0 for P slices, L0 and L1 for B slices).

[0165] 4. Alternatively, when the above does not apply, the following may apply.

[0166] a. The signaling of weighted prediction is independent of the signaling of the reference picture list. That is, in the signaling of the prediction weight table, there is no assumption about the order of the reference pictures in the reference picture list.

[0167] b. There is no signaling of weighted prediction values ​​for reference pictures in L0 and L1. For reference pictures, weighted values ​​are provided directly.

[0168] c. Only one loop, rather than two, may be used to signal the weighting values ​​of reference pictures. In each loop, the reference picture associated with the signaled weighting value is first identified.

[0169] d. Reference picture identification is based on the picture order count (POC) value.

[0170] e. For bit saving, the delta POC value between the reference picture and the current picture may be signaled instead of signaling the POC value of the reference picture.

[0171] 5. In addition to item 4, to signal the delta POC value between the reference picture and the current picture, the following may be applied so that the absolute delta POC value may be signaled as follows.

[0172] a. The first signaled delta POC is the delta between the POC of the reference picture and the POC of the current picture.

[0173] b. The remaining signaled delta POCs (ie, where i starts at 1) are the deltas between the POC of the i-th reference picture and the POC of the (i-1)-th reference picture.

[0174] 6. The two flags in PPS can be unified into a single control flag (eg, pps_weighted_pred_flag). This flag can be used to indicate the presence of additional flags in the picture header.

[0175] a. The flag in PH can be conditional on the PPS flag and can further indicate the presence of pred_weighted_table() data (prediction weighting table syntax) when the NAL unit type is not instantaneous decoding refresh (IDR).

[0176] 7. The two flags signaled in PPS (pps_weighted_pred_flag and pps_weighted_bipred_flag) can be unified into one flag. The one flag can use the existing name of pps_weighted_pred_flag.

[0177] 8. A flag may be signaled in the picture header to indicate whether weighted prediction is applied to the picture associated with the picture header. The flag may be called pic_weighted_pred_flag.

[0178] a. The presence of pic_weighted_pred_flag may be conditional on the value of pps_weighted_pred_flag. When the value of pps_weighted_pred_flag is 0, pic_weighted_pred_flag is not present and its value may be inferred to be 0.

[0179] b. When the value of pic_weighted_pred_flag is 1, the signaling of pred_weighted_table() may exist in the picture header.

[0180] 9. Alternatively, when weighted prediction is enabled (ie, the value of pps_weighted_pred_flag is 1 or the value of pps_weighted_bipred_flag is 1), information about weighted prediction may still exist in the slice header, and the following may apply.

[0181] a. A new flag may be signaled to indicate whether information about weighted prediction is present in the slice header. The flag may be called slice_weighted_pred_present_flag.

[0182] b. The presence of slice_weighted_pred_present_flag can be determined according to the slice type and the values ​​of pps_weighted_pred_flag and pps_weighted_bipred_flag.

[0183] In this document, information about weighted prediction may include information / syntax elements related to weighted prediction described in Tables 1 to 4. Video / image information may include various inter-frame prediction information, such as information about weighted prediction, residual information, and inter-frame prediction mode information. The inter-frame prediction mode information may include information / syntax elements, such as information indicating whether merge mode or MVP mode is applied to the current block, and selection information for selecting one of the motion candidates in the motion candidate list. For example, when merge mode is applied to the current block, a merge candidate list is constructed based on neighboring blocks of the current block, and one candidate for deriving motion information about the current block may be selected / used from the merge candidate list (based on the merge index). In another example, when MVP mode is applied to the current block, an MVP candidate list may be constructed based on neighboring blocks of the current block, and one candidate for deriving motion information about the current block may be selected / used from the MVP candidate list (based on the MVP flag).

[0184] In one embodiment, for weighted prediction in inter prediction, the PPS may include the syntax elements shown in Table 5 below, and the semantics of the syntax elements may be as shown in Table 6 below.

[0185] [Table 5]

[0186]

[0187] [Table 6]

[0188]

[0189] Referring to Tables 5 and 6, a value of pps_weighted_pred_flag equal to 0 may indicate that weighted prediction is not applied to a P or B slice referencing a PPS, and a value of pps_weighted_pred_flag equal to 1 may indicate that weighted prediction is applied to a P or B slice referencing a PPS.

[0190] In addition, the picture header may include the syntax elements shown in Table 7 below, and the semantics of the syntax elements may be as shown in Table 8 below.

[0191] [Table 7]

[0192]

[0193] [Table 8]

[0194]

[0195] Referring to Tables 7 and 8, a value of pic_weighted_pred_flag equal to 0 may indicate that weighted prediction is not applied to a P or B slice referencing a picture header. A value of pic_weighted_pred_flag equal to 1 may indicate that weighted prediction is applied to a P or B slice referencing a picture header.

[0196] When the value of pic_weighted_pred_flag is 1, all slices associated with the picture header in the picture may have the same reference picture list. Otherwise, when the value of pic_weighted_pred_flag is 1, the value of pic_rpl_present_flag may be 1.

[0197] In the absence of the above conditions, pic_weighted_pred_flag may be signaled as shown in Table 9 below.

[0198] [Table 9]

[0199]

[0200] In addition, the slice header may include the syntax elements shown in Table 10 below.

[0201] [Table 10]

[0202]

[0203] In addition, the prediction weighting table syntax may include the syntax elements shown in Table 11 below, and the semantics of the syntax elements may be as shown in Table 12 below.

[0204] [Table 11]

[0205]

[0206] [Table 12]

[0207]

[0208] Referring to Table 11 and Table 12, num_10_weighted_ref_pics may indicate the number of weighted reference pictures in reference picture list 0. The value of num_10_weighted_ref_pics is within the range from 0 to MaxDecPicBuffMinus1+14, included.

[0209] num_11_weighted_ref_pics may indicate the number of weighted reference pictures in reference picture list 1.

[0210] The value of num_11_weighted_ref_pics is within the range from 0 to MaxDecPicBuffMinus1+14, inclusive.

[0211] A value of luma_weight_10_flag[i] equal to 1 indicates the presence of a weighting factor for the luma component predicted using List 0 (L0) of RefPicList[0][i].

[0212] A value of chroma_weight_10_flag[i] equal to 1 indicates that weighting factors for the chroma prediction values ​​using L0 prediction of RefPicList[0][i] are present. A value of chroma_weight_10_flag[i] equal to 0 indicates that these weighting factors are not present.

[0213] A value of luma_weight_11_flag[i] equal to 1 indicates the presence of a weighting factor for the luma component predicted using List 1 (L1) of RefPicList[0][i].

[0214] chroma_weight_11_flag[i] indicates the presence of weighting factors for chroma prediction values ​​using L1 prediction of RefPicList[0][i]. A value of chroma_weight_10_flag[i] equal to 0 indicates that these weighting factors are not present.

[0215] For example, when weighted prediction is applied to the current block, the encoding device may generate quantity information about weighted reference pictures in the reference picture list of the current block based on the weighted prediction. The quantity information may refer to quantity information about the weights signaled for the items (reference pictures) in the L0 reference picture list and / or the L1 reference picture list. That is, the value of the quantity information may be equal to the number of weighted reference pictures in the reference picture list (L0 and / or L1). Therefore, when the value of the quantity information is n, the prediction weighting table syntax may include n weighting factor-related flags for the reference picture list. The weighting factor-related flags may correspond to luma_weight_l0_flag, luma_weight_l1_flag, chroma_weight_l0_flag, and / or chroma_weight_l0_flag of Table 11. The weight of the current picture may be derived based on the weighting factor-related flags.

[0216] When weighted bi-prediction is applied to the current block, the prediction weighting table syntax may independently include information about the number of weighted reference pictures in the L1 reference picture list and information about the number of weighted reference pictures in the L0 reference picture list, as shown in Table 11. A weighting factor-related flag may be independently included for each of the information about the number of weighted reference pictures in the L1 reference picture list and the information about the number of weighted reference pictures in the L0 reference picture list. That is, the prediction weighting table syntax may include the same number of luma_weight_l0_flags and / or chroma_weight_l0_flags as the number of weighted reference pictures in the L0 reference picture list, and may include the same number of luma_weight_l1_flags and / or chroma_weight_l1_flags as the number of weighted reference pictures in the L1 reference picture list.

[0217] The encoding device can encode image information including quantity information and weighting factor related flags, and can output the encoded image information in the form of a bitstream. Here, the quantity information and weighting factor related flags can be included in the prediction weighting table syntax in the image information as shown in Table 11. The prediction weighting table syntax can be included in the picture header in the image information or in the slice header of the image information. In order to indicate whether the prediction weighting table syntax is included in the picture header, that is, to indicate whether the information about weighted prediction is present in the picture header, the weighted prediction related flag can be included in the picture parameter set and / or the picture header. When the weighted prediction related flag is included in the picture parameter set, the weighted prediction related flag can correspond to pps_weighted_pred_flag of Table 5. When the weighted prediction related flag is included in the picture header, the weighted prediction related flag can correspond to pic_weighted_pred_flag of Table 7. Alternatively, both pps_weighted_pred_flag and pic_weighted_pred_flag may be included in the picture information to indicate whether the prediction weighting table syntax is included in the picture header.

[0218] When a weighted prediction related flag is parsed from a bitstream, the decoding device may parse the prediction weighting table syntax from the bitstream based on the parsed flag. The weighted prediction related flag may be parsed from a picture parameter set and / or a picture header of the bitstream. In other words, the weighted prediction related flag may include pps_weighted_pred_flag and / or pic_weighted_pred_flag. When the value of pps_weighted_pred_flag and / or pic_weighted_pred_flag is 1, the decoding device may parse the prediction weighting table syntax from the picture header of the bitstream.

[0219] When the prediction weighting table syntax is parsed from a picture header (when the value of pps_weighted_pred_flag and / or pic_weighted_pred_flag is 1), the decoding device may apply the information about weighted prediction included in the prediction weighting table syntax to all slices in the current picture. In other words, when the prediction weighting table syntax is parsed from a picture header, all slices in the picture associated with the picture header may have the same reference picture list.

[0220] The decoding device may parse information about the number of weighted reference pictures in the reference picture list for the current block based on the prediction weighting table syntax. The value of this information may be equal to the number of weighted reference pictures in the reference picture list. When weighted bi-prediction is applied to the current block, the decoding device may independently parse information about the number of weighted reference pictures in the L1 reference picture list and the number of weighted reference pictures in the L0 reference picture list from the prediction weighting table syntax.

[0221] The decoding device may parse the weighting factor related flags of the reference picture list from the prediction weighting table syntax based on the quantity information. The weighting factor related flags may correspond to luma_weight_l0_flag, luma_weight_l1_flag, chroma_weight_l0_flag and / or chroma_weight_l0_flag of Table 11. For example, when the value of the quantity information is n, the decoding device may parse n weighting factor related flags from the prediction weighting table syntax. The decoding device may derive the weight of the reference picture of the current block based on the weighting factor related flags, and may perform weighted prediction on the current block based on the weights, thereby generating or deriving prediction samples. Subsequently, the decoding device may generate or derive reconstructed samples of the current block based on the prediction samples, and may reconstruct the current picture based on the reconstructed samples.

[0222] In another embodiment, for weighted prediction in inter-frame prediction, the picture header may include syntax elements shown in Table 13 below, and the semantics of the syntax elements may be as shown in Table 14 below.

[0223] [Table 13]

[0224]

[0225] [Table 14]

[0226]

[0227]

[0228] Referring to Tables 13 and 14, a value of pic_weighted_pred_flag equal to 0 may indicate that weighted prediction is not applied to the P or B slice referenced by the picture header. A value of pic_weighted_pred_flag equal to 1 may indicate that weighted prediction is applied to the P or B slice referenced by the picture header. When the value of sps_weighted_pred_flag is 0, the value of pic_weighted_pred_flag is 0.

[0229] The slice header may include the syntax elements shown in Table 15 below.

[0230] [Table 15]

[0231]

[0232]

[0233] Referring to Table 15, the weighted prediction related flag (pic_weighted_pred_flag) can indicate whether the prediction weighting table syntax (information about weighted prediction) is present in the picture header or the slice header. A value of pic_weighted_pred_flag equal to 1 can indicate that the prediction weighting table syntax (information about weighted prediction) may be present in the picture header instead of the slice header. A value of pic_weighted_pred_flag equal to 0 can indicate that the prediction weighting table syntax (information about weighted prediction) may be present in the slice header instead of the picture header. Although Tables 13 and 14 indicate that the weighted prediction related flag is signaled in the picture header, the weighted prediction related flag can also be signaled in the picture parameter set.

[0234] For example, when applying weighted prediction to the current block, the encoding device performs weighted prediction and may encode image information including a weighted prediction-related flag and a prediction weighting table syntax based on the weighted prediction. Here, when the prediction weighting table syntax is included in a picture header of the image information, the encoding device may determine that the flag has a value of 1, and when the prediction weighting table syntax is included in a slice header of the image information, the encoding device may determine that the flag has a value of 0. When the flag has a value of 1, the information regarding weighted prediction included in the prediction weighting table syntax may be applied to all slices in the current picture. When the flag has a value of 0, the information regarding weighted prediction included in the prediction weighting table syntax may be applied to slices associated with the slice header among the slices in the current picture. Therefore, when the prediction weighting table syntax is included in the picture header, all slices in the picture associated with the picture header may have the same reference picture list, and when the prediction weighting table syntax is included in the slice header, the slices associated with the slice header may have the same reference picture list.

[0235] The prediction weighting table syntax may include information about the number of weighted reference pictures in the reference picture list of the current block, weighting factor-related flags, and the like. As described above, the quantity information may refer to information about the number of weights signaled for entries (reference pictures) in the L0 reference picture list and / or the L1 reference picture list, and the value of the quantity information may be equal to the number of weighted reference pictures in the reference picture list (L0 and / or L1). Therefore, when the value of the quantity information is n, the prediction weighting table syntax may include n weighting factor-related flags for the reference picture list. The weighting factor-related flags may correspond to luma_weight_l0_flag, luma_weight_l1_flag, chroma_weight_l0_flag, and / or chroma_weight_l0_flag of Table 11.

[0236] When weighted bi-prediction is applied to the current block, the encoding device may generate a prediction weighting table syntax including information about the number of weighted reference pictures in the L1 reference picture list and information about the number of weighted reference pictures in the L0 reference picture list. The prediction weighting table syntax may independently include a weighting factor-related flag for each of the information about the number of weighted reference pictures in the L1 reference picture list and the information about the number of weighted reference pictures in the L0 reference picture list. That is, the prediction weighting table syntax may include the same number of luma_weight_l0_flags and / or chroma_weight_l0_flags as the number of weighted reference pictures in the L0 reference picture list, and may include the same number of luma_weight_l1_flags and / or chroma_weight_l1_flags as the number of weighted reference pictures in the L1 reference picture list.

[0237] When the weighted prediction related flag is parsed from the bitstream, the decoding device can parse the prediction weighting table syntax from the bitstream based on the parsed flag. The weighted prediction related flag can be parsed from the picture parameter set and / or picture header of the bitstream. In other words, the weighted prediction related flag can correspond to pps_weighted_pred_flag and / or pic_weighted_pred_flag. When the value of the weighted prediction related flag is 1, the decoding device can parse the prediction weighting table syntax from the picture header of the bitstream. When the value of the weighted prediction related flag is 0, the decoding device can parse the prediction weighting table syntax from the slice header of the bitstream.

[0238] When the prediction weighting table syntax is parsed from a picture header, the decoding device may apply the information regarding weighted prediction included in the prediction weighting table syntax to all slices in the current picture. In other words, when the prediction weighting table syntax is parsed from a picture header, all slices in the picture associated with the picture header may have the same reference picture list. When the prediction weighting table syntax is parsed from a slice header, the decoding device may apply the information regarding weighted prediction included in the prediction weighting table syntax to slices associated with the slice header among the slices in the current picture. In other words, when the prediction weighting table syntax is parsed from a picture header, slices associated with the slice header may have the same reference picture list.

[0239] The decoding device may parse information about the number of weighted reference pictures in the reference picture list for the current block based on the prediction weighting table syntax. The value of this information may be equal to the number of weighted reference pictures in the reference picture list. When weighted bi-prediction is applied to the current block, the decoding device may independently parse information about the number of weighted reference pictures in the L1 reference picture list and information about the number of weighted reference pictures in the L0 reference picture list from the prediction weighting table syntax.

[0240] The decoding device can parse the weighting factor related flag of the reference picture list from the prediction weighting table syntax based on the quantity information. The weighting factor related flag can correspond to the above-mentioned luma_weight_l0_flag, luma_weight_l1_flag, chroma_weight_l0_flag and / or chroma_weight_l0_flag. For example, when the value of the quantity information is n, the decoding device can parse n weighting factor related flags from the prediction weighting table syntax. The decoding device can derive the weight of the reference picture of the current block based on the weighting factor related flag, and can perform inter-frame prediction on the current block based on the weight, thereby generating or deriving prediction samples. The decoding device can generate or derive reconstructed samples of the current block based on the prediction samples, and can generate a reconstructed picture of the current picture based on the reconstructed samples.

[0241] In yet another embodiment, the prediction weighting table syntax may include the syntax elements shown in Table 16 below, and the semantics of the syntax elements may be as shown in Table 17 below.

[0242] [Table 16]

[0243]

[0244] [Table 17]

[0245]

[0246]

[0247] In Tables 16 and 17, when pic_poc_delta_sign[i] is not present, pic_poc_delta_sign[i] is inferred to be 0. DeltaPocWeightedRefPic[i] may be derived as follows, where i is in the range from 0 to num_weighted_ref_pics_minus1, inclusive.

[0248] [Formula 1]

[0249] DeltaPocWeightedRefPic[i]=pic_poc_abs_delta[i]*(1-2*pic_poc_delta_sign[i])

[0250] Chromaweight[i][j] can be derived as (1<<Chromalog2WeightDenom)+delta_chroma_weight[i][j]. When the value of chroma_weight_flag[i] is 1, the value of delta_chroma_weight[i][j] is included in the range from -128 to 127. When the value of chroma_weight_flag[i] is 0, Chromaweight[i][j] can be derived as 2Chromalog2WeightDenom.

[0251] ChromaOffset[i][j] can be derived as follows.

[0252] [Formula 2]

[0253] ChromaOffset[i][j]=Clip3(-128,127,(128+delta chroma offset[i][j]-((128*ChromaWeight[i][j])>>ChromaLog2WeightDenom)))

[0254] The value of delta_chroma_offset[i][j] may be within the range of -4*128 to 4*127. When the value of chroma_weight_flag[i] is 0, the value of ChromaOffset[i][j] is inferred to be 9.

[0255] sumWeightflags can be derived as the sum of luma_weight_flag[i]+2*chroma_weight_flag[i]. i is in the range from 0 to num_weighted_ref_pics_minus1, inclusive. When slice_type is P, sumWeightL0Flags is less than or equal to 24.

[0256] When the current slice is a P slice or a B slice and the value of pic_weighted_pred_flag is 1, L0ToWeightedRefIdx[i] may represent the mapping between the index in the weighted reference picture list and the i-th reference picture L0. i is included in the range from 0 to NumRefIdxActive[0]-1 and may be derived as follows.

[0257] [Formula 3]

[0258]

[0259] When the current slice is a B slice and the value of pic_weighted_pred_flag is 1, L1ToWeightedRefIdx[i] may represent the mapping between the index in the weighted reference picture list and the i-th active reference picture L1. i is included in the range from 0 to NumRefIdxActive[1]-1 and may be derived as follows.

[0260] [Formula 4]

[0261]

[0262] When luma_weight_l0_flag[i] appears, luma_weight_l0_flag[i] is replaced by luma_weight_flag[L0ToWeightedRefIdx[i]], and when luma_weight_l1_flag[i] appears, luma_weight_l1_flag[i] is replaced by luma_weight_flag[L1ToWeightedRefIdx[i]].

[0263] When LumaWeightL0[i] appears, LumaWeightL0[i] is replaced by LumaWeight[L0ToWeightedRefIdx[i]], and when LumaWeightL1[i] appears, LumaWeightL1[i] is replaced by LumaWeight[L1ToWeightedRefIdx[i]].

[0264] When luma_offset_l0[i] appears, luma_offset_l0[i] is replaced by luma_offset[L0ToWeightedRefIdx[i]], and when luma_offset_l1[i] appears, luma_offset_l1[i] is replaced by luma_offset[L1ToWeightedRefIdx[i]].

[0265] When ChromaWeightL0[i] occurs, ChromaWeightL0[i] is replaced by ChromaWeight[L0ToWeightedRefIdx[i]], and when ChromaWeightL1[i] occurs, ChromaWeightL1[i] is replaced by ChromaWeight[L1ToWeightedRefIdx[i]].

[0266] In yet another embodiment, the slice header syntax may include the syntax elements shown in Table 18 below, and the semantics of the syntax elements may be as shown in Table 19 below.

[0267] [Table 18]

[0268]

[0269] [Table 19]

[0270]

[0271] Referring to Table 18 and Table 19, a flag indicating whether the prediction weight table syntax is present in the slice header may be signaled. This flag may be signaled in the slice header and may be referred to as slice_weight_pred_present_flag.

[0272] A value of slice_weight_pred_present_flag equal to 1 may indicate that the prediction weight table syntax is present in the slice header. A value of slice_weight_pred_present_flag equal to 0 may indicate that the prediction weight table syntax is not present in the slice header. That is, slice_weight_pred_present_flag equal to 0 may indicate that the prediction weight table syntax is present in the picture header.

[0273] In yet another embodiment, the prediction weight table syntax is parsed from the slice header, but an adaptation parameter set including the syntax elements shown in Table 20 below may be signaled.

[0274] [Table 20]

[0275]

[0276] Each APS RBSP needs to be available to the decoding process before being included for use as a reference in at least one access unit whose TemporalId is less than or equal to the TemporalId of the coded slice NAL unit that references the APS RBSP or is provided by external methods.

[0277] aspLayerId may be referred to as the nuh_layer_id of the APS NAL unit. When the layer with nuh_layer_id equal to aspLayerId is an independent layer (i.e., when vps_independent_layer_flag[GeneralLayerIdx[aspLayerId]] is 1), the APS NAL unit including the APS RBSP has the same nuh_layer_id as the nuh_layer_id of the coded slice NAL that references the APS RBSP. Otherwise, the APS NAL unit including the APS RBSP has the same nuh_layer_id as the nuh_layer_id of the coded slice NAL unit that references the APS RBSP or the nuh_layer_id of a directly dependent layer of the layer including the coded slice NAL unit that references the APS RBSP.

[0278] All APS NAL units with a specific value of adaptation_parameter_set_id and a specific value of aps_params_type in an access unit have the same content.

[0279] adaptation_parameter_set_id provides an identifier of the APS so that other syntax elements can refer to the identifier.

[0280] When aps_params_type is ALF_APS, SCALING_APS, or PRED_WEIGHT_APS, the value of adaptation_parameter_set_id is in the range from 0 to 7, inclusive.

[0281] When aps_params_type is LMCS_APS, the value of adaptation_parameter_set_id is within the range from 0 to 3 inclusive.

[0282] aps_params_type indicates the type of APS parameters included in the APS, as shown in the following Table 21. When the value of aps_params_type is 1 (LMCS_APS), the value of adaptation_parameter_set_id is included in the range from 0 to 3.

[0283] [Table 21]

[0284] aps_params_type The name of aps_params_type Types of APS parameters 0 ALF_APS ALF parameters 1 LMCS_APS LMCS parameters 2 SCALING_APS Scaling List Parameters 3 PRED_WEIGHT_APS Prediction weighting parameters 4..7 reserve reserve

[0285] Each type of APS uses a separate value space for adaptation_parameter_set_id.

[0286] APS NAL units (with a specific value of adaptation_parameter_set_id and a specific value of aps_params_type) can be shared between pictures, and different slices in a picture can reference different ALF APSs.

[0287] A value of aps_extension_flag equal to 0 indicates that the aps_extension_data_flag syntax element is not present in the APS RBSP syntax structure. A value of aps_extension_flag equal to 1 indicates that the aps_extension_data_flag syntax element is present in the APS RBSP syntax structure.

[0288] aps_extension_data_flag may have a random value.

[0289] As described above, a new aps_params_type (PRED_WEIGHT_APS) can be added to the existing types. In addition, the slice header can be modified to signal the APS ID instead of pred_weight_table(), as shown in Table 22 below.

[0290] [Table 22]

[0291]

[0292] In Table 22, slice_pred_weight_aps_id indicates adaptation_parameter_set_id of prediction weight table APS. TemporalId of the APS NAL unit having the same aps_params_type as PERD_WEIGHT_APS and the same adaptation_parameter_set_id as slice_pred_weight_aps_id is less than or equal to TemporalId of the coded slice NAL unit.

[0293] When the slice_pred_weight_aps_id syntax element is present in the slice header, the value of slice_pred_weight_aps_id is the same for all slices of the picture.

[0294] In this case, the prediction weighting table syntax shown in Table 23 below can be signaled.

[0295] [Table 23]

[0296]

[0297]

[0298] In Table 23, a value of num_lists_active_flag equal to 1 may indicate that prediction weight table information is signaled for one reference picture list. A value of num_lists_active_flag equal to 0 may indicate that prediction weight table information of two reference picture lists L0 and L1 is not signaled.

[0299] numRefIdxActive[i] may be used to indicate the number of active reference indices. The value of numRefIdxActive[i] ranges from 0 to 14.

[0300] The syntax of Table 23 indicates whether information on one or two lists is parsed in the APS when num_lists_active_flag is parsed.

[0301] Instead of Table 23, the prediction weight table syntax shown in Table 24 below may be used.

[0302] [Table 24]

[0303]

[0304] In Table 24, a value of num_lists_active_flag equal to 1 may indicate that prediction weighting table information is signaled for one reference picture list, and a value of num_lists_active_flag equal to 0 may indicate that prediction weighting table information is not signaled for two reference picture lists.

[0305] Figure 8 and Figure 9 An example of a video / image encoding method and related components according to an embodiment of this document is schematically illustrated.

[0306] Figure 8 The video / image encoding method disclosed in Figure 2 and Figure 9 Specifically, for example, Figure 8 S800 to S820 may be performed by the predictor 220 of the encoding apparatus 200 , and S830 may be performed by the entropy encoder 240 of the encoding apparatus 200 . Figure 8 The video / image encoding method disclosed in may include the above-mentioned embodiments of this document.

[0307] Specifically, refer to Figure 8 and Figure 9 , the predictor 220 of the encoding device can derive motion information about the current block in the current picture based on motion estimation (S800). For example, the encoding device can use the original block in the original picture to search for a similar reference block with high correlation with the current block in units of fractional pixels within a predetermined search range in the reference picture, and thus can derive motion information. The similarity of the blocks can be derived based on the difference between the sample values ​​based on the stage. For example, the similarity of the blocks can be calculated based on the sum of the absolute differences (SAD) between the current block (or the template of the current block) and the reference block (or the template of the reference block). In this case, the motion information can be derived based on the reference block with the smallest SAD in the search area. According to various methods, based on the inter-frame prediction mode, the derived motion information can be notified to the decoding device with a signal.

[0308] The predictor 220 of the encoding device can perform weighted (sample) prediction on the current block based on the motion information about the current block (S810), and can generate prediction samples (prediction blocks) and prediction-related information for the current block. The prediction-related information may include prediction mode information (merge mode, skip mode, etc.), information about motion information, etc. The information about motion information may include candidate selection information (e.g., merge index, MVP flag, or MVP index), which is information used to derive the motion vector. In addition, the information about motion information may include information about the above-mentioned MVD and / or reference picture index information. In addition, the information about motion information may include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. For example, when the slice type of the current slice is a P slice or a B slice, the predictor 220 may perform weighted prediction on the current block in the current slice. Weighted prediction can be used not only when bi-prediction is applied to the current block, but also when uni-prediction is applied to the current block.

[0309] The predictor 220 of the encoding device may generate information regarding the number of weighted reference pictures in a reference picture list for weighted prediction based on weighted prediction according to motion information (S820). In this case, the entropy encoder 240 of the encoding device may encode the image information including the information regarding the number. The information regarding the number may be included in a prediction weighting table syntax in the image information. Even in this case, the prediction weighting table syntax may be included in a picture header in the image information. Here, the value of the number information may be the same as the number of weighted reference pictures in the reference picture list. The prediction weighting table syntax may include as many weighting factor-related flags as the value of the number information. For example, when the value of the number information is n, the prediction weighting table syntax may include n weighting factor-related flags. When weighted bi-prediction is applied, the prediction weighting table syntax may include the number information and / or weighting factor-related flags independently for each of L0 and L1. In other words, the information regarding the number of weighted reference pictures in L0 and the information regarding the number of weighted reference pictures in L1 may be signaled independently in the prediction weighting table syntax without relying on each other (independent of the number of active reference pictures in each list).

[0310] The residual processor 230 of the encoding device can generate residual samples and residual information based on the predicted samples generated by the predictor 220 and the original picture (original block and original sample). Here, the residual information is information about the residual samples and may include information about (quantized) transform coefficients for the residual samples.

[0311] The adder (or reconstructor) of the encoding apparatus may generate a reconstructed sample (a reconstructed picture, a reconstructed block, or a reconstructed sample array) by adding the residual sample generated by the residual processor 230 and the prediction sample generated by the predictor 220 .

[0312] The entropy encoder 240 of the encoding apparatus may encode image information (S830), the image information including prediction-related information generated by the predictor 220, residual information generated by the residual processor 230, a flag related to weighted prediction, a prediction weighting table syntax, etc. The prediction-related information may include a flag related to weighted prediction, a prediction weighting table syntax, etc., and the prediction weighting table syntax may include information on the number of weighted reference pictures, a weighting factor-related flag, etc.

[0313] For example, the entropy encoder 240 of the encoding device may encode the image information based on at least one of Tables 5 to 23, and may output the encoded image information in the form of a bitstream. Specifically, the entropy encoder 240 of the encoding device may determine the value of a flag related to weighted prediction to be 1 based on the prediction weighting table syntax of this document being included in the picture header of the image information, and may determine the value of the flag related to weighted prediction to be 0 based on the prediction weighting table syntax being included in the slice header of the image information. Alternatively, the entropy encoder 240 of the encoding device may determine the value of the flag related to weighted prediction to be 1 based on the information about weighted prediction included in the prediction weighting table syntax being applied to all slices in the current picture including the current block, and may determine the value of the flag related to weighted prediction to be 0 based on the information about weighted prediction included in the prediction weighting table syntax being applied to slices associated with the slice header among the slices in the current picture. When the prediction weighting table syntax is included in a picture header, all slices in a picture associated with the picture header can have the same reference picture list, and when the prediction weighting table syntax is included in a slice header, all slices associated with the slice header can have the same reference picture list. A flag related to weighted prediction can be included in a picture parameter set or a picture header of image information and transmitted to a decoding device. The flag related to weighted prediction can be information indicating whether information regarding weighted prediction is present in a picture header.

[0314] Figure 10 and Figure 11 An example of a video / image decoding method and related components according to an embodiment of this document is schematically illustrated.

[0315] Figure 10 The video / image decoding method disclosed in Figure 3 and Figure 11 Specifically, for example, Figure 10 S1000 and S1010 may be performed by the entropy decoder 310 of the decoding device, Figure 10S1020 and S1020 may be performed by the predictor 330 of the decoding device, and S1030 may be performed by the adder 340 of the decoding device. Figure 10 The video / image decoding method disclosed in may include the above-mentioned embodiments of this document.

[0316] refer to Figure 10 and Figure 11 , the entropy decoder 310 of the decoding device can parse a flag related to weighted prediction from the bitstream, and can parse a prediction weighting table syntax from the bitstream based on the flag related to weighted prediction (S1000). The flag related to weighted prediction can be parsed from a picture parameter set or a picture header of the bitstream, and can indicate whether information about weighted prediction (prediction weighting table syntax) exists in the picture header. For example, when the value of the flag related to weighted prediction is 1, the entropy decoder 310 of the decoding device can parse the prediction weighting table syntax from the picture header of the bitstream, and when the value of the flag related to weighted prediction is 0, the entropy decoder 310 of the decoding device can parse the prediction weighting table syntax from the slice header of the bitstream. When the value of the flag related to weighted prediction is 1, the information about weighted prediction included in the prediction weighting table syntax can be applied to all slices in the current picture, and when the value of the flag related to weighted prediction is 0, the information about weighted prediction included in the prediction weighting table syntax can be applied to slices associated with the slice header among the slices in the current picture. When the prediction weighting table syntax is parsed from the picture header, all slices in the picture associated with the picture header can have the same reference picture list, and when the prediction weighting table syntax is parsed from the slice header, the slices associated with the slice header can have the same reference picture list.

[0317] The entropy decoder 310 of the decoding device can parse quantity information from the prediction weighting table syntax. The value of the quantity information can be the same as the number of weighted reference pictures in the reference picture list. The entropy decoder 310 of the decoding device can parse as many weighting factor-related flags as the value of the quantity information. For example, when the value of the quantity information is n, the prediction weighting table syntax may include n weighting factor-related flags. When weighted bi-prediction is applied, the quantity information and / or weighting factor-related flags can be independently included in the prediction weighting table syntax for each of L0 and L1. In one example, the quantity information about the weighted reference pictures in L0 and the quantity information about the weighted reference pictures in L1 can be independently parsed in the prediction weighting table syntax without relying on each other (not relying on the number of active reference pictures in each list).

[0318] The decoding device can perform weighted prediction on the current block in the current picture based on the prediction-related information (inter / intra prediction classification information, intra prediction mode information, inter prediction mode information, information about weighted prediction, etc.) obtained from the bitstream, thereby reconstructing the current picture. Here, the information about weighted prediction may include a prediction weighting table syntax. For example, the predictor 330 of the decoding device can derive the weight of the weighted prediction based on the weighting factor related flag parsed according to the quantity information in the prediction weighting table syntax (S1020). Specifically, the value of the quantity information in the prediction weighting table syntax is n, and the predictor 330 of the decoding device can parse n weighting factor related flags from the prediction weighting table syntax. The predictor 330 of the decoding device can perform weighted prediction on the current block based on the weight, thereby deriving the predicted sample of the current block (S1030).

[0319] The residual processor 320 of the decoding device may generate residual samples based on the residual information obtained from the bitstream. The adder 340 of the decoding device may generate reconstructed samples based on the predicted samples generated by the predictor 330 and the residual samples generated by the residual processor 320. The adder 340 of the decoding device may generate a reconstructed picture (reconstructed block) based on the reconstructed samples (S1040).

[0320] Subsequently, if necessary, an in-loop filtering process (such as deblocking filtering, SAO and / or ALF) may be applied to the reconstructed picture to improve the subjective / objective picture quality.

[0321] Although the methods have been described in the above embodiments based on flowcharts in which steps or blocks are listed in sequence, the steps of this document are not limited to a specific order, and a step may be performed in a different step, in a different order, or simultaneously relative to the above order. In addition, it should be understood by those skilled in the art that the steps in the flowchart are not exclusive, and another step may be included or one or more steps in the flowchart may be deleted without affecting the scope of this document.

[0322] The above-mentioned method according to the present document may be in the form of software, and the encoding device and / or decoding device according to the present document may be included in an apparatus for performing image processing (e.g., a TV, a computer, a smart phone, a set-top box, a display device, etc.).

[0323] When the embodiments of this document are implemented with software, the above-mentioned methods can be implemented with modules (processing or functions) that perform the above-mentioned functions. The modules can be stored in a memory and executed by a processor. The memory can be installed inside or outside the processor and can be connected to the processor via various well-known devices. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits and / or data processing devices. The memory may include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium and / or other storage devices. In other words, according to the embodiments of this document, it can be implemented and executed on a processor, a microprocessor, a controller or a chip. For example, the functional units illustrated in the corresponding figures can be implemented and executed on a computer, a processor, a microprocessor, a controller or a chip. In this case, information about the implementation (e.g., information about instructions) or the algorithm can be stored in a digital storage medium.

[0324] In addition, the decoding device and the encoding device to which the embodiments of this document are applied may be included in a multimedia broadcast transceiver, a mobile communication terminal, a home theater video device, a digital theater video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a portable camera, a video on demand (VoD) service provider, an over-the-top (OTT) video device, an Internet streaming service provider, a 3D video device, a virtual reality (VR) device, an augmented reality (AR) device, an image phone video device, an in-vehicle terminal (e.g., an in-vehicle (including an autonomous vehicle) terminal, an airplane terminal, or a ship terminal), and a medical video device; and may be used to process image signals or data. For example, an OTT video device may include a game console, a Blueray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, and a digital video recorder (DVR).

[0325] In addition, the processing method of the embodiment of the present document can be generated in the form of a program executed by a computer and can be stored in a computer-readable recording medium. The multimedia data with a data structure according to the embodiment of the present document can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all kinds of storage devices and distributed storage devices storing computer-readable data. The computer-readable recording medium may include, for example, a Blu-ray disc (BD), a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. The computer-readable recording medium also includes a medium implemented in the form of a carrier wave (e.g., transmission on the Internet). In addition, the bit stream generated by the encoding method can be stored in a computer-readable recording medium or can be transmitted through a wired or wireless communication network.

[0326] In addition, the embodiments of this document can be implemented as a computer program product based on a program code, and the program code can be executed on a computer according to the embodiments of this document.The program code can be stored on a computer-readable carrier.

[0327] Figure 12 represents an example of a content streaming system to which embodiments of this document may be applied.

[0328] refer to Figure 12 A content streaming system to which embodiments of this document are applied may generally include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0329] The encoding server is used to compress content input from a multimedia input device such as a smartphone, camera, or camcorder into digital data, generate a bitstream, and transmit it to the streaming server. In another example, if the multimedia input device such as a smartphone, camera, or camcorder directly generates a bitstream, the encoding server can be omitted.

[0330] The bitstream may be generated by an encoding method or a bitstream generation method to which the embodiments of this document are applied, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.

[0331] The streaming server transmits multimedia data to the user device via a network server based on the user's request. The network server serves as a tool for notifying the user of available services. When the user requests a desired service, the network server transfers the request to the streaming server, which then transmits the multimedia data to the user. In this regard, the content streaming system may include a separate control server, in which case the control server is used to control commands and responses between the various devices in the content streaming system.

[0332] The streaming server may receive content from a media storage and / or encoding server. For example, when receiving content from an encoding server, the content may be received in real time. In this case, the streaming server may store the bitstream for a predetermined period of time to provide a smooth streaming service.

[0333] For example, user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation, tablet PCs, tablet PCs, ultrabooks, wearable devices (e.g., watch-type terminals (smart watches), glasses-type terminals (smart glasses), head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc.

[0334] Each server in the content streaming system may be operated as a distributed server, and in such case, data received by each server may be processed in a distributed manner.

Claims

1. A decoding device for image decoding, the decoding device comprising: Memory; as well as at least one processor coupled to the memory, the at least one processor configured to: Parse the L0 quantity information about the weighted reference pictures in reference picture list 0 from the prediction weight table syntax; Parsing L1 quantity information about weighted reference pictures in reference picture list 1 from the prediction weighting table syntax; deriving a weight for weighted prediction based on the L0 quantity information and the L1 quantity information; deriving a prediction sample for the current block by performing the weighted prediction on the current block based on the weight; and Reconstruct the current picture based on the predicted samples, wherein the prediction weighting table syntax is included in a picture header of a bitstream, and Wherein, the L1 quantity information is different from the L0 quantity information, wherein the L0 quantity information and the L1 quantity information are present in the prediction weighting table syntax included in the picture header, The weights are derived based on the following operations: Parsing L0 flag information related to the luma weighting factor for the reference picture list 0 from the prediction weighting table syntax based on the L0 quantity information, Parsing the L0 flag information related to the chroma weighting factor for the reference picture list 0 from the prediction weighting table syntax, Parsing incremental luminance weighting factor related L0 information from the prediction weighting table based on the luminance weighting factor related L0 flag information, and The parsing order of the L0 flag information related to the chroma weighting factor is after the parsing order of the L0 flag information related to the luminance weighting factor and before the parsing order of the L0 information related to the incremental luminance weighting factor.

2. A coding device for image coding, the coding device comprising: Memory; as well as at least one processor coupled to the memory, the at least one processor configured to: Derive motion information about the current block; performing weighted prediction on the current block based on the motion information; generating L0 quantity information about weighted reference pictures in reference picture list 0 and L1 quantity information about weighted reference pictures in reference picture list 1; and encoding the image information including the L0 quantity information and the L1 quantity information, The L0 quantity information and the L1 quantity information are included in the prediction weight table syntax in the image information. wherein the prediction weighting table syntax is included in a picture header of the image information, and Wherein, the L1 quantity information is different from the L0 quantity information, wherein the L0 quantity information and the L1 quantity information are present in the prediction weighting table syntax included in the picture header, The prediction weighting table syntax includes: L0 flag information related to the luminance weighting factor for the reference picture list 0 based on the L0 quantity information, Chroma weighting factor related L0 flag information for the reference picture list 0, and Incremental luma weighting factor-related L0 information based on the luma weighting factor-related L0 flag information, and wherein the parsing order of the chroma weighting factor-related L0 flag information is after the parsing order of the luma weighting factor-related L0 flag information and before the parsing order of the incremental luma weighting factor-related L0 information.

3. A device for transmitting data for an image, the device comprising: At least one processor configured to obtain a bitstream for the image, wherein the bitstream is generated based on the following operations: Derive motion information about the current block, performing weighted prediction on the current block based on the motion information, Generate L0 quantity information about weighted reference pictures in reference picture list 0 and L1 quantity information about weighted reference pictures in reference picture list 1, and encoding the image information including the L0 quantity information and the L1 quantity information; and a transmitter configured to transmit the data comprising the bit stream, wherein the L0 quantity information and the L1 quantity information are included in the prediction weighting table syntax in the image information, and wherein the prediction weighting table syntax is included in the picture header of the image information, Wherein, the L1 quantity information is different from the L0 quantity information, wherein the L0 quantity information and the L1 quantity information are present in the prediction weighting table syntax included in the picture header, The prediction weighting table syntax includes: L0 flag information related to the luminance weighting factor for the reference picture list 0 based on the L0 quantity information, Chroma weighting factor related L0 flag information for the reference picture list 0, and Incremental luma weighting factor-related L0 information based on the luma weighting factor-related L0 flag information, and wherein the parsing order of the chroma weighting factor-related L0 flag information is after the parsing order of the luma weighting factor-related L0 flag information and before the parsing order of the incremental luma weighting factor-related L0 information.