Image encoding / decoding method and apparatus, and recording medium in which bitstream is stored
By using multiple prediction modes to perform inter prediction in the image decoding method and equipment, the efficiency problem of the multi-reference block mode in image compression in the prior art is solved, and higher prediction accuracy and compression efficiency are achieved.
Patent Information
- Application Number
- CN202380071596.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-12
- Filing Date
- 2023-10-12
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to effectively use the multi-reference block mode for inter prediction in image compression, resulting in low prediction accuracy and compression efficiency.
By performing inter prediction based on the multi-prediction mode in the image decoding method and the device, the basic prediction block and additional reference block of the current block are generated and a weighted sum thereof is calculated to obtain the final prediction block.
Improve the accuracy of inter-frame prediction, reduce signaling overhead, and improve compression efficiency.
Smart Images

Figure CN119948863A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image encoding / decoding method and apparatus and a recording medium storing a bit stream. Background Art
[0002] Recently, demands for high-resolution and high-quality images such as HD (High Definition) images and UHD (Ultra High Definition) images have been increasing in various application fields, and therefore, efficient image compression technology is being discussed.
[0003] There are various technologies, such as inter-frame prediction technology that uses video compression technology to predict pixel values included in the current picture from pictures before or after the current picture, intra-frame prediction technology that predicts pixel values included in the current picture by using pixel information in the current picture, entropy coding technology that assigns short symbols to values with high occurrence frequency and assigns long symbols to values with low occurrence frequency, etc., and these image compression technologies can be used to effectively compress image data and transmit or store it. Summary of the invention
[0004] Technical issues
[0005] The present disclosure provides a method and apparatus for performing inter-frame prediction based on a multiple reference block mode.
[0006] The present disclosure provides a signaling method and device for determining a multiple reference block mode.
[0007] The present disclosure provides a method and apparatus for configuring a motion information candidate list for a multiple reference block mode.
[0008] Technical Solution
[0009] According to the image decoding method and device of the present disclosure, prediction can be performed based on the first prediction mode to generate a basic prediction block of the current block, an additional reference block of the current block can be derived based on the second prediction mode, and a weighted sum of the basic prediction block and the additional reference block can be calculated to generate a final prediction block of the current block.
[0010] The image decoding method and apparatus according to the present disclosure may acquire a flag indicating whether an additional reference block is used for a current block.
[0011] In the image decoding method and apparatus according to the present disclosure, the first prediction mode may include at least one of a merge mode, a skip mode, an AMVP mode, an intra block copy (IBC) mode, or an AMVP-merge combined mode.
[0012] In the image decoding method and apparatus according to the present disclosure, the second prediction mode may include at least one of a merge mode, a skip mode, an AMVP mode, an IBC mode, or an AMVP-merge combined mode.
[0013] In the image decoding method and apparatus according to the present disclosure, the first prediction mode and the second prediction mode may be determined as a specific combination within a predefined prediction mode combination set.
[0014] In the image decoding method and apparatus according to the present disclosure, the predefined prediction mode combination set may include a plurality of combination candidates.
[0015] In the image decoding method and device according to the present disclosure, multiple combination candidates can be configured by including at least two of merge mode, skip mode, AMVP mode, IBC mode, AMVP-merge combination mode, geometric partition mode (GPM), combined inter-intra prediction (CIIP) mode, sub-block merge mode or affine mode.
[0016] The image decoding method and apparatus according to the present disclosure may perform bidirectional prediction to derive a first reference block and a second reference block of a current block, and calculate a weighted sum of the first reference block and the second reference block to generate a basic prediction block.
[0017] The image decoding method and apparatus according to the present disclosure may perform unidirectional prediction to derive a third reference block of a current block and generate a basic prediction block.
[0018] In the image decoding method and apparatus according to the present disclosure, when a plurality of additional reference blocks are derived, a final prediction block may be generated by sequentially weighting and summing the plurality of additional reference blocks to a basic prediction block.
[0019] In the image decoding method and apparatus according to the present disclosure, the information about the second prediction mode may include at least one of weight information or prediction information.
[0020] In the image decoding method and apparatus according to the present disclosure, the weight information may represent information indicating a weight used for a weighted sum of the additional reference block, and the prediction information may represent information used to derive the additional reference block.
[0021] According to the image encoding method and device of the present disclosure, unidirectional or bidirectional prediction can be performed based on the first prediction mode to generate a basic prediction block of the current block, an additional reference block of the current block is derived based on the second prediction mode, and a weighted sum of the basic prediction block and the additional reference block is calculated to generate a final prediction block of the current block.
[0022] A computer-readable digital storage medium is provided, which stores encoded video / image information, resulting in execution of an image decoding method due to a decoding device according to the present disclosure.
[0023] A computer-readable digital storage medium storing video / image information generated according to an image encoding method according to the present disclosure is provided.
[0024] A method and apparatus for transmitting video / image information generated according to an image encoding method according to the present disclosure are provided.
[0025] Beneficial Effects
[0026] The present disclosure may improve the accuracy of prediction by performing inter-frame prediction based on multiple hypothesis prediction modes.
[0027] The present disclosure may reduce signaling overhead and increase compression efficiency by effectively defining a prediction mode determination method for multiple hypothesis prediction modes.
[0028] The present disclosure may improve compression efficiency by effectively configuring a motion information candidate list for a multi-hypothesis prediction mode. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 A video / image compilation system according to the present disclosure is shown.
[0030] Figure 2 A rough block diagram of an encoding device to which an embodiment of the present disclosure can be applied and which performs encoding of a video / image signal is shown.
[0031] Figure 3 A rough block diagram of a decoding device to which an embodiment of the present disclosure can be applied and which performs decoding of a video / image signal is shown.
[0032] Figure 4 An example of a video / image encoding method based on inter-frame prediction to which an embodiment of the present disclosure can be applied is shown.
[0033] Figure 5 An example of a video / image decoding method based on inter-frame prediction to which an embodiment of the present disclosure can be applied is shown.
[0034] Figure 6 An inter-frame prediction process to which the embodiments of the present disclosure can be applied is illustratively shown.
[0035] Figure 7 An inter-frame prediction method performed by the decoding device 300 according to an embodiment of the present disclosure is shown.
[0036] Figure 8 is a diagram illustrating reference blocks used for a multiple reference block mode according to an embodiment of the present disclosure.
[0037] Fig. 9 is a diagram illustrating a method for signaling information about additional reference blocks used in a multiple reference block mode according to an embodiment of the present disclosure.
[0038] Fig.10is a flowchart illustrating a syntax parsing structure to which an embodiment of the present disclosure can be applied.
[0039] Fig.11 is a flowchart illustrating a syntax parsing structure and a prediction mode determination method according to an embodiment of the present disclosure.
[0040] Fig.12 is a flowchart illustrating a prediction mode determination method according to an embodiment of the present disclosure.
[0041] Fig.13 is a flowchart illustrating a prediction mode determination method according to an embodiment of the present disclosure.
[0042] Fig.14 is a flowchart illustrating a prediction method based on multiple reference blocks according to an embodiment of the present disclosure.
[0043] Fig.15 is a flowchart illustrating a prediction method based on multiple reference blocks according to an embodiment of the present disclosure.
[0044] Fig.16 is a diagram for describing a candidate list configuration method according to an embodiment of the present disclosure.
[0045] Fig.17 A rough configuration of the inter-frame predictor 332 that performs the inter-frame prediction method according to the present disclosure is shown.
[0046] Fig.18 An inter-frame prediction method performed by the encoding apparatus 200 as an embodiment according to the present disclosure is shown.
[0047] Fig.19 A rough configuration of the inter-frame predictor 221 performing the inter-frame prediction method according to the present disclosure is shown.
[0048] Fig. 20 An example of a content streaming system to which an embodiment of the present disclosure can be applied is shown. DETAILED DESCRIPTION
[0049] Because the present disclosure can make various changes and has several embodiments, specific embodiments will be illustrated in the drawings and described in detail in the detailed description. However, it is not intended to limit the present disclosure to specific embodiments, and it should be understood to include all changes, equivalents and substitutes included in the spirit and technical scope of the present disclosure. While describing each of the accompanying drawings, similar reference numerals are used for similar components.
[0050] Terms such as first, second, etc. can be used to describe various components, but components should not be limited by these terms. These terms are only used to distinguish one component from other components. For example, without departing from the scope of the present disclosure, a first component can be referred to as a second component, and similarly, a second component can also be referred to as a first component. Terms and / or combinations of any one or more related statement items in a plurality of related statement items are included.
[0051] When a component is referred to as being "connected" or "linked" to another component, it should be understood that it can be directly connected or linked to another component, but another component may also exist in between. On the other hand, when a component is referred to as being "directly connected" or "directly linked" to another component, it should be understood that another component does not exist in between.
[0052] The terms used in this application are only used to describe specific embodiments and are not intended to limit the present disclosure. Unless the context clearly indicates otherwise, singular expressions include plural expressions. In this application, it should be understood that terms such as "including" or "having" are intended to designate the existence of features, numbers, steps, operations, components, parts or combinations thereof described in the specification, but do not exclude the possibility of the existence or addition of one or more other features, numbers, steps, operations, components, parts or combinations thereof in advance.
[0053] The present disclosure relates to video / image coding. For example, the methods / embodiments disclosed herein may be applied to methods disclosed in the Universal Video Coding (VVC) standard. In addition, the methods / embodiments disclosed herein may be applied to methods disclosed in the Basic Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation audio video coding standard (AVS2), or the next generation video / image coding standard (e.g., H.267 or H.268, etc.).
[0054] This specification proposes various embodiments of video / image coding, and unless otherwise stated, these embodiments may be performed in combination with each other.
[0055] Here, video may refer to a collection of a series of images over time. A picture generally refers to a unit representing an image within a specific time period, and a slice / tile is a unit that forms a part of a picture in coding. A slice / tile may include at least one coding tree unit (CTU). A picture may consist of at least one slice / tile. A tile is a rectangular area consisting of multiple CTUs within a specific tile column and a specific tile row of a picture. A tile column is a rectangular area of a CTU having the same height as the picture and a width assigned by the syntax requirements of the picture parameter set. A tile row is a rectangular area of a CTU having a height assigned by the picture parameter set and the same width as the width of the picture. The CTU within a tile may be arranged continuously according to a CTU raster scan, and the tiles within a picture may be arranged continuously according to a raster scan of the tile. A slice may include an integer number of complete tiles or an integer number of continuous complete CTU rows within a tile of a picture that may be exclusively included in a single NAL unit. At the same time, a picture may be divided into at least two sub-pictures. A sub-picture may be a rectangular area of at least one slice within a picture.
[0056] Pixel, pixel or picture element may refer to the smallest unit constituting a picture (or image). In addition, "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luminance component, or only a pixel / pixel value of a chrominance component.
[0057] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information associated with the corresponding region. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, a unit may be used interchangeably with terms such as a block or region. In general, an MxN block may include a set (or array) of transform coefficients or samples (or sample arrays) consisting of M columns and N rows.
[0058] Here, "A or B" may refer to "only A", "only B", or "both A and B". In other words, herein, "A or B" may be interpreted as "A and / or B". For example, herein, "A, B or C" may refer to "only A", "only B", "only C", or "any combination of A, B, and C".
[0059] As used herein, a slash ( / ) or a comma may mean "and / or". For example, "A / B" may mean "A and / or B". Thus, "A / B" may mean "only A", "only B", or "both A and B". For example, "A, B, C" may mean "A, B, or C".
[0060] Here, "at least one of A and B" may refer to "only A", "only B", or "both A and B". In addition, herein, expressions such as "at least one of A or B" or "at least one of A and / or B" can be interpreted in the same manner as "at least one of A and B".
[0061] In addition, herein, “at least one of A, B, and C” may refer to “only A”, “only B”, “only C”, or “any combination of A, B, and C”. In addition, “at least one of A, B, or C” or “at least one of A, B and / or C” may refer to “at least one of A, B, and C”.
[0062] In addition, the brackets used herein may refer to "for example". Specifically, when the indication is "prediction (intra-frame prediction)", "intra-frame prediction" may be proposed as an example of "prediction". In other words, the "prediction" here is not limited to "intra-frame prediction", and "intra-frame prediction" may be proposed as an example of "prediction". In addition, even when the indication is "prediction (ie, intra-frame prediction)", "intra-frame prediction" may be proposed as an example of "prediction".
[0063] Here, technical features described individually in one drawing may be implemented individually or simultaneously.
[0064] Figure 1 A video / image compilation system according to the present disclosure is shown.
[0065] refer to Figure 1 , a video / image coding system may include a first device (source device) and a second device (receiving device).
[0066] The source device may send the encoded video / image information or data to the receiving device in the form of a file or stream transmission through a digital storage medium or a network. The source device may include a video source, an encoding device, and a sending unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be composed of a separate device or an external component.
[0067] The video source may obtain the video / image through the process of capturing, synthesizing or generating the video / image. The video source may include a device for capturing the video / image and a device for generating the video / image. The device for capturing the video / image may include at least one camera, a video / image archive including previously captured videos / images, etc. The device for generating the video / image may include a computer, a tablet computer, a smart phone, etc., and may (electronically) generate the video / image. For example, a virtual video / image may be generated by a computer, etc., and in this case, the process of capturing the video / image may be replaced by the process of generating the relevant data.
[0068] The encoding device can encode the input video / image. The encoding device can perform a series of processes such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0069] The sending unit may send the encoded video / image information or data output in the form of a bit stream to the receiving unit of the receiving device in the form of a file or stream transmission through a digital storage medium or a network. The digital storage medium may include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The sending unit may include an element for generating a media file in a predetermined file format and may include an element for transmission through a broadcast / communication network. The receiving unit may receive / extract a bit stream and send it to a decoding device.
[0070] The decoding device may decode the video / image by performing a series of processes corresponding to the operations of the encoding device, such as dequantization, inverse transformation, prediction, etc.
[0071] The renderer may render the decoded video / image. The rendered video / image may be displayed through a display unit.
[0072] Figure 2 A rough block diagram of an encoding device to which an embodiment of the present disclosure can be applied and which performs encoding of a video / image signal is shown.
[0073] refer to Figure 2, the encoding device 200 may be composed of an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the above-mentioned image segmenter 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or processor). In addition, the memory 270 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include a memory 270 as an internal / external component.
[0074] The image divider 210 may divide the input image (or picture, frame) input to the encoding device 200 into at least one processing unit. As an example, the processing unit may be referred to as a coding unit (CU). In this case, the coding unit may be recursively divided from a coding tree unit (CTU) or a maximum coding unit (LCU) according to a quadtree binary tree ternary tree (QTBTTT) structure.
[0075] For example, one coding unit may be segmented into a plurality of coding units having a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quadtree structure may be applied first, and the binary tree structure and / or the ternary structure may be applied later. Alternatively, the binary tree structure may be applied before the quadtree structure. The coding process according to this specification may be performed based on a final coding unit that is no longer segmented. In this case, based on coding efficiency according to image characteristics, etc., the maximum coding unit may be directly used as the final coding unit, or if necessary, the coding unit may be recursively segmented into coding units of a deeper depth, and the coding unit with the best size may be used as the final coding unit. Here, the coding process may include processes such as prediction, transformation, and reconstruction described later.
[0076] As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be divided or partitioned from the above-mentioned final coding unit, respectively. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from the transform coefficient.
[0077] In some cases, a unit may be used interchangeably with terms such as a block or region. In general, an MxN block may represent a set of transform coefficients or samples consisting of M columns and N rows. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luma component, or only a pixel / pixel value of a chroma component. A sample may be used as a term to make a picture (or image) correspond to a pixel or a picture element.
[0078] The encoding device 200 may subtract the prediction signal (prediction block, prediction sample array) output from the inter predictor 221 or the intra predictor 222 from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual sample array), and the generated residual signal is transmitted to the transformer 232. In this case, a unit that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) within the encoding device 200 may be referred to as a subtractor 231.
[0079] The predictor 220 may perform prediction on a block to be processed (hereinafter referred to as a current block) and generate a predicted block including a prediction sample for the current block. The predictor 220 may determine whether intra prediction or inter prediction is applied in units of a current block or CU. The predictor 220 may generate various information about the prediction, such as prediction mode information, and send it to the entropy encoder 240, as described later in the description of each prediction mode. The information about the prediction may be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0080] The intra-frame predictor 222 can predict the current block by referring to the samples in the current picture. Depending on the prediction mode, the referenced sample can be located near the current block or can be located at a certain distance away from the current block. In intra-frame prediction, the prediction mode may include at least one non-directional mode and multiple directional modes. The non-directional mode may include at least one of the DC mode or the plane mode. Depending on the detail level of the prediction direction, the directional mode may include 33 directional modes or 65 directional modes. However, this is only an example, and more or less directional modes may be used depending on the configuration. The intra-frame predictor 222 may determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0081] The inter-frame predictor 221 may derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information sent in the inter-frame prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter-frame prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be referred to as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block may be referred to as a collocated picture (colPic). For example, the inter-frame predictor 221 may configure a motion information candidate list based on the neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, for skip mode and merge mode, the inter predictor 221 may use motion information of neighboring blocks as motion information of the current block. For skip mode, unlike merge mode, a residual signal may not be transmitted. For motion vector prediction (MVP) mode, motion vectors of surrounding blocks are used as motion vector predictors, and motion vector differences are signaled to indicate the motion vector of the current block.
[0082] The predictor 220 may generate a prediction signal based on various prediction methods described later. For example, the predictor may not only apply intra prediction or inter prediction to predict a block, but may also apply intra prediction and inter prediction at the same time. It may be referred to as a combined inter and intra prediction (CIIP) mode. In addition, the predictor may be based on an intra block copy (IBC) prediction mode or may be based on a palette mode for prediction of a block. The IBC prediction mode or the palette mode may be used for content image / video coding of games, such as screen content coding (SCC), etc. IBC basically performs prediction within the current picture, but it may be performed similarly to inter prediction because it derives a reference block within the current picture. In other words, IBC may use at least one of the inter prediction techniques described herein. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, the sample values within the picture may be signaled based on information about the palette table and the palette index. The prediction signal generated by the predictor 220 may be used to generate a reconstructed signal or a residual signal.
[0083] The transformer 232 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, GBT refers to a transform obtained from a graph when relationship information between pixels is expressed as a graph. CNT refers to a transform obtained based on generating a prediction signal using all previously reconstructed pixels. In addition, the transform process may be applied to square pixel blocks of the same size or may be applied to non-square blocks of variable size.
[0084] The quantizer 233 may quantize the transform coefficients and send them to the entropy encoder 240, and the entropy encoder 240 may encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantizer 233 may rearrange the quantized transform coefficients in the block form into a one-dimensional vector form based on the coefficient scanning order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0085] The entropy encoder 240 may perform various encoding methods such as Exponential Golomb, Context Adaptive Variable Length Coding (CAVLC), Context Adaptive Binary Arithmetic Coding (CABAC), etc. The entropy encoder 240 may encode information necessary for video / video image reconstruction (e.g., values of syntax elements, etc.) in addition to transform coefficients quantized together or individually.
[0086] The encoded information (e.g., encoded video / image information) can be transmitted or stored in units of network abstraction layer (NAL) units in the form of a bitstream. The video / image information may further include information about various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. Here, information and / or syntax elements transmitted / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information may be encoded and included in the bitstream through the above-mentioned encoding process. The bitstream may be transmitted through a network or may be stored in a digital storage medium. Here, the network may include a broadcast network and / or a communication network, etc., and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) for transmission and / or a storage unit (not shown) for storing a signal output from the entropy encoder 240 may be configured as an internal / external element of the encoding device 200, or the transmission unit may also be included in the entropy encoder 240.
[0087] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients through the dequantizer 234 and the inverse transformer 235. The adder 250 can add the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). When there is no residual of the block to be processed, such as when the skip mode is applied, the prediction block can be used as a reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture, and can also be used for inter-frame prediction of the next picture through filtering described later. At the same time, luminance mapping with chroma scaling (LMCS) can be applied in the picture encoding and / or reconstruction process.
[0088] The filter 260 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and the modified reconstructed picture can be stored in the memory 270, specifically in the DPB of the memory 270. Various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 260 can generate various information about filtering and send it to the entropy encoder 240. The information about filtering can be encoded in the entropy encoder 240 and output in the form of a bit stream.
[0089] The modified reconstructed picture transmitted to the memory 270 may be used as a reference picture in the inter predictor 221. When inter prediction is applied therethrough, the encoding apparatus can avoid prediction mismatch in the encoding apparatus 200 and the decoding apparatus, and can also improve encoding efficiency.
[0090] The DPB of the memory 270 may store the modified reconstructed picture to be used as a reference picture in the inter-frame predictor 221. The memory 270 may store the motion information of the block from which the motion information in the current picture is derived (or encoded) and / or the motion information of the block in the pre-reconstructed picture. The stored motion information may be sent to the inter-frame predictor 221 to be used as the motion information of the spatial neighboring block or the motion information of the temporal neighboring block. The memory 270 may store the reconstructed samples of the reconstructed block in the current picture and send them to the intra-frame predictor 222.
[0091] Figure 3 A rough block diagram of a decoding device to which an embodiment of the present disclosure can be applied and which performs decoding of a video / image signal is shown.
[0092] refer to Figure 3 , the decoding device 300 may be configured by including an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 331 and an intra-frame predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321.
[0093] According to an embodiment, the above-mentioned entropy decoder 310, residual processor 320, predictor 330, adder 340 and filter 350 may be configured by one hardware component (e.g., decoder chipset or processor). In addition, the memory 360 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory 360 as an internal / external component.
[0094] When a bit stream including video / image information is input, the decoding device 300 may respond to the Figure 2 The image is reconstructed by a process of processing video / image information in an encoding device. For example, the decoding device 300 can derive a unit / block based on relevant information of block segmentation obtained from a bitstream. The decoding device 300 can perform decoding by using a processing unit applied in the encoding device. Therefore, the processing unit of decoding can be a coding unit, and the coding unit can be divided from a coding tree unit or a maximum coding unit according to a quadtree structure, a binary tree structure and / or a ternary tree structure. At least one transform unit can be derived from the coding unit. And, the reconstructed image signal decoded and output by the decoding device 300 can be played by a playback device.
[0095] The decoding device 300 may receive the bit stream from Figure 2 The received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information (e.g., video / image information) necessary for image reconstruction (or picture reconstruction). The video / image information may further include information about various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The decoding device may further decode the picture based on the information about the parameter set and / or the general constraint information. The signaled / received information and / or the syntax elements described later in this document may be decoded and obtained from the bitstream by a decoding process. For example, the entropy decoder 310 may decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, CABAC, etc., and output the values of the syntax elements necessary for image reconstruction and the quantized values of the transform coefficients of the residual. In more detail, the CABAC entropy decoding method may receive a bin corresponding to each syntax element from a bitstream, determine a context model by using information of the syntax element to be decoded, decoding information of surrounding blocks and blocks to be decoded, or information of a symbol / bin decoded in a previous step, perform arithmetic decoding on the bin by predicting the probability of occurrence of the bin according to the determined context model, and generate a symbol corresponding to the value of each syntax element. In this case, after determining the context model, the CABAC entropy decoding method may update the context model by using information about the decoded symbol / bin of the context model for the next symbol / bin. Among the information decoded in the entropy decoder 310, information about the prediction is provided to the predictor (inter-frame predictor 332 and intra-frame predictor 331), and the residual value on which entropy decoding is performed in the entropy decoder 310, that is, the quantized transform coefficient and related parameter information may be input to the residual processor 320. The residual processor 320 may derive a residual signal (residual block, residual sample, residual sample array). In addition, information about filtering among the information decoded in the entropy decoder 310 may be provided to the filter 350. Meanwhile, a receiving unit (not shown) receiving a signal output from the encoding device may be further configured as an internal / external element of the decoding device 300 or the receiving unit may be a component of the entropy decoder 310.
[0096] Meanwhile, the decoding device according to this specification may be referred to as a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0097] The dequantizer 321 may dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 321 may rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding device. The dequantizer 321 may dequantize the quantized transform coefficients by using a quantization parameter (e.g., quantization step size information) and obtain the transform coefficients.
[0098] The inverse transformer 322 inversely transforms the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0099] The predictor 320 may perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor 320 may determine whether to apply intra prediction or inter prediction to the current block based on the information on prediction output from the entropy decoder 310, and determine a specific intra / inter prediction mode.
[0100] The predictor 320 can generate a prediction signal based on various prediction methods described later. For example, the predictor 320 can not only apply intra prediction or inter prediction to predict a block, but also apply intra prediction and inter prediction at the same time. It can be called a combined inter and intra prediction (CIIP) mode. In addition, the predictor can be based on an intra block copy (IBC) prediction mode or can be based on a palette mode for prediction of a block. The IBC prediction mode or the palette mode can be used for content image / video coding of games, such as screen content coding (SCC), etc. IBC basically performs prediction within the current picture, but it can be performed similarly to inter prediction because it derives a reference block within the current picture. In other words, IBC can use at least one of the inter prediction techniques described herein. The palette mode can be considered as an example of intra coding or intra prediction. When the palette mode is applied, information about the palette table and the palette index can be included in the video / image information and sent with a signal.
[0101] The intra-frame predictor 331 can predict the current block by referring to samples within the current picture. Depending on the prediction mode, the referenced sample can be located near the current block or can be located at a certain distance away from the current block. In intra-frame prediction, the prediction mode may include at least one non-directional mode and multiple directional modes. The intra-frame predictor 331 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring block.
[0102] The inter-frame predictor 332 may derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information sent in the inter-frame prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter-frame prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter-frame predictor 332 may configure a motion information candidate list based on the neighboring blocks, and derive a motion vector and / or a reference picture index for the current block based on the received candidate selection information. Inter-frame prediction may be performed based on various prediction modes, and information about the prediction may include information indicating an inter-frame prediction mode for the current block.
[0103] The adder 340 may add the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including the inter-frame predictor 332 and / or the intra-frame predictor 331) to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). When there is no residual of the block to be processed, such as when the skip mode is applied, the prediction block may be used as the reconstructed block.
[0104] The adder 340 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next block to be processed in the current picture, may be output through filtering described later, or may be used for inter prediction of the next picture. At the same time, luminance mapping with chroma scaling (LMCS) may be applied during picture decoding.
[0105] The filter 350 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and send the modified reconstructed picture to the memory 360, specifically the DPB of the memory 360. The various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0106] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter-frame prediction unit 332. The memory 360 can be derived from the motion information in its current picture (or decoded) The motion information of the block and / or the motion information of the block in the pre-reconstructed picture. The stored motion information can be sent to the inter-frame predictor 260 to be used as the motion information of the spatial neighboring block or the motion information of the temporal neighboring block. The memory 360 can store the reconstructed samples of the reconstructed block in the current picture and send them to the intra-frame predictor 331.
[0107] Here, the embodiments described in the filter 260, the inter-frame predictor 221, and the intra-frame predictor 222 of the encoding device 200 may also be equally or correspondingly applied to the filter 350, the inter-frame predictor 332, and the intra-frame predictor 331 of the decoding device 300, respectively.
[0108] Meanwhile, when inter prediction is applied, the predictor of the encoding device / decoding device can perform inter prediction in units of blocks to derive prediction samples. Inter prediction may refer to prediction derived in a manner that depends on data elements (e.g., sample values or motion information) of pictures other than the current picture. When inter prediction is applied to the current block, a prediction block (prediction sample array) for the current block may be derived based on a reference block (reference sample array) specified by a motion vector on a reference picture indicated by a reference picture index.
[0109] In this case, in order to reduce the amount of motion information sent in the inter-frame prediction mode, the motion information of the current block can be predicted in units of blocks, sub-blocks or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information may include a motion vector and / or a reference picture index. The motion information may further include information about the inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). When inter-frame prediction is applied, the neighboring blocks may include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture.
[0110] The reference picture including the reference block may be the same as or different from the reference picture including the temporally neighboring block. The temporally neighboring block may be referred to as a co-located reference block, a co-located CU (colCU), etc., and the reference picture including the temporally neighboring block may be referred to as a co-located picture (colPic). For example, a motion information candidate list may be configured based on the neighboring blocks of the current block, and a flag or index information indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block may be signaled.
[0111] Inter-frame prediction may be performed based on various prediction modes, and for example, for skip mode and merge mode, the motion information of the current block may be the same as the motion information of the selected neighboring block. For skip mode, unlike merge mode, a residual signal may not be sent. For motion vector prediction (MVP) mode, the motion vector of the selected neighboring block may be used as a motion vector predictor, and the motion vector difference may be signaled. In this case, the motion vector of the current block may be derived by using the sum of the motion vector predictor and the motion vector difference.
[0112] Depending on the inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.), the motion information may include L0 motion information and / or L1 motion information. The motion vector in the L0 direction may be referred to as an L0 motion vector or MVL0, and the motion vector in the L1 direction may be referred to as an L1 motion vector or MVL1. Prediction based on the L0 motion vector may be referred to as an L0 prediction, prediction based on the L1 motion vector may be referred to as an L1 prediction, and prediction based on both the L0 motion vector and the L1 motion vector may be referred to as a Bi prediction. Here, the L0 motion vector may represent a motion vector associated with the reference picture list L0 (L0), and the L1 motion vector may represent a motion vector associated with the reference picture list L1 (L1). The reference picture list L0 may include a picture before the current picture as a reference picture in output order, and the reference picture list L1 may include a picture after the current picture in output order. The previous picture may be referred to as a forward (reference) picture, and the subsequent picture may be referred to as a backward (reference) picture.
[0113] The reference picture list L0 may further include a picture after the current picture as a reference picture in the output order. In this case, within the reference picture list L0, the previous picture may be indexed first, and the subsequent picture may be indexed next. The reference picture list L1 may further include a picture before the current picture as a reference picture in the output order. In this case, within the reference picture list 1, the subsequent picture may be indexed first, and the previous picture may be indexed next. Here, the output order may correspond to a picture order count (POC) order.
[0114] For example, the video / image encoding process based on inter-frame prediction may generally include the following.
[0115] Figure 4 An example of a video / image encoding method based on inter-frame prediction to which an embodiment of the present disclosure can be applied is shown.
[0116] The encoding device may perform inter prediction on the current block S400. The encoding device may derive the inter prediction mode and motion information of the current block and generate prediction samples of the current block. Here, the processes of determining the inter prediction mode, deriving the motion information, and generating the prediction samples may be performed simultaneously, or any one of the processes may be performed before the other processes. For example, the inter predictor of the encoding device may include a prediction mode determination unit, a motion information deriving unit, and a prediction sample deriving unit, the prediction mode determination unit may determine the prediction mode for the current block, the motion information deriving unit may derive the motion information of the current block, and the prediction sample deriving unit may derive the prediction samples of the current block.
[0117] For example, the inter-frame predictor of the encoding device can search for a block similar to the current block in a specific area (search area) of the reference picture through motion estimation, and derive a reference block whose difference with the current block is the smallest or less than or equal to a specific standard. Based on this, a reference picture index indicating the reference picture where the reference block is located can be derived, and a motion vector can be derived based on the position difference between the reference block and the current block. The encoding device can determine the mode applied to the current block among various prediction modes. The encoding device can compare the RD costs for various prediction modes and determine the best prediction mode for the current block.
[0118] For example, when a skip mode or a merge mode is applied to the current block, the encoding device may configure a merge candidate list described below, and derive a reference block whose difference with the current block is the smallest or less than or equal to a specific standard among the reference blocks indicated by the merge candidates included in the merge candidate list. In this case, a merge candidate associated with the derived reference block may be selected, and merge index information indicating the selected merge candidate may be generated and sent to the decoding device by a signal. The motion information of the current block may be derived by using the motion information of the selected merge candidate.
[0119] As another example, when the (A) MVP mode is applied to the current block, the encoding device may configure the (A) MVP candidate list described below, and use the motion vector of the motion vector predictor (MVP) candidate selected from the motion vector predictor candidates included in the (A) MVP candidate list as the motion vector predictor of the current block. In this case, for example, the motion vector indicating the reference block derived by the above-mentioned motion estimation may be used as the motion vector of the current block, and the motion vector predictor candidate having the smallest difference between the motion vector of the current block and the motion vector predictor candidate among the motion vector predictor candidates may become the selected motion vector predictor candidate. A motion vector difference (MVD) may be derived, which is the difference obtained by subtracting the motion vector predictor from the motion vector of the current block. In this case, information about the MVD may be signaled to the decoding device. In addition, when the (A) MVP mode is applied, the value of the reference picture index may be configured as reference picture index information and signaled separately to the decoding device.
[0120] The encoding apparatus may derive residual samples based on the prediction samples S410. The encoding apparatus may derive residual samples by comparing original samples of the current block with the prediction samples.
[0121] The encoding device may encode the image information including prediction information and residual information S420. The encoding device can output the encoded image information in the form of a bitstream. The prediction information is information related to the prediction process, and may include prediction mode information (e.g., a skip flag, a merge flag, or a merge index, etc.) and / or motion information. The motion information may include candidate selection information (e.g., a merge index, an mvp flag, or an mvp index), which is information for deriving a motion vector. In addition, the motion information may include information about the above-mentioned MVD and / or reference picture index information. In addition, the motion information may include information indicating whether L0 prediction, L1 prediction, or dual prediction is applied. The residual information is information about the residual sample. The residual information may include information about the quantized transform coefficients used for the residual sample.
[0122] The output bitstream may be stored in a (digital) storage medium and sent to a decoding device, or may be sent to a decoding device over a network.
[0123] At the same time, as described above, the encoding device can generate a reconstructed picture (including a reconstructed sample and a reconstructed block) based on the reference sample and the residual sample. To derive the same prediction result as that performed in the decoding device in the encoding device, the coding efficiency can be improved. Therefore, the encoding device can store the reconstructed picture (or reconstructed sample, reconstructed block) in a memory and use it as a reference picture for inter-frame prediction. As described above, a loop filtering process, etc. can be further applied to the reconstructed picture.
[0124] For example, the video / image decoding process based on inter-frame prediction may generally include the following.
[0125] Figure 5 An example of a video / image decoding method based on inter-frame prediction to which an embodiment of the present disclosure can be applied is shown.
[0126] refer to Figure 5 The decoding device may perform operations corresponding to the operations performed in the encoding device. The decoding device may perform prediction on the current block based on the received prediction information and derive a prediction sample.
[0127] Specifically, the decoding device may determine a prediction mode for the current block based on the received prediction information S500. The decoding device may determine which inter prediction mode is applied to the current block based on prediction mode information in the prediction information.
[0128] For example, whether to apply the merge mode to the current block or whether to determine the (A) MVP mode may be determined based on the merge flag. Alternatively, one of various inter prediction mode candidates may be selected based on the mode index. The inter prediction mode candidate may include a skip mode, a merge mode, and / or an (A) MVP mode, or may include various inter prediction modes described below.
[0129] The decoding device may derive motion information of the current block based on the determined inter prediction mode S510. For example, when the skip mode or merge mode is applied to the current block, the decoding device may configure a merge candidate list described below and select one of the merge candidates included in the merge candidate list. The selection may be performed based on the above-mentioned selection information (merge index). The motion information of the current block may be derived by using the motion information of the selected merge candidate. The motion information of the selected merge candidate may be used as the motion information of the current block.
[0130] As another example, when the (A) MVP mode is applied to the current block, the decoding device may configure the (A) MVP candidate list described below, and use the motion vector of the motion vector predictor (MVP) candidate selected from among the motion vector predictor (MVP) candidates included in the (A) MVP candidate list as the MVP of the current block. The selection may be performed based on the selection information (MVP flag or MVP index) described above. In this case, the MVD of the current block may be derived based on the information about the MVD, and the motion vector of the current block may be derived based on the MVP and MVD of the current block. In addition, the reference picture index of the current block may be derived based on the reference picture index information. The picture indicated by the reference picture index in the reference picture list for the current block may be derived as a reference picture referenced for inter-frame prediction of the current block.
[0131] Meanwhile, as described below, the motion information of the current block may be derived without configuring the candidate list, and in this case, the motion information of the current block may be derived according to the process initiated in the prediction mode described below. In this case, the candidate list configuration as described above may be omitted.
[0132] The decoding device may generate a prediction sample for the current block based on the motion information of the current block S520. In this case, the reference picture may be derived based on the reference picture index of the current block, and the prediction sample of the current block may be derived by using the sample of the reference block indicated by the motion vector of the current block on the reference picture. In this case, as described below, in some cases, a prediction sample filtering process may be further performed on all or part of the prediction samples of the current block.
[0133] For example, the inter-frame predictor of the decoding device may include a prediction mode determination unit, a motion information export unit and a prediction sample export unit. The prediction mode determination unit can determine the prediction mode for the current block based on the received prediction mode information, the motion information export unit can export the motion information (motion vector and / or reference picture index, etc.) of the current block based on the received motion information, and the prediction sample export unit can export the prediction sample of the current block.
[0134] The decoding device generates residual samples for the current block based on the received residual information S530. The decoding device can generate reconstructed samples for the current block based on the prediction samples and the residual samples, and generate a reconstructed picture based thereon S540. As described above, thereafter, an in-loop filtering process or the like can be further applied to the reconstructed picture.
[0135] Figure 6 The inter-frame prediction process to which the embodiments of the present disclosure can be applied is exemplified.
[0136] refer to Figure 6 As described above, the inter-frame prediction process may include determining an inter-frame prediction mode, deriving motion information according to the determined prediction mode, and performing prediction (generating prediction samples) based on the derived motion information. The inter-frame prediction process may be performed in the encoding device and the decoding device as described above. In this document, the encoding device may include an encoding device and / or a decoding device.
[0137] refer to Figure 6, the coding device determines the inter prediction mode for the current block S600. Various inter prediction modes can be used for the prediction of the current block in the picture. For example, various modes such as merge mode, skip mode, motion vector prediction (MVP) mode, affine mode, sub-block merge mode, merge with MVD (MMVD) mode, etc. can be used. Decoder-side motion vector refinement (DMVR) mode, adaptive motion vector resolution (AMVR) mode, dual prediction with CU level weight (BCW), bidirectional optical flow (BDOF), etc. can be used additionally or alternatively as incidental modes. In addition, according to an embodiment of the present disclosure, the above-mentioned inter prediction mode may include a multi-hypothesis prediction (MHP) mode. The multi-hypothesis prediction mode represents a method for performing prediction by weighted summing of prediction blocks generated based on additional motion information for bidirectional prediction (or bi-prediction) blocks. The multi-hypothesis prediction mode will be described in detail later.
[0138] In the present disclosure, the affine mode may be referred to as an affine motion prediction mode. In addition, the MVP mode may be referred to as an advanced motion vector prediction (AMVP) mode. In the present disclosure, some modes and / or motion information candidates derived from some modes may be included as one of the motion information related candidates of another mode. For example, an HMVP candidate may be added as a merge candidate for a merge / skip mode, or may be added as a motion vector predictor candidate for an AMVP mode. When an HMVP candidate is used as a motion information candidate for a merge mode or a skip mode, the HMVP candidate may be referred to as an HMVP merge candidate.
[0139] Prediction mode information indicating an inter prediction mode of a current block may be sent from an encoding device to a decoding device using a signal. The prediction mode information may be included in a bitstream and received in a decoding device. The prediction mode information may include index information indicating one of a plurality of candidate modes. Alternatively, the inter prediction mode may be indicated by hierarchical signaling of flag information.
[0140] In this case, the prediction mode information may include at least one flag. For example, a skip flag may be signaled to indicate whether the skip mode is applied, a merge flag may be signaled to indicate whether the merge mode is applied when the skip mode is not applied, and a flag indicating that the MVP mode is applied or used for additional partitioning may be further signaled when the merge mode is not applied. The affine mode may be signaled as an independent mode, or may be signaled as a mode dependent on the merge mode or the MVP mode, etc. For example, the affine mode may include an affine merge mode and an affine MVP mode.
[0141] The coding device may derive motion information for the current block S610. The motion information may be derived based on an inter prediction mode. The coding device may perform inter prediction by using the motion information of the current block. The coding device may derive optimal motion information for the current block through a motion estimation process.
[0142] For example, the encoding device may use the original block in the original picture for the current block to search for a similar reference block with high correlation in the fractional pixel unit within the determined search range within the reference picture, and derive motion information therefrom. The similarity of the blocks may be derived based on the difference between the sample values based on the phase. For example, the similarity of the blocks may be calculated based on the SAD between the current block (or the template of the current block) and the reference block (or the template of the reference block). In this case, the motion information may be derived based on the reference block with the minimum SAD within the search range. The derived motion information may be signaled to the decoding device according to various methods based on the inter-frame prediction mode.
[0143] The coding apparatus may perform inter prediction based on the motion information for the current block S620. The coding apparatus may generate a prediction sample for the current block based on the motion information. The current block including the prediction sample may be referred to as a prediction block.
[0144] Hereinafter, the multi-hypothesis prediction mode is described in detail. As described above, the multi-hypothesis prediction mode represents a prediction method using an additional prediction block (or predictor) for a basic prediction block. The multi-hypothesis prediction mode can be selectively used as one of the various inter-frame prediction modes mentioned above. Of course, the multi-hypothesis prediction mode according to an embodiment of the present disclosure is not limited to this name. In the present disclosure, the multi-hypothesis prediction mode may also be referred to as a multi-reference mode, a multi-reference prediction, a multi-reference prediction mode, a multi-reference block mode, a multi-hypothesis (MHP) mode, a multi-hypothesis inter-frame prediction mode, an inter-frame combined prediction mode, a combined inter-frame prediction mode, a combined prediction mode, a multi-frame prediction mode, a multi-prediction mode, an additional reference prediction mode, an additional reference mode, a multi-reference block, etc.
[0145] Figure 7 An inter-frame prediction method performed by the decoding device 300 according to an embodiment of the present disclosure is illustrated.
[0146] refer to Figure 7, the decoding device may perform unidirectional or bidirectional prediction to generate a basic prediction block (or reference block) S700. When the multi-reference block mode is applied, the decoding device may generate and combine additional prediction blocks in addition to the prediction blocks generated (or derived) by the unidirectional or bidirectional prediction. As an example, the basic prediction block may include an L0 prediction block and / or an L1 prediction block. In the present disclosure, the basic prediction block may be referred to as a basic reference block, an initial prediction block, an initial reference block, a temporary prediction block, a temporary reference block, a reference prediction block, a regular prediction block, a regular reference block, etc. In addition, as an example, the basic prediction block may refer to an L0 prediction block or an L1 prediction block, or may refer to a block obtained by weighted summation of L0 and L1 prediction blocks. In the present embodiment, the case where the basic prediction block is a block obtained by weighted summation of L0 and L1 prediction blocks is mainly described, but the present invention is not limited thereto, and before performing the weighted summation, the basic prediction block may be a reference block.
[0147] In other words, in an embodiment, the image decoding device according to the present disclosure may derive the first reference block and the second reference block of the current block by performing bidirectional prediction, and generate a basic prediction block by performing weighted summation on the first reference block and the second reference block. Alternatively, in an embodiment, the image decoding device according to the present disclosure may perform unidirectional prediction to derive the third reference block of the current block and generate a basic prediction block. In other words, the basic prediction block may be derived by performing weighted summation on a plurality of reference blocks, or may be derived by using a single reference block.
[0148] In an embodiment, the weights may be used for weighted summation of L0 and L1 prediction blocks. As weights for weighted prediction, the weights may be collectively referred to as dual prediction with CU-based weights (BCW) or CU-based weights (CW). The weights may be derived from a weight candidate list. The weight candidate list may include multiple weight candidates and may be predefined in an encoding / decoding device.
[0149] The weight candidate may be a set of weights (i.e., a first weight and a second weight) indicating a weight applied to each bidirectional prediction block, or may be a weight applied to a prediction block in either direction. When only a weight applied to a prediction block in any one direction is derived from the weight candidate list, a weight applied to a prediction block in the other direction may be derived based on the weight derived from the weight candidate list. For example, the weight applied to a prediction block in the other direction may be derived by subtracting a weight derived from the weight candidate list from a predetermined value.
[0150] In an embodiment, a weight index indicating a weight for weighted prediction of a current block may be derived in a weight candidate list. In the present disclosure, a weight index may be referred to as bcw_idx, bcw index. The weight index may be derived by a decoding device, or may be signaled from an encoding device. When derived by a decoding device, the weight index may be derived as a weight index for a specific merge candidate in a merge candidate list. As an example, a specific merge candidate may be specified by a merge index in a merge candidate list.
[0151] The decoding apparatus may derive (or generate) an additional reference block (or prediction block) based on the multiple reference block mode S710. The decoding apparatus may derive an additional reference block in addition to the basic prediction block, and combine (or weighted sum) the basic prediction block and the derived additional reference block.
[0152] As an embodiment, when the multiple reference block mode is applied, the decoding device may derive and combine up to a predefined number of additional reference blocks. In other words, the decoding device may combine (or weighted sum) less than or equal to the predefined number of additional reference blocks to the basic prediction block. As an example, the predefined number may be 2. Alternatively, as an example, the predefined number may be one of 1, 2, 3, and 4. The predefined number may be referred to as the maximum number of the multiple reference block mode.
[0153] In addition, when a plurality of additional reference blocks are combined, the plurality of additional reference blocks may be weighted and summed sequentially into a basic prediction block. For example, when up to 2 additional reference blocks are generated, the basic prediction block and the first additional reference block may be weighted and summed to generate a prediction block, and the generated prediction block and the second additional reference block may be weighted and summed to generate a final prediction block. The prediction block generated by weighted summing the basic prediction block and the first additional reference block may be referred to as an intermediate prediction block.
[0154] Alternatively, when a plurality of additional reference blocks are combined, the basic prediction block and the plurality of generated additional reference blocks may be weighted and summed at once. In other words, after the plurality of additional reference blocks are generated, weights may be applied to each of the plurality of additional reference blocks and the basic prediction block (or the L0 prediction block and the L1 prediction block) and weighted and summed at once.
[0155] In addition, as an embodiment, the decoding device may determine whether to apply the multi-reference block mode. In this case, a step of determining whether to apply the multi-reference block mode may be added before S710. As an example, whether to apply the multi-reference block mode may be explicitly signaled or may be implicitly derived (or determined) by the decoding device.
[0156] In addition, as an embodiment, whether the multi-reference block mode is applied can be sent from the encoding device to the decoding device by signal. For example, a multi-reference block mode flag indicating whether the multi-reference block mode is applied can be sent from the encoding device to the decoding device by signal. In this case, the conditions for signaling / parsing the multi-reference block mode flag can be defined in advance. The signaling / parsing conditions of the multi-reference block mode flag can be the availability conditions of the multi-reference block mode. When the signaling / parsing conditions are met, the decoding device can parse the multi-reference block mode flag from the bitstream. Alternatively, as an embodiment, whether the multi-reference block mode is applied can be derived by the decoding device based on predefined encoding information. As an example, whether the multi-reference block mode is applied can be defined in the same manner as the multi-reference block mode availability condition (or signaling / parsing condition) described below.
[0157] In addition, as an embodiment, the decoding device may obtain multi-reference block mode information (or may be referred to as multi-reference block mode prediction information) to generate an additional reference block. As an example, the multi-reference block mode information may include weight information and / or prediction information. A reference block according to the multi-reference block mode, that is, an additional reference block, may be derived based on the prediction information, and the additional reference block derived based on the weight information may be weighted and summed with a basic prediction block (or an intermediate prediction block). In addition, as an example, the multi-reference block mode information may further include a multi-reference block mode flag indicating whether the multi-reference block mode is applied.
[0158] As an example, the prediction information may include mode information used to derive the additional reference block and motion information according to the mode. The mode information may be inter-prediction mode information indicating whether it is a merge mode or an AMVP mode. For example, the mode information may be a merge flag. In other words, the additional reference block may be derived using the merge mode or the AMVP mode, and a flag syntax element indicating it may be signaled. Alternatively, the additional reference block may be derived using a predefined mode among the merge mode or the AMVP mode. Alternatively, the merge mode or the AMVP mode may be selected based on predefined encoding information.
[0159] As an example, when the merge mode is used to derive an additional reference block, the prediction information may include a merge index. The merge index may specify a merge candidate in a merge candidate list. When the AMVP mode is used to derive an additional reference block, the prediction information may include a motion vector predictor flag, a reference index, and motion vector difference information. The motion vector predictor flag may specify a candidate in a motion vector predictor candidate list.
[0160] The decoding device may generate a final prediction block by weighted summing the basic prediction block and the additional reference block S720. As described above, the number of additional reference blocks may be less than or equal to the predefined number. For example, when the number of additional reference blocks is 2, the final prediction block may be a block in which the basic prediction block and the two additional reference blocks are weighted summed. As an embodiment, weight information for the weighted sum may be sent or derived using a signal.
[0161] As described above, when a plurality of additional reference blocks are combined, the plurality of additional reference blocks may be weighted and summed sequentially to form a basic prediction block, or the basic prediction block and the plurality of generated additional reference blocks may be weighted and summed at one time.
[0162] Generally, the inter prediction process supports unidirectional prediction or bidirectional prediction, but when it includes more than one prediction block, it can be regarded as multi-hypothesis prediction. In other words, multi-reference block is a method for prediction using multiple reference blocks (or prediction blocks), and the signaling or deriving method can be considered as follows.
[0163] As an embodiment, information about the additional reference block can be signaled by using a merge index in the same or similar manner as the merge mode. Alternatively, information about the additional reference block can be signaled by using a reference index, a motion vector predictor (MVP) flag (or index), a motion vector difference, etc. in the same or similar manner as the AMVP mode. Alternatively, information about the additional reference block can be inherited from an already decoded neighboring block to derive motion information. In an embodiment, whether to signal the information about the additional reference block can be determined depending on the number of additional reference blocks used. For example, when there is one additional reference block, information about the additional reference block can be signaled, and when there are two additional reference blocks, all or part of the information about the additional reference block can be derived on the decoder side without signaling.
[0164] According to the multi-reference block mode according to an embodiment of the present disclosure, weight information and motion information of a prediction block can be signaled / derived to generate a block having different characteristics from an existing prediction block, and various reference blocks can be used for prediction to improve prediction accuracy.
[0165] Figure 8 is a diagram illustrating reference blocks used for a multiple reference block mode according to an embodiment of the present disclosure.
[0166] refer to Figure 8 , showing the case where multiple reference blocks (prediction blocks) are used for prediction when the multi-reference block mode is applied. In other words, Figure 8 , reference blocks P0 and P1 represent basic reference blocks (or regular reference blocks), and reference blocks P2 and P3 represent additional reference blocks (or additional prediction blocks).
[0167] exist Figure 8 In the example, for ease of description, a case where two reference blocks exist for each prediction direction is shown, but the present invention is not limited thereto. In other words, the number of reference blocks for each prediction direction may be changed, and the motion compensation order may also be changed.
[0168] Fig. 9 is a diagram illustrating a method for signaling information about additional reference blocks used in a multiple reference block mode according to an embodiment of the present disclosure.
[0169] When multiple reference blocks can exist and information about each additional reference block is signaled and parsed, it can be done as Fig. 9 The motion information of the added prediction blocks is signaled and parsed in the order shown in . Here, MaxNum may represent the maximum number of additional reference blocks.
[0170] refer to Fig. 9 , a loop can be performed until the number of additional reference blocks reaches a maximum number. When the number of additional reference blocks is less than or equal to the maximum number, mhp_flag can be parsed (signaled). mhp_flag represents a syntax element indicating whether multiple reference blocks are used. When multiple reference blocks are used, mhp_mrg can be parsed. mhp_mrg represents a syntax element indicating whether merge mode or AMVP mode is used to derive additional reference blocks. When merge mode is applied according to the mhp_mrg value, merge index and weight index can be signaled, and when AMVP mode is applied, reference index, motion vector predictor index, motion vector difference data and weight index can be signaled.
[0171] In other words, Fig. 9 As shown in , when the maximum number of additional reference blocks that can be generated is MaxNum, it can be determined by mhp_flag whether there is MHP-related syntax. As an example, when MaxNum is 2, mhp_flag can have the value in Table 1 below, and identify whether there is a first additional block and a second additional block according to its value.
[0172] [Table 1]
[0173]
[0174] Referring to Table 1, when mhp_flag is '1', additional reference blocks may exist. Additional reference blocks may be distinguished by mhp_mrg whether they are in MHP_MERGE mode or MHP_AMVP mode. When it is in MHP_MERGE mode, merge index and weight index may be signaled. When it is in MHP_AMVP mode, reference index, motion vector predictor index (or flag), motion vector difference data and weight index may be signaled.
[0175] The syntax name described in the present disclosure is an example, and its name may be changed. In addition, in the present disclosure, the names as modes when there are multiple reference blocks are described as MHP_MERGE and MHP_AMVP, which can be distinguished from the merge mode and AMVP mode representing the motion information of the conventional reference block. Specifically, the motion vector predictor candidate list for the conventional reference block and the motion vector predictor candidate list for the additional reference block can be configured independently. In addition, the MHP_MERGE and MHP_AMVP modes for the additional reference block may include weight index information.
[0176] In addition, as an embodiment, in addition to the method for applying multiple reference blocks by the above-mentioned signaling, multiple reference blocks by the derivation method can also be applied. Specifically, because the merge mode uses motion information inherited from the decoded adjacent / non-adjacent blocks, the information used for the multiple reference blocks can also use the motion information inherited from the adjacent / non-adjacent blocks. As an example, when the current block is in merge mode and the adjacent / non-adjacent blocks used to obtain motion information include MHP information, the corresponding information can be inherited and used to generate an additional reference block for the current block.
[0177] As an embodiment, as described above, the multi-reference block information can be obtained through signaling or a derivation process, and both methods can be applied. As an example, when N multi-reference blocks are configured, when there are M multi-reference blocks (M<=N) obtained through a derivation process, NM multi-reference block information can be obtained through signaling.
[0178] The additional reference blocks obtained by signaling or a derivation method can generate a final prediction block by weighted sum in the following manner. As an example, when multiple reference blocks are applied, the final prediction block can be calculated as in the following formula 1. Formula 1 assumes that P0 and P1 exist as regular reference blocks, and P2 and P3 exist additionally.
[0179] [Formula 1]
[0180] Step 1: P' = (P0 + P1) / 2
[0181] Step 2: P'' = W0 * P2 + (1-W0) * P'
[0182] Step 3: P = W1 * P3 + (1-W1) * P''
[0183] In Formula 1, W0 and W1 represent weights applied to the additional reference blocks P2 and P3, respectively. In the first step of Formula 1, the basic prediction block may be generated by the weighted sum of the conventional reference blocks, and in the second and third steps of Formula 1, the weighted sum of the additional reference blocks may be performed. As described above, the basic prediction block may refer to a reference conventional reference block before the weighted sum, or may refer to a weighted summed prediction block.
[0184] The above-mentioned multi-reference block mode generates a final prediction block by using the prediction block information obtained by the signaling and derivation method. However, the existing multi-reference block mode is only applied to specific modes, such as merge mode, AMVP mode, etc., and because the motion information for the additional reference block is also obtained by using the limited method for MHP_MERGE and MHP_AMVP, it is difficult to reflect the characteristics of various images. Therefore, in an embodiment of the present disclosure, a method for supporting multiple reference blocks in the auxiliary prediction mode is described.
[0185] Fig.10 is a flowchart illustrating a syntax parsing structure to which an embodiment of the present disclosure can be applied.
[0186] According to an embodiment of the present disclosure, the syntax element transmitted by the signal may be Fig.10 The branch shown in determines the mode applied to the current processing block. In this embodiment, the syntax element names used in the existing image compression technology (VVC) are borrowed, but the present disclosure is not limited thereto, and it is natural that the syntax name, application order, whether the technology is supported, etc. can be changed.
[0187] refer to Fig.10 , the modes applied to the coding unit can be roughly divided into intra / inter / IBC modes, and Fig.10 The parsing structure of the inter / IBC mode is shown. The inter prediction mode can be divided into a merge / skip mode and an inter mode. The merge / skip mode, i.e., the general merge mode, can be selected as one of the subblock merge mode (merge_subblock_flag == 1), the regular merge mode (regular_merge_flag == 1), the merge mode with motion vector difference (MMVD) mode (mmvd_merge_flag == 1), the combined inter-intra prediction (CIIP) mode (ciip_flag == 1), and the geometric partition mode (GPM) mode (ciip_flag == 0) to obtain prediction information through signaling or a predetermined derivation method.
[0188] Specifically, in general merge mode, whether subblock merge mode is applied may be checked based on merge_subblock_flag. When subblock merge mode is selected, an MVP candidate list for subblock merge mode may be configured, and a motion vector may be obtained by using a signaled merge index (ie, merge_subblock_idx).
[0189] In addition, when the normal merge mode is selected, the MVP candidate list for the merge mode may be configured, and then the motion vector may be obtained by using the merge index (ie, merge_idx) transmitted by the signal. In addition, when the MMVD mode is selected, the MVP candidate list may be configured, and the motion vector may be obtained by using the MVD information (ie, mmvd_cand_flag, mmvd_distance_idx, mmvd_direction_idx) transmitted by the signal. When the CIIP mode is selected, the MVP candidate list for the CIIP mode may be configured, and the motion vector may be obtained by using the merge index transmitted by the signal.
[0190] In an embodiment, the intra prediction mode for CIIP mode may be signaled or derived. When GPM mode is selected, an MVP candidate list for GPM mode may be configured, and a motion vector for each partition may be obtained by using the signaled partition information and merge index (ie, merge_gpm_partition_idx, merge_gpm_idx0, merge_gpm_idx1).
[0191] When the general merge mode is not applied, the IBC mode or the inter mode may be applied. When the IBC mode is selected (ie, MODE_IBC == 1), the motion vector (or block vector) may be obtained by using the MVP flag and MVD (ie, mvp_l0_flag, motion vector difference data) for the IBC block. When the inter mode is selected (ie, MODE_IBC == 0), the information sent by the signal may vary depending on whether the affine mode is used. As an embodiment, when the affine mode is selected (ie, inter_affine_flag == 1), the final motion information may be derived by using the reference index, the motion vector predictor flag, and the motion vector difference data for each direction (ie, ref_idx_lX, mvp_lX_flag, mvdLX for each control point). When the affine mode is not selected, the final motion information may be derived by using the reference index, the motion vector predictor flag, and the motion vector difference data for each direction (ie, ref_idx_lX, mvp_lX_flag, mvdLX). In this case, X represents a prediction direction and may be expressed as a value of 0 or 1.
[0192] When symmetric motion vector difference (SMVD) is applied (ie, sym_mvd_flag == 1), L0 motion information may be signaled (ie, ref_idx_l0, mvp_l0_flag, mvdL0), and L1 motion information may derive the final motion information by mirroring the L0 motion information.
[0193] The modes described in this embodiment refer to prediction methods included in existing image compression technologies, but are not limited to the listed modes. In other words, unlisted intra prediction methods / inter prediction methods / intra prediction methods, etc. may be included. As an example, the AMVP-MERGE mode may be included. The AMVP-MERGE mode represents a method for signaling motion information in one direction by including MVD information like AMVP and signaling motion information in another direction by including only a merge index like the merge mode.
[0194] As another example, the GPM-INTRA mode may be included. The GPM-INTRA mode has two prediction blocks based on angle partitioning like the GPM, but it indicates a mode for performing inter-intra or intra-inter combined prediction. For ease of description, the parsing structure of the intra mode is omitted, but is not limited to the inter technique, and the method proposed in this embodiment may be applied by including the intra mode.
[0195] Fig.11 is a flowchart illustrating a syntax parsing structure and a prediction mode determination method according to an embodiment of the present disclosure.
[0196] According to an embodiment of the present disclosure, multiple prediction blocks can be generated based on multiple prediction modes, and a final prediction block can be generated by weighted sum. In other words, the prediction information for the current block can be sent N times with a signal, and the final prediction block can be generated using the prediction information sent N times with a signal. In the present disclosure, for the convenience of description, the case where N is 2 is mainly described, but the present disclosure is not limited thereto. When N is 2, the first prediction mode can be referred to as the primary prediction mode, and the second prediction mode can be referred to as the secondary prediction mode.
[0197] To signal the secondary prediction mode, Fig.11 The changes shown above in Fig.10 The grammatical structure described in . It can be applied in this embodiment basically the same Fig.11 The method 10 described in the present invention is described herein, and the related redundant description is omitted.
[0198] refer to Fig.11 , a flag for determining whether multiple reference blocks are applied to each coding unit or prediction unit can be signaled. As an example, multiple_pred_flag represents a syntax element indicating whether multiple reference blocks are applied. When multiple_pred_flag is "true", the corresponding loop can be repeated N times (ie, LoopIdx = 1..N). This means that when there are N modes, each mode can have a prediction mode tree within an independent general merge mode. For ease of description, when N is configured to be 2, there can be a primary prediction mode and a secondary prediction mode. As described above, this is an example, and naturally there can be a fixed number of modes.
[0199] In an embodiment, the primary prediction mode and the secondary prediction mode may be determined as follows.
[0200] - The primary prediction mode and the secondary prediction mode may include one of the modes included in the general merge mode.
[0201] - The primary prediction mode and the secondary prediction mode can have the same mode.
[0202] In addition, in an embodiment, the primary prediction mode and the secondary prediction mode may be changed and determined as follows to improve compression performance and reduce encoding / decoding complexity. It is naturally possible to apply a combination of the methods listed below.
[0203] - When not B_SLICE, multiple_pred_flag may be inferred to be 0. In other words, the application of the secondary prediction mode may be limited. In other words, only the primary prediction mode may be included.
[0204] - There may be a restriction that the primary prediction mode and the secondary prediction mode do not have the same mode. In other words, the mode determined to be the primary prediction mode may be excluded from the candidate modes for the secondary prediction mode.
[0205] - A set limited to combinations of primary prediction mode and secondary prediction mode can be defined. As an example, when it is {primary prediction mode, secondary prediction mode}, the combination of {GPM, CIIP} and {CIIP, GPM} can be limited. Combinations between partition-based prediction blocks can be limited to increase the compression efficiency of blocks including corresponding characteristics, and specific modes can be limited to reduce encoding / decoding complexity. This is an example, and for the same reason, a specific combination of at least one of {IBC, CIIP} and {CIIP, IBC}, {AFFINE (sub-block merge mode or affine inter-frame mode), INTRA} and {INTRA, AFFINE}, or {AFFINE, CIIP} and {CIIP, AFFINE} can be limited.
[0206] - The secondary prediction mode can be restricted to the normal merge mode. This can reduce the bit usage of the prediction mode tree such as merge_subblock_flag, ciip_flag, etc., and increase compression efficiency.
[0207] - Whether to signal a mode for the secondary prediction mode may be determined according to the primary prediction mode. As an example, when the primary prediction mode is the GPM mode, the secondary prediction mode may allow only the normal merge mode. In other words, when multiple_pred_flag is 1, only the merge index may be included without including a flag for determining the second prediction mode. As another example, when the primary prediction mode is the CIIP mode, the secondary prediction mode may allow only the normal merge mode.
[0208] - When the primary prediction mode is not skip mode, the secondary prediction mode may exist. In other words, when the primary prediction mode is skip mode, multiple_pred_flag may be inferred to be 0 without being signaled.
[0209] - The secondary prediction mode may be limited based on the size and / or shape of the block. For example, when the size (width x height) of the current block is less than a predefined threshold, multiple_pred_flag may be inferred to be 0 without being signaled. Here, width indicates the width of the current block, and height indicates the height of the current block.
[0210] - When the primary prediction mode and the secondary prediction mode have the same mode, different methods may be used to configure the MVP candidate list for deriving each motion information.
[0211] In addition, in an embodiment, whether the above-mentioned multiple reference blocks and the number of prediction modes are applied may be defined and determined (or signaled) as follows. As an example, whether multiple reference blocks are applied may be signaled in a sequence parameter set (SPS), a picture parameter set (PPS), a picture header (PH), a slice header (SH), or a coding unit (CU). In addition, as an example, when multiple reference blocks are applied, the number of prediction modes may be signaled in SPS, PPS, PH, SH, and CU.
[0212] Fig.12 is a flowchart illustrating a prediction mode determination method according to an embodiment of the present disclosure.
[0213] According to an embodiment of the present disclosure, a method for converting the above Fig.11 Another method for applying the secondary prediction mode in a specific mode. The method described in this embodiment can be used to reduce the signaling bits for the secondary prediction mode. In order to signal the information about the secondary prediction mode, it can be as follows Fig.12 Change the above as shown in Fig.10 and Fig.11 The grammatical parsing structure in . Can be applied in this embodiment in substantially the same manner Fig.10 and Fig.11 The method described in is used herein and the related redundant description is omitted.
[0214] refer to Fig.12 , multiple_pred_flag can be sent by signal to indicate whether multiple reference blocks are applied, such as Fig.12 In other words, without a flag for the prediction mode tree, the secondary prediction mode can be derived using only the merge index (ie, merge_idx_2nd).
[0215] In other words, a flag (eg, multiple_pred_flag) for determining whether to apply multiple reference blocks to each coding unit or prediction unit is signaled, and when the corresponding flag is "true", N prediction modes may exist. When N is configured as 2 for convenience of description, the primary prediction mode may derive the prediction mode by using an existing method, and the secondary prediction mode may include prediction information derived by using merge_idx_2nd indicated in the determined MVP candidate list.
[0216] This embodiment describes a case where an auxiliary prediction mode is additionally present in an existing prediction mode, but at least two modes may be added, and each mode may include one of a merge mode or an inter mode. In other words, when the maximum number of additional reference blocks is MaxNum, it is possible to check whether an additional reference block exists through multiple_pred_flag. When multiple_pred_flag is '1', an additional reference block exists, and the additional reference block can be distinguished by multiple_pred_mrg as to whether it is an MHP_MERGE mode or an MHP_AMVP mode. When it is an MHP_MERGE mode, merge_index_2nd may be signaled, and when it is an MHP_AMVP mode, refIdx_2nd, mvp_idx_2nd, and mvd_data_2nd may be signaled.
[0217] In addition, in an embodiment, the primary prediction mode and the secondary prediction mode may be changed and determined as follows to improve compression performance and reduce encoding / decoding complexity. It is naturally possible to apply a combination of the methods listed below.
[0218] - When not B_SLICE, multiple_pred_flag may be inferred to be 0. The secondary prediction mode may be restricted. In other words, only the primary prediction mode may be included.
[0219] - There may be a restriction that the primary prediction mode and the secondary prediction mode do not have the same mode. In other words, the mode determined as the primary prediction mode may be excluded from the candidate modes for the secondary prediction mode. As an example, when the primary prediction mode is determined as the AMVP mode, the secondary prediction mode may be determined as the merge mode without the multiple_pred_mrg flag, which may be applied in the same manner for the opposite case.
[0220] - When the primary prediction mode is not skip mode, the secondary prediction mode may exist. In other words, when the primary prediction mode is skip mode, multiple_pred_flag may be inferred to be 0 without being signaled.
[0221] - The secondary prediction mode may be limited according to the size and / or shape of the block. As an example, when the size of the current block is less than a threshold, multiple_pred_flag may be called 0 without being signaled.
[0222] - When the main prediction mode is AMVP mode and the value of MVD is greater than a predefined threshold, multiple_pred_flag can be called 0 without being signaled. In this case, the value of MVD (MvdX, MvdY), abs (MvdX) ^ 2 + abs (MvdY) ^ 2, abs (MvdX), abs (MvdY) or abs (MvdX) + abs (MvdY), etc., can be compared with the threshold.
[0223] - When the primary prediction mode and the secondary prediction mode have the same mode, different methods may be used to configure the MVP candidate list for deriving each motion information.
[0224] In addition, in an embodiment, whether the above-mentioned multiple reference blocks and the number of prediction modes are applied may be defined and determined (or signaled) as follows. As an example, whether multiple reference blocks are applied may be signaled in SPS, PPS, PH, SH, and CU. In addition, as an example, when multiple reference blocks are applied, the number of prediction modes may be signaled in SPS, PPS, PH, SH, and CU.
[0225] Fig.13 is a flowchart illustrating a prediction mode determination method according to an embodiment of the present disclosure.
[0226] According to the embodiment of the present disclosure, when there is the above Fig.11 and Fig.12 When multiple reference blocks are described in the description, a method for applying weights when generating a final prediction block by using each reference block is described. In order to signal information about the secondary prediction mode, it is possible to Fig.13 Changes as shown above Figures 10 to 12 The syntax parsing structure described in . It can also be applied in this embodiment in a substantially similar manner Figures 10 to 12 The method described in is used herein and the related redundant description is omitted.
[0227] refer to Fig.13 , shows the structure when the weight index is signaled in the above method. As an embodiment, the weight index may be signaled regardless of the mode. In other words, there may be weight indexes for the primary prediction mode and the secondary prediction mode, respectively. Alternatively, Fig.13 Unlike shown in , the weight index can be signaled only for inter mode. For merge mode, the derived (or predefined) weight index can be used.
[0228] In addition, in an embodiment, the weight index may not be signaled for the IBC mode or INTRA mode with unidirectional (single) prediction information. However, the weight index may be signaled for the IBC mode or INTRA mode with bidirectional (multiple) prediction information.
[0229] In addition, in the embodiment, in the above Fig.11 In the method shown in , the weight index may be signaled regardless of the mode for the general merge mode only, and the weight index for the primary prediction mode and the secondary prediction mode may be signaled separately. In this case, for a specific mode, changes including using the derived weight index may be applied. Alternatively, in the embodiment above Fig.12 In the method described in , the weight index for the secondary prediction mode may be signaled. In this case, when the primary prediction mode is a specific mode, changes including using the derived weight index may be applied.
[0230] In an embodiment of the present disclosure, a weight index candidate set may be defined according to the precision of the weight index. For example, when the precision level is 8, weight candidates including at least one of (1 / 8, 2 / 8, 3 / 8, 4 / 8, -1 / 8, -2 / 8, -3 / 8, -4 / 8) may be used. In addition, when the precision level is 16, weight candidates including at least one of (1 / 16, 2 / 16, 3 / 16, 4 / 16, 5 / 16, 6 / 16, 7 / 16, 8 / 16, -1 / 16, -2 / 16, -3 / 16, -4 / 16, -5 / 16, -6 / 16, -7 / 16, -8 / 16) may be used. In addition, when the precision level is 32, weight candidates including at least one of (1 / 32, 2 / 32, 3 / 32, 4 / 32, 5 / 32, 6 / 32, 7 / 32, 8 / 32, 9 / 32, 10 / 32, 11 / 32, 12 / 32, 13 / 32, 14 / 32, 15 / 32, 16 / 32, -1 / 32, -2 / 32, -3 / 32, -4 / 32, -5 / 32, -6 / 32, -7 / 32, -8 / 32, -9 / 32, -10 / 32, -11 / 32, -12 / 32, -13 / 32, -14 / 32, -15 / 32, -16 / 32) can be used.
[0231] The additional reference blocks obtained by the signaling or deriving method can generate the final prediction block by weighted sum, as shown in Table 2 below. In other words, the final prediction block can be generated as follows by using the weights W0 and W1 obtained for each additional reference block by the signaling or deriving method. In this case, whether to signal / derive the weight index can be determined according to whether the primary prediction mode and the secondary prediction mode are bidirectional.
[0232] [Table 2]
[0233]
[0234] Referring to Table 2, the process of performing weighted sum on the prediction block according to whether the prediction block is bidirectional can be changed and applied as follows. As an example, when the primary prediction mode is bidirectional prediction and the secondary prediction mode is unidirectional prediction, the weighted sum can be performed as in the following formula 2 according to the weight index for the secondary prediction mode.
[0235] [Formula 2]
[0236] Step 1: P' = W0 * P0 + (1-W0) * P1
[0237] Step 2: P = W1 * P2 + (1-W1) * P'
[0238] As another example, when the primary prediction mode is unidirectional prediction and the secondary prediction mode is unidirectional prediction, a weighted sum may be performed as in Formula 3 below according to a weight index for the secondary prediction mode.
[0239] [Formula 3]
[0240] Step 1: P = W1*P2 + (1-W1)*P0
[0241] As another example, when the primary prediction mode and the secondary prediction mode are bi-directional predictions and there is one weight, a weighted sum may be performed as in Formula 4 below.
[0242] [Formula 4]
[0243] Step 1: P' = (P0 + P1) / 2
[0244] Step 2: P'' = (P2 + P3) / 2
[0245] Step 3: P = W0*P' + (1-W0)*P''
[0246] The above weight index can be signaled under certain conditions, and the following conditions can be considered. Naturally, one or more combinations of the methods listed below can be applied. The weight when signaling is not performed can be configured as a default mode. In this case, the default mode means that the weight for the bidirectional prediction block is 1 / 2.
[0247] - When it is B_SLICE, it can be sent with a signal.
[0248] - When the primary prediction mode and the secondary prediction mode are bi-directionally predicted (when they have two prediction blocks), they can be signaled separately.
[0249] - Can be signaled when not in IBC mode with one motion information (MV or BV).
[0250] - When not in INTRA mode, signaling can be used.
[0251] - When not in GPM mode, you can send signals.
[0252] - When not in CIIP mode, signaling can be used.
[0253] Additionally, in an embodiment, when deriving weights from adjacent / non-adjacent blocks in merge mode, the following method may be applied.
[0254] - When the Temporal Motion Vector Predictor (TMVP) is selected as the merge candidate, the weights for multiple reference blocks can be configured as the default mode.
[0255] - When the MVP is determined from spatial neighboring blocks, the weight index for multiple reference blocks can be derived and determined from the information of the corresponding blocks.
[0256] - When to select a history-based candidate may be derived and determined from the information stored in the buffer for HMVP.
[0257] - When a combined candidate such as a paired candidate, a combined bi-prediction candidate, or a zero candidate is determined as the MVP, the weight index for multiple reference blocks may be configured as a default mode.
[0258] In addition, in an embodiment, the weights may be derived as follows without signaling. As an example, the following method may be applied in a template-based prediction method. The template-based cost may be calculated by using the sum of absolute transform differences (SATD) between neighboring samples of the current block and neighboring samples of the reference block. Specifically, the template-based cost between the current block and each reference block may be calculated, and the weight W may be derived and applied as in Formula 5 below.
[0259] [Formula 5]
[0260] W = cost2 / (cost1+ cost2)
[0261] In Formula 5, cost1 represents the SATD between the adjacent samples of the current block and the adjacent samples of the first prediction block, and cost2 represents the SATD between the adjacent samples of the current block and the adjacent samples of the second prediction block. The first prediction block and the second prediction block may refer to the bidirectional prediction blocks P0 and P1 of the main prediction mode or the bidirectional prediction blocks P2 and P3 of the auxiliary prediction mode for performing bidirectional prediction of the main prediction mode and the auxiliary prediction mode or the unidirectional prediction blocks P0 (P1) and P2 (P3). In addition, the cost is not limited to SATD, and SAD or MR-SAD, etc. may be used.
[0262] Fig.14 is a flowchart illustrating a prediction method based on multiple reference blocks according to an embodiment of the present disclosure.
[0263] According to the embodiments of the present disclosure, in the application Figures 7 to 13 In the case of a prediction method based on multiple reference blocks described in , a method for reducing the signaling bits for each additional reference block is described. In other words, because candidate indexes and / or weight indices are signaled for each reference mode in the multiple reference block mode, the bits used for candidate index signaling can be reduced as in this embodiment. When there are a normal reference mode and an additional reference mode, the method described in this embodiment can be applied to each mode. In addition, it can be applied to multiple additional reference modes in addition to the normal reference mode.
[0264] refer to Fig.14 , when there are a normal reference mode (ie, primary prediction mode) and an additional reference mode (ie, secondary prediction mode), if the MVP information for each reference block is derived from the same prediction mode, different MVP candidate lists can be configured to derive the MVP or a shared MVP candidate list. Fig.14 In the example, it is assumed that the MVP candidate list is shared. In this case, a reordering of the candidate list can be performed.
[0265] In other words, Fig.14 As shown in , after configuring an MVP candidate list, motion information can be derived by using the MVP index, mvp_idx_1st, and mvp_idx_2nd for the primary prediction mode and the secondary prediction mode. When there are multiple additional reference modes, the same method can be applied to each additional reference mode except the normal reference mode. In this case, when it is applied only to the additional reference mode, for ease of description, by assuming that there are two additional reference modes, each is described as a primary prediction mode and a secondary prediction mode.
[0266] In addition, in an embodiment, the above embodiment may be changed and applied as follows. The MVP candidate list is shared, but the candidate index for the secondary prediction mode (i.e., mvp_idx_2nd) may not be signaled. In this case, the candidate index for the secondary prediction mode may be applied as a minimum value other than mvp_idx_1st+1 or mvp_idx_1st as follows. Alternatively, the signaling bits for the candidate indexes for the primary prediction mode and the secondary prediction mode may be reduced by determining the first candidate in the candidate list as mvp_idx_1st and the second candidate as mvp_idx_2nd.
[0267] Fig.15 is a flowchart illustrating a prediction method based on multiple reference blocks according to an embodiment of the present disclosure.
[0268] refer to Fig.15 , it is possible to allow only MHP_MERGE mode without signaling the mhp_mrg flag to simplify the signaling for multiple reference blocks. Fig.15 A flow chart for a simplified multiple reference block that always only allows MHP_MERGE mode without mhp_mrg is shown. According to this embodiment, the signaling bits can be reduced by signaling only the mhp_flag, merge index and weight index for the multiple reference block.
[0269] As an embodiment, the multi-reference block signaling method according to this embodiment may be variably applied according to the prediction method. As an example, when an MVP improvement method such as decoder-side motion vector refinement (DMVR) is applied, only MHP_MERGE may be allowed by taking into account that the accuracy of motion information is improved. Here, the DMVR method is an example, and the same method may also be applied to another prediction method to which MVP improvement is applied. In addition, when MVD information such as MMVD is included, only MHP_MERGE may be allowed by taking into account that the accuracy of motion information is improved. On the contrary, when MVD information is mirrored and exported like SMVD, the accuracy of motion information is low, so only MHP_AMVP may be allowed.
[0270] In addition, in an embodiment of the present disclosure, a method for reducing signaling information for each additional reference block when a plurality of additional reference blocks are applied is described.
[0271] In an embodiment, when there are two additional reference blocks and they are MHP_MERGE or have the same mode, one candidate index may be signaled for the two additional reference blocks. For example, the candidate indexes in Table 3 below may be defined.
[0272] [Table 3]
[0273]
[0274] Referring to Table 3, the candidate index may indicate a candidate for each additional reference block. Since two indexes are paired and applied as one index, signaling bits may be reduced. However, the indexes in Table 3 are examples, and the combination of indexes may be changed.
[0275] In addition, in an embodiment, the weight index for multiple reference blocks may be defined as shown in Table 4 below.
[0276] [Table 4]
[0277]
[0278] Referring to Table 4, the index may indicate a weight candidate for each additional reference block. Since two indexes are paired and applied as one index, signaling bits may be reduced. However, the indexes in Table 4 are examples, and the combination of indexes may be changed.
[0279] In addition, in an embodiment, candidate indexes and weight indexes for a multi-reference block may be paired and defined as in the following Table 5. However, the indexes in Table 5 are examples, and combinations of indexes may be changed.
[0280] [Table 5]
[0281]
[0282] Fig.16 is a diagram for describing a candidate list configuration method according to an embodiment of the present disclosure.
[0283] According to an embodiment of the present disclosure, when multiple additional reference blocks are applied, bidirectional prediction candidates can be configured to reduce signaling information for each additional reference block. Because each additional reference block allows one reference block and the candidate index for each block is signaled, the candidate index can be saved by applying the bidirectional prediction block as the additional reference block.
[0284] refer to Fig.16 , the MVP candidate list may include spatial neighboring blocks, temporal neighboring blocks, non-neighboring blocks, history-based candidates, etc. The configured list may include the following process for additional reference blocks. However, Fig.16 The embodiment shown in shows an example of a method for configuring bidirectional prediction candidates, and the order of each process may be changed and applied.
[0285] In an embodiment, the MVP candidates can be reordered according to whether the MVP candidate is a unidirectional prediction candidate or a bidirectional prediction candidate. In this case, the bidirectional prediction candidate can be positioned at the front of the list. In other words, a high priority can be assigned.
[0286] Alternatively, in an embodiment, the candidates in the list may be reordered based on the cost. In this case, the bidirectional prediction candidates among the candidates included in the candidate list may be reordered. Alternatively, the candidates within a specific index range may be reordered. Here, the cost may be calculated based on the SAD between the template area of the current block and the template area of the reference block. Alternatively, when bidirectional prediction is performed, the cost may be calculated based on the SAD between two reference blocks. In this case, SAD is an example and may be changed and applied as SATD, MR-SAD, etc.
[0287] Alternatively, in an embodiment, a unidirectional prediction block may be applied by being converted into a bidirectional prediction candidate. When the current picture is a B_SLICE, the unidirectional motion information may be derived by mirroring the motion information in the other direction. In the present disclosure, mirroring may refer to being derived as a value having the same size and a different direction (or a different sign).
[0288] Fig.17 A rough configuration of the inter-frame predictor 332 that performs the inter-frame prediction method according to the present disclosure is shown.
[0289] refer to Fig.17 , the inter predictor 332 may include a basic prediction block generating unit 1700 , an additional reference block deriving unit 1710 , and a final prediction block generating unit 1720 .
[0290] The basic prediction block generation unit 1700 may perform unidirectional prediction or bidirectional prediction to generate a basic prediction block. When a multi-reference block mode is applied, the basic prediction block generation unit 1700 may generate and combine additional prediction blocks in addition to the prediction blocks generated (or derived) by the unidirectional or bidirectional prediction. As an example, the basic prediction block may include an L0 prediction block and / or an L1 prediction block. In the present disclosure, the basic prediction block may be referred to as an initial prediction block, a temporary prediction block, a reference prediction block, etc. In addition, as an example, the basic prediction block may be a prediction block obtained by weighted summing of the L0 and L1 prediction blocks.
[0291] In an embodiment, weights may be used for the weighted sum of L0 and L1 prediction blocks. As weights used for weighted prediction, weights may be collectively referred to as bi-prediction with CU-based weights (BCW) or CU-based weights (CW). Weights may be derived from a weight candidate list. The weight candidate list may include multiple weight candidates and may be predefined in an encoding / decoding device.
[0292] The weight candidate may be a weight set (i.e., a first weight and a second weight) representing a weight applied to each bidirectional prediction block, or may be a weight applied to a prediction block in either direction. When only the weight applied to a prediction block in any one direction is derived from the weight candidate list, the weight applied to the prediction block in the other direction may be derived based on the weight derived from the weight candidate list. For example, the weight applied to the prediction block in the other direction may be derived by subtracting the weight derived from the weight candidate list from a predetermined value.
[0293] In an embodiment, a weight index indicating a weight used for weighted prediction of the current block may be derived in a weight candidate list. In the present disclosure, the weight index may be referred to as bcw_idx, bcw index. The weight index may be derived by a decoding device, or may be signaled from an encoding device. When derived by a decoding device, the weight index may be derived as a weight index of a specific merge candidate in a merge candidate list. As an example, a specific merge candidate may be specified by a merge index in a merge candidate list.
[0294] The additional reference block deriving unit 1710 may derive an additional reference block based on the MHP mode. The additional reference block deriving unit 1710 may derive an additional reference block other than the basic prediction block, and combine (or weighted sum) the basic prediction block and the generated additional reference block.
[0295] As an embodiment, when the MHP mode is applied, the additional reference block deriving unit 1710 may derive up to a predefined number of additional reference blocks. In other words, the additional reference block deriving unit 1710 may combine (or weighted sum) additional reference blocks less than or equal to the predefined number to the basic prediction block. As an example, the predefined number may be 2. Alternatively, as an example, the predefined number may be one of 1, 2, 3, and 4. The predefined number may be referred to as the maximum number of MHP.
[0296] In addition, when a plurality of additional reference blocks are combined, the plurality of additional reference blocks may be weighted-added to the basic prediction block in sequence. For example, when up to 2 additional reference blocks are derived, the basic prediction block and the first additional reference block may be weighted-added to generate a prediction block, and the generated prediction block and the second additional reference block may be weighted-added to generate a final prediction block. The prediction block generated by weighted-adding the basic prediction block and the first additional reference block may be referred to as an intermediate prediction block.
[0297] Alternatively, when a plurality of additional reference blocks are combined, the basic prediction block and the plurality of generated additional reference blocks may be weighted and summed at once. In other words, after the plurality of additional reference blocks are derived, weights may be applied to each of the plurality of additional reference blocks and the basic prediction block (or the L0 prediction block and the L1 prediction block) and weighted and summed at once.
[0298] In addition, as an embodiment, the additional reference block derivation unit 1710 may determine whether to apply MHP. As an example, whether to apply MHP may be explicitly signaled, or may be implicitly derived by the decoding device.
[0299] In addition, as an embodiment, whether MHP is applied may be signaled from the encoding device to the decoding device. For example, an MHP flag indicating whether MHP is applied may be signaled from the encoding device to the decoding device. In this case, conditions for signaling / parsing the MHP flag may be defined in advance. The signaling / parsing condition of the MHP flag may be an availability condition of the MHP. When the signaling / parsing condition is met, the decoding device may parse the MHP flag from the bitstream. Alternatively, as an embodiment, whether MHP is applied may be derived by the decoding device based on predefined encoding information. As an example, whether MHP is applied may be defined in the same manner as the MHP availability condition (or signaling / parsing condition) described below.
[0300] In addition, as an embodiment, the additional reference block deriving unit 1710 may obtain MHP information (or may be referred to as MHP prediction information) to derive an additional reference block. As an example, the MHP information may include weight information and / or prediction information. A reference block according to an MHP mode, that is, an additional reference block, may be derived based on the prediction information, and the additional reference block derived based on the weight information may be weighted summed with a basic prediction block (or an intermediate prediction block). In addition, as an example, the MHP information may further include an MHP flag indicating whether MHP is applied.
[0301] As an example, the prediction information may include mode information for deriving an additional reference block and motion information according to the mode. The mode information may be inter-prediction mode information indicating whether it is a merge mode or an AMVP mode. For example, the mode information may be a merge flag. In other words, the additional reference block may be derived using the merge mode or the AMVP mode, and a flag syntax element indicating it may be signaled. Alternatively, the additional reference block may be derived using a predefined mode among the merge mode or the AMVP mode. Alternatively, the merge mode or the AMVP mode may be selected based on predefined encoding information.
[0302] As an example, when the merge mode is used to derive the additional reference block, the prediction information may include a merge index. The merge index may specify a merge candidate in the merge candidate list. When the AMVP mode is used to derive the additional reference block, the prediction information may include a motion vector predictor flag, a reference index, and motion vector difference information. The motion vector predictor flag may specify a candidate in the motion vector predictor candidate list.
[0303] The final prediction block generation unit 1720 may generate a final prediction block by weighted summing the basic prediction block and the additional reference block. As described above, the number of additional reference blocks may be less than or equal to the predefined number. For example, when the number of additional reference blocks is 2, the final prediction block may be a block in which the basic prediction block and the two additional reference blocks are weighted summed.
[0304] As described above, when a plurality of additional reference blocks are combined, the plurality of additional reference blocks may be weighted-added to the basic prediction block in sequence, or the basic prediction block and the plurality of generated additional reference blocks may be weighted-added at once.
[0305] exist Figures 7 to 16 The above-described embodiments in can be similarly applied, and repeated descriptions related thereto will be omitted.
[0306] Fig.18 An inter-frame prediction method performed by the encoding apparatus 200 as an embodiment according to the present disclosure is shown.
[0307] refer to Fig.18 , the encoding device may perform unidirectional or bidirectional prediction to generate a basic prediction block S1800. When the multi-reference block mode is applied, the encoding device may generate and combine additional prediction blocks in addition to the prediction blocks generated (or derived) by the unidirectional or bidirectional prediction. As an example, the basic prediction block may include an L0 prediction block and / or an L1 prediction block. In the present disclosure, the basic prediction block may be referred to as an initial prediction block, a temporary prediction block, a reference prediction block, etc. In addition, as an example, the basic prediction block may be a prediction block obtained by weighted summing the L0 and L1 prediction blocks.
[0308] In an embodiment, the weights may be used for weighted summation of L0 and L1 prediction blocks. As weights for weighted prediction, the weights may be collectively referred to as bi-prediction with CU-based weights (BCW) or CU-based weights (CW). The weights may be derived from a weight candidate list. The weight candidate list may include multiple weight candidates and may be predefined in an encoding / decoding device.
[0309] The weight candidate may be a weight set (i.e., a first weight and a second weight) indicating a weight applied to each bidirectional prediction block, or may be a weight applied to a prediction block in either direction. When only a weight applied to a prediction block in any one direction is derived from the weight candidate list, a weight applied to a prediction block in the other direction may be derived based on the weight derived from the weight candidate list. For example, the weight applied to a prediction block in the other direction may be derived by subtracting the weight derived from the weight candidate list from a predetermined value.
[0310] In an embodiment, a weight index indicating a weight for weighted prediction of a current block may be derived in a weight candidate list. In the present disclosure, a weight index may be referred to as bcw_idx, bcw index. The weight index may be derived by a decoding device, or may be signaled from an encoding device. When derived by a decoding device, the weight index may be derived as a weight index for a specific merge candidate in a merge candidate list. As an example, a specific merge candidate may be specified by a merge index in a merge candidate list.
[0311] The encoding apparatus may derive an additional reference block based on the multiple reference block mode S1810. The encoding apparatus may derive an additional reference block other than the basic prediction block, and combine (or weighted sum) the basic prediction block and the derived additional reference block.
[0312] As an embodiment, when the multi-reference block mode is applied, the encoding device may derive up to a predefined number of additional reference blocks. In other words, the encoding device may combine (or weighted sum) additional reference blocks less than or equal to the predefined number to the basic prediction block. As an example, the predefined number may be 2. Alternatively, as an example, the predefined number may be one of 1, 2, 3, and 4. The predefined number may be referred to as the maximum number of multi-reference blocks.
[0313] In addition, when a plurality of additional reference blocks are combined, the plurality of additional reference blocks may be weighted-added to the basic prediction block in sequence. For example, when up to 2 additional reference blocks are derived, the basic prediction block and the first additional reference block may be weighted-added to generate a prediction block, and the generated prediction block and the second additional reference block may be weighted-added to generate a final prediction block. The prediction block generated by weighted-adding the basic prediction block and the first additional reference block may be referred to as an intermediate prediction block.
[0314] Alternatively, when a plurality of additional reference blocks are combined, the basic prediction block and the plurality of derived additional reference blocks may be weighted summed at once. In other words, after the plurality of additional reference blocks are derived, weights may be applied to each of the plurality of additional reference blocks and the basic prediction block (or the L0 prediction block and the L1 prediction block) and weighted summed at once.
[0315] In addition, as an embodiment, the encoding device may determine whether to apply multiple reference blocks. In this case, a step of determining whether to apply multiple reference blocks may be added before S1810. As an example, whether to apply multiple reference blocks may be explicitly signaled or may be implicitly derived by the decoding device.
[0316] In addition, as an embodiment, whether to apply multiple reference blocks can be signaled from the encoding device to the decoding device. For example, a multiple reference block flag indicating whether to apply multiple reference blocks can be signaled from the encoding device to the decoding device. In this case, the conditions for signaling / parsing the multiple reference block flag can be defined in advance. The signaling / parsing conditions of the multiple reference block flag can be the availability conditions of the multiple reference blocks. When the signaling / parsing conditions are met, the encoding device can signal the multiple reference block flag from the bitstream. Alternatively, as an embodiment, whether to apply multiple reference blocks can be derived by the decoding device based on predefined encoding information. As an example, whether to apply multiple reference blocks can be defined in the same manner as the multiple reference block availability conditions (or signaling / parsing conditions) described below.
[0317] In addition, as an embodiment, the encoding device may obtain multi-reference block information (or may be referred to as multi-reference block prediction information) to derive an additional reference block. As an example, the multi-reference block information may include weight information and / or prediction information. A reference block according to a multi-reference block mode, that is, an additional reference block, may be derived based on the prediction information, and the additional reference block derived based on the weight information may be weighted summed with a basic prediction block (or an intermediate prediction block). In addition, as an example, the multi-reference block information may further include a multi-reference block flag indicating whether the multi-reference block is applied.
[0318] As an example, the prediction information may include mode information for deriving an additional reference block and motion information according to the mode. The mode information may be inter-prediction mode information indicating whether it is a merge mode or an AMVP mode. For example, the mode information may be a merge flag. In other words, the additional reference block may be derived using the merge mode or the AMVP mode, and a flag syntax element indicating it may be signaled. Alternatively, the additional reference block may be derived using a predefined mode among the merge mode or the AMVP mode. Alternatively, the merge mode or the AMVP mode may be selected based on predefined encoding information.
[0319] As an example, when the merge mode is used to derive the additional reference block, the prediction information may include a merge index. The merge index may specify a merge candidate in the merge candidate list. When the AMVP mode is used to derive the additional reference block, the prediction information may include a motion vector predictor flag, a reference index, and motion vector difference information. The motion vector predictor flag may specify a candidate in the motion vector predictor candidate list.
[0320] The encoding device may generate a final prediction block by weighted summing the basic prediction block and the additional reference block S1820. As described above, the number of additional reference blocks may be less than or equal to the predefined number. For example, when the number of additional reference blocks is 2, the final prediction block may be a block in which the basic prediction block and the two additional reference blocks are weighted summed.
[0321] As described above, when a plurality of additional reference blocks are combined, the plurality of additional reference blocks may be weighted-added to the basic prediction block in sequence, or the basic prediction block and the plurality of derived additional reference blocks may be weighted-added at once.
[0322] exist Figures 7 to 16 The above-described embodiments in may be substantially equally applied, and repeated descriptions related thereto will be omitted.
[0323] Fig.19 A rough configuration of the inter-frame predictor 221 performing the inter-frame prediction method according to the present disclosure is shown.
[0324] refer to Fig.19 , the inter-frame predictor 221 may include a basic prediction block generating unit 1900 , an additional reference block deriving unit 1910 , and a final prediction block generating unit 1920 .
[0325] The basic prediction block generation unit 1900 may perform unidirectional or bidirectional prediction to generate a basic prediction block. When a multi-reference block mode is applied, the basic prediction block generation unit 1900 may generate and combine additional prediction blocks in addition to the prediction blocks generated (or derived) by unidirectional or bidirectional prediction. As an example, the basic prediction block may include an L0 prediction block and / or an L1 prediction block. In the present disclosure, the basic prediction block may be referred to as an initial prediction block, a temporary prediction block, a reference prediction block, etc. In addition, as an example, the basic prediction block may be a prediction block obtained by weighted summing of the L0 and L1 prediction blocks.
[0326] In an embodiment, weights may be used for the weighted sum of L0 and L1 prediction blocks. As weights used for weighted prediction, weights may be collectively referred to as bi-prediction with CU-based weights (BCW) or CU-based weights (CW). Weights may be derived from a weight candidate list. The weight candidate list may include multiple weight candidates and may be predefined in an encoding / decoding device.
[0327] The weight candidate may be a weight set (i.e., a first weight and a second weight) representing a weight applied to each bidirectional prediction block, or may be a weight applied to a prediction block in either of the two directions. When only the weight applied to the prediction block in any one direction is derived from the weight candidate list, the weight applied to the prediction block in the other direction may be derived based on the weight derived from the weight candidate list. For example, the weight applied to the prediction block in the other direction may be derived by subtracting the weight derived from the weight candidate list from a predetermined value.
[0328] In an embodiment, a weight index indicating a weight used for weighted prediction of the current block may be derived in a weight candidate list. In the present disclosure, the weight index may be referred to as bcw_idx, bcw index. The weight index may be derived by a decoding device, or may be signaled from an encoding device. When derived by a decoding device, the weight index may be derived as a weight index of a specific merge candidate in a merge candidate list. As an example, a specific merge candidate may be specified by a merge index in a merge candidate list.
[0329] The additional reference block deriving unit 1910 may derive an additional reference block based on the multiple reference block mode. The additional reference block deriving unit 1910 may derive an additional reference block in addition to the basic prediction block and combine (or weighted sum) the basic prediction block and the derived additional reference block.
[0330] As an embodiment, when the multi-reference block mode is applied, the additional reference block deriving unit 1910 may derive up to a predefined number of additional reference blocks. In other words, the additional reference block deriving unit 1910 may combine (or weighted sum) additional reference blocks less than or equal to the predefined number to the basic prediction block. As an example, the predefined number may be 2. Alternatively, as an example, the predefined number may be one of 1, 2, 3, and 4. The predefined number may be referred to as the maximum number of multi-reference blocks.
[0331] In addition, when a plurality of additional reference blocks are combined, the plurality of additional reference blocks may be weighted-added to the basic prediction block in sequence. For example, when up to 2 additional reference blocks are derived, the basic prediction block and the first additional reference block may be weighted-added to generate a prediction block, and the generated prediction block and the second additional reference block may be weighted-added to generate a final prediction block. The prediction block generated by weighted-adding the basic prediction block and the first additional reference block may be referred to as an intermediate prediction block.
[0332] Alternatively, when a plurality of additional reference blocks are combined, the basic prediction block and the plurality of derived additional reference blocks may be weighted summed at once. In other words, after the plurality of additional reference blocks are derived, weights may be applied to each of the plurality of additional reference blocks and the basic prediction block (or the L0 prediction block and the L1 prediction block) and weighted summed at once.
[0333] In addition, as an embodiment, the additional reference block derivation unit 1910 may determine whether to apply multiple reference blocks. As an example, whether to apply multiple reference blocks may be explicitly signaled or may be implicitly derived by the decoding device.
[0334] In addition, as an embodiment, whether to apply multiple reference blocks can be sent from the encoding device to the decoding device by signal. For example, a multiple reference block flag indicating whether to apply multiple reference blocks can be sent from the encoding device to the decoding device by signal. In this case, the conditions for signaling / parsing the multiple reference block flag can be defined in advance. The signaling / parsing conditions of the multiple reference block flag can be the availability conditions of the multiple reference blocks. When the signaling / parsing conditions are met, the decoding device can parse the multiple reference block flag from the bitstream. Alternatively, as an embodiment, whether to apply multiple reference blocks can be derived by the decoding device based on predefined encoding information. As an example, whether to apply multiple reference blocks can be defined in the same manner as the multiple reference block availability conditions (or signaling / parsing conditions) described below.
[0335] In addition, as an embodiment, the additional reference block deriving unit 1910 may obtain multi-reference block information (or may be referred to as multi-reference block prediction information) to derive an additional reference block. As an example, the multi-reference block information may include weight information and / or prediction information. A reference block according to a multi-reference block mode, i.e., an additional reference block, may be derived based on the prediction information, and the additional reference block derived based on the weight information may be weighted summed with a basic prediction block (or an intermediate prediction block). In addition, as an example, the multi-reference block information may further include a multi-reference block flag indicating whether the multi-reference block is applied.
[0336] As an example, the prediction information may include mode information for deriving an additional reference block and motion information according to the mode. The mode information may be inter-prediction mode information indicating whether it is a merge mode or an AMVP mode. For example, the mode information may be a merge flag. In other words, the additional reference block may be derived using the merge mode or the AMVP mode, and a flag syntax element indicating it may be signaled. Alternatively, the additional reference block may be derived using a predefined mode among the merge mode or the AMVP mode. Alternatively, the merge mode or the AMVP mode may be selected based on predefined encoding information.
[0337] As an example, when the merge mode is used to derive the additional reference block, the prediction information may include a merge index. The merge index may specify a merge candidate in the merge candidate list. When the AMVP mode is used to derive the additional reference block, the prediction information may include a motion vector predictor flag, a reference index, and motion vector difference information. The motion vector predictor flag may specify a candidate in the motion vector predictor candidate list.
[0338] The final prediction block generation unit 1920 may generate a final prediction block by weighted summing the basic prediction block and the additional reference block. As described above, the number of additional reference blocks may be less than or equal to the predefined number. For example, when the number of additional reference blocks is 2, the final prediction block may be a block in which the basic prediction block and two additional reference blocks are weighted summed.
[0339] As described above, when a plurality of additional reference blocks are combined, the plurality of additional reference blocks may be weighted-added to the basic prediction block in sequence, or the basic prediction block and the plurality of generated additional reference blocks may be weighted-added at once.
[0340] exist Figures 7 to 16 The above-described embodiments in can be similarly applied, and repeated descriptions related thereto will be omitted.
[0341] In the above embodiments, the method is described as a series of steps or boxes based on the flowchart, but the corresponding embodiments are not limited to the order of the steps, and some steps may occur simultaneously or in a different order than other steps described above. In addition, those skilled in the art will appreciate that the steps shown in the flowchart are not exclusive, and other steps may be included or one or more steps in the flowchart may be deleted without affecting the scope of the embodiments of the present disclosure.
[0342] The above-mentioned method according to the embodiment of the present disclosure can be implemented in the form of software, and the encoding device and / or decoding device according to the present disclosure can be included in a device that performs image processing, such as a TV, a computer, a smart phone, a set-top box, a display device, etc.
[0343] In the present disclosure, when the embodiment is implemented as software, the above method can be implemented as a module (process, function, etc.) that performs the above functions. The module can be stored in a memory and can be executed by a processor. The memory can be located inside or outside the processor and can be connected to the processor by various well-known means. The processor may include an application-specific integrated circuit (ASIC), another chipset, a logic circuit, and / or a data processing device. The memory may include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium, and / or other storage devices. In other words, the embodiments described herein can be implemented on a processor, a microprocessor, a controller, or a chip. For example, the functional unit shown in each of the figures can be implemented on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information (e.g., information about instructions) or an algorithm for implementation can be stored in a digital storage medium.
[0344] In addition, the decoding device and the encoding device of the embodiment of the present disclosure can be included in multimedia broadcast sending and receiving devices, mobile communication terminals, home theater video devices, digital theater video devices, surveillance cameras, video conversation devices, real-time communication devices such as video communication, mobile streaming devices, storage media, cameras, devices for providing video on demand (VoD) services, over-the-top video (OTT) devices, devices for providing Internet streaming services, three-dimensional (3D) video devices, virtual reality (VR) devices, augmented reality (AR) devices, videophone video devices, transportation terminals (e.g., vehicle (including autonomous driving vehicle) terminals, aircraft terminals, ship terminals, etc.) and medical video devices, etc., and can be used to process video signals or data signals. For example, over-the-top video (OTT) devices can include game consoles, Blu-ray players, networked TVs, home theater systems, smart phones, tablet computers, digital video recorders (DVRs), etc.
[0345] In addition, the processing method of the embodiment of the present disclosure can be generated in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to an embodiment of the present disclosure can also be stored in a computer-readable recording medium. Computer-readable recording media include all types of storage devices and distributed storage devices that store computer-readable data. Computer-readable recording media may include, for example, Blu-ray discs (BD), universal serial buses (USB), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, tapes, floppy disks, and optical media storage devices. In addition, computer-readable recording media include media implemented in the form of carrier waves (e.g., transmitted via the Internet). In addition, the bit stream generated by the encoding method may be stored in a computer-readable recording medium or may be sent via a wired or wireless communication network.
[0346] In addition, the embodiments of the present disclosure may be implemented by a computer program product through a program code, and the program code may be executed on a computer by the embodiments of the present disclosure. The program code may be stored on a computer-readable carrier.
[0347] Fig. 20 An example of a content streaming system to which an embodiment of the present disclosure can be applied is shown.
[0348] refer to Fig. 20 The content streaming transmission system to which the embodiments of the present disclosure are applied may mainly include an encoding server, a streaming transmission server, a web server, a media storage, a user device, and a multimedia input device.
[0349] The encoding server generates a bitstream by compressing content input from a multimedia input device such as a smartphone, a camera, a camcorder, etc. into digital data and transmits it to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, a camcorder, etc. directly generates a bitstream, the encoding server may be omitted.
[0350] A bitstream may be generated by applying the encoding method or the bitstream generating method of the embodiment of the present disclosure, and the streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0351] The streaming server sends multimedia data to the user device through the web server based on the user's request, and the web server serves as a medium to inform the user what services are available. When the user requests the required service from the web server, the web server delivers it to the streaming server, and the streaming server sends the multimedia data to the user. In this case, the content streaming system may include a separate control server, and in this case, the control server controls the command / response between each device in the content streaming system.
[0352] The streaming server may receive content from a media storage and / or encoding server. For example, when receiving content from an encoding server, the content may be received in real time. In this case, in order to provide a smooth streaming service, the streaming server may store the bitstream for a specific period of time.
[0353] Examples of user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation, tablet PCs, tablet computers, ultrabooks, wearable devices (e.g., smart watches, smart glasses, head-mounted displays (HMDs), digital televisions, desktops, digital signage, etc.).
[0354] Each server in the content streaming system may be operated as a distributed server, and in this case, data received from each server may be distributed and processed.
[0355] The claims set forth herein can be combined in various ways. For example, the technical features of the method claims of the present disclosure can be combined and implemented as a device, and the technical features of the device claims of the present disclosure can be combined and implemented as a method. In addition, the technical features of the method claims of the present disclosure and the technical features of the device claims can be combined and implemented as a device, and the technical features of the method claims of the present disclosure and the technical features of the device claims can be combined and implemented as a method.
Claims
1. An image decoding method, the method comprising: performing prediction based on the first prediction mode to generate a basic prediction block of the current block; deriving an additional reference block for the current block based on a second prediction mode; as well as A weighted sum of the basic prediction block and the additional reference block is calculated to generate a final prediction block of the current block.
2. The method according to claim 1, wherein: The method further includes obtaining a flag indicating whether the additional reference block is used for the current block.
3. The method according to claim 1, wherein: The first prediction mode includes at least one of a merge mode, a skip mode, an advanced motion vector prediction (AMVP) mode, an intra-block copy (IBC) mode, or an AMVP-merge combined mode.
4. The method according to claim 3, wherein: The second prediction mode includes at least one of the merge mode, the skip mode, the AMVP mode, the IBC mode, or the AMVP-merge combined mode.
5. The method according to claim 1, wherein: The first prediction mode and the second prediction mode are determined as specific combinations within a predefined prediction mode combination set.
6. The method according to claim 5, wherein: The predefined prediction mode combination set includes a plurality of combination candidates. The plurality of combination candidates are configured by including at least two of a merge mode, a skip mode, an AMVP mode, an IBC mode, an AMVP-merge combination mode, a geometric partition mode (GPM), a combined inter-intra prediction (CIIP) mode, a subblock merge mode, or an affine mode.
7. The method according to claim 1, wherein: Generating the basic prediction block comprises: performing bidirectional prediction to derive a first reference block and a second reference block of the current block; and A weighted sum of the first reference block and the second reference block is calculated to generate the basic prediction block.
8. The method according to claim 1, wherein: Generating the basic prediction block comprises: performing unidirectional prediction to derive a third reference block for the current block; and The basic prediction block is generated based on the third reference block.
9. The method according to claim 1, wherein: When a plurality of additional reference blocks are derived, the final prediction block is generated by sequentially weighting and summing the plurality of additional reference blocks to the basic prediction block.
10. The method according to claim 9, wherein: The information about the second prediction mode includes at least one of weight information or prediction information, The weight information represents information indicating a weight used for the weighted sum of the additional reference blocks, The prediction information represents information used to derive the additional reference block.
11. A method for image encoding, the method comprising: performing bidirectional prediction based on the first prediction mode to generate a basic prediction block of the current block; deriving an additional reference block for the current block based on a second prediction mode; as well as A weighted sum of the basic prediction block and the additional reference block is calculated to generate a final prediction block of the current block. 12 . A computer-readable storage medium storing a bit stream generated by the image encoding method according to claim 11 .
13. A method for transmitting data for image information, the method comprising: performing bidirectional prediction based on the first prediction mode to generate a basic prediction block of the current block; deriving an additional reference block for the current block based on a second prediction mode; Calculating a weighted sum of the basic prediction block and the additional reference block to generate a final prediction block of the current block; Encoding the current block based on the final prediction block to generate a bitstream; as well as Data including the bit stream is transmitted.