Image encoding / decoding method and apparatus, and recording medium for storing bit stream
By deriveing inter prediction mode and sub-block transformation information in the image compression technology, optimizing intra prediction mode and weight, the problem of inefficient encoding in the prior art is solved, and more efficient image encoding is achieved.
Patent Information
- Application Number
- CN202380088919.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-29
- Filing Date
- 2023-12-29
- Publication Date
- 2025-08-12
AI Technical Summary
The existing image compression technology is difficult to effectively derive the inter prediction mode in the merge mode or the skip mode, and the CIIP mode lacks an efficient intra prediction mode and a signal transmission method of weight information, resulting in insolubilization of encoding efficiency.
By generating prediction samples of the current block based on a predetermined inter prediction mode, and derive the residual samples of the current block using the information of sub-block transformation (SBT), it allows the CIIP mode to derive weights based on the intra prediction mode, omit signal transmission of part of SBT information, and optimize information transmission of geometric partition merging mode.
The encoding efficiency of inter-frame prediction is improved, the encoding efficiency of CIIP mode is improved, and the number of bits required to encode SBT information is reduced, thereby improving the overall encoding efficiency.
Smart Images

Figure CN120476583A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and a recording medium storing a bitstream. Background Art
[0002] Recently, demands for high-resolution and high-quality images such as HD (High Definition) images and UHD (Ultra High Definition) images have been increasing in various application fields, and therefore, efficient image compression technology is being discussed.
[0003] There are various technologies, such as inter-frame prediction technology that uses video compression technology to predict pixel values included in the current picture from pictures before or after the current picture, intra-frame prediction technology that predicts pixel values included in the current picture by using pixel information in the current picture, entropy coding technology that assigns short symbols to values with high frequency of occurrence and long symbols to values with low frequency of occurrence, etc., and these image compression technologies can be used to effectively compress image data and transmit or store it. Summary of the Invention
[0004] Technical issues
[0005] The present disclosure seeks to provide a method and apparatus for deriving inter-prediction mode in merge mode or skip mode.
[0006] The present disclosure seeks to provide a method and apparatus for deriving intra prediction modes and / or weights in CIIP mode.
[0007] The present disclosure seeks to provide a method and apparatus for signaling information about GPM mode.
[0008] The present disclosure seeks to provide a method and apparatus for signaling / deriving information about sub-block transform based on inter prediction mode in merge mode or skip mode.
[0009] Technical Solution
[0010] According to the image decoding method and apparatus of the present disclosure, a prediction sample of a current block may be generated based on a predetermined inter-frame prediction mode, a residual sample of the current block may be derived based on information about a sub-block transform (SBT), and the current block may be reconstructed based on the prediction sample and the residual sample of the current block. Here, at least one of the information about the SBT may be derived based on the inter-frame prediction mode, and the information about the SBT may include at least one of an SBT flag, an SBT size flag, an SBT direction flag, or an SBT position flag.
[0011] In the image decoding method and apparatus according to the present disclosure, when the inter prediction mode of the current block is the CIIP mode, SBT may be allowed for the current block.
[0012] In the image decoding method and apparatus of the present disclosure, when the inter prediction mode of the current block is the CIIP mode, at least one of the information about the SBT may not be signaled and may be derived based on the intra prediction mode for the CIIP mode.
[0013] In the image decoding method and apparatus according to the present disclosure, the intra prediction mode of the current block may be derived as any one of one or more candidate modes belonging to the candidate list.
[0014] In the image decoding method and apparatus according to the present disclosure, when the inter prediction mode of the current block is the CIIP mode, at least one of the information about the SBT may not be signaled and may be derived based on the weight for the CIIP mode.
[0015] In the image decoding method and apparatus according to the present disclosure, the weight may be determined based on at least one of an intra prediction mode of the current block, a position of a subregion to which a prediction sample of the current block belongs, or a weight index.
[0016] In the image decoding method and apparatus according to the present disclosure, the CIIP mode may be allowed regardless of a flag indicating whether the current block is a block coded in the skip mode.
[0017] In the image decoding method and apparatus according to the present disclosure, when the inter prediction mode of the current block is the geometric partition merging mode, sub-block transformation may be allowed for the current block.
[0018] In the image decoding method and device according to the present disclosure, when the inter-frame prediction mode of the current block is the geometric partition merge mode, at least one of the information about the SBT may not be sent with a signal and may be derived based on at least one of the angle of the boundary line of the geometric partition of the current block or the distance from the center position of the current block to the boundary line.
[0019] According to the image encoding method and apparatus of the present disclosure, prediction samples of a current block may be generated based on a predetermined inter-frame prediction mode, residual samples of the current block may be derived based on the prediction samples of the current block, information about a sub-block transform (SBT) for encoding the residual samples of the current block may be determined, the residual samples of the current block may be encoded to generate residual information, and the residual information of the current block may be encoded to generate a bitstream. Here, at least one of the information about the SBT may be derived based on the inter-frame prediction mode, and the information about the SBT may include at least one of an SBT flag, an SBT size flag, an SBT direction flag, or an SBT position flag.
[0020] A computer-readable digital storage medium is provided, which stores encoded video / image information, thereby causing the image decoding method to be executed by the decoding device according to the present disclosure.
[0021] A computer-readable digital storage medium storing video / image information generated according to an image encoding method is provided according to the present disclosure.
[0022] Provided are a method and apparatus for transmitting video / image information generated according to an image encoding method according to the present disclosure.
[0023] Beneficial effects
[0024] According to the present disclosure, by defining an inter prediction mode available in a merge mode or a skip mode and efficiently signaling it, encoding efficiency of inter prediction can be improved.
[0025] According to the present disclosure, by determining an optimal intra prediction mode and a weight for the CIIP mode, the encoding efficiency of the CIIP mode can be improved.
[0026] According to the present disclosure, by more efficiently signaling information about the geometric partition merging mode, the number of bits required to encode the information can be reduced.
[0027] According to the present disclosure, encoding efficiency can be improved by omitting signaling of all or part of information about subblock transforms and deriving it as a specific value depending on an inter prediction mode in a merge mode or a skip mode. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 A video / image coding system according to the present disclosure is shown.
[0029] Figure 2 A schematic block diagram illustrating an encoding device to which embodiments of the present disclosure are applicable and which performs encoding of a video / image signal.
[0030] Figure 3 A schematic block diagram illustrating a decoding device to which embodiments of the present disclosure are applicable and which performs decoding of a video / image signal.
[0031] Figure 4 The diagram illustrates an image decoding method performed in an image decoding device (300) according to the present disclosure.
[0032] Figure 5 The predefined intra prediction modes and prediction directions thereof according to the present disclosure are exemplarily shown.
[0033] Figure 6 The angle of the boundary line corresponding to each angleIdx according to the present disclosure is illustrated.
[0034] Figure 7 The diagram illustrates a schematic configuration of a decoding device (300) for performing an image decoding method according to the present disclosure.
[0035] Figure 8 The diagram illustrates an image encoding method performed in an encoding device (200) according to the present disclosure.
[0036] Figure 9 The diagram illustrates a schematic configuration of an encoding device (200) that performs an image encoding method according to the present disclosure.
[0037] Figure 10 An example of a content streaming system to which embodiments of the present disclosure can be applied is shown. DETAILED DESCRIPTION
[0038] Because the present disclosure can be modified in various ways and has several embodiments, specific embodiments will be illustrated in the drawings and described in detail in the detailed description. However, it is not intended to limit the present disclosure to specific embodiments, and it should be understood that all variations, equivalents, and alternatives are included in the spirit and technical scope of the present disclosure. When describing each of the drawings, similar reference numerals are used for similar components.
[0039] Terms such as first, second, etc. may be used to describe various components, but components should not be limited by these terms. These terms are only used to distinguish one component from other components. For example, a first component may be referred to as a second component, and similarly, a second component may be referred to as a first component without departing from the scope of the present disclosure. Terms and / or combinations of any one or more related statement items include multiple related statement items.
[0040] When a component is referred to as being "connected" or "linked" to another component, it should be understood that it can be directly connected or linked to the other component, but another component may also exist in between. On the other hand, when a component is referred to as being "directly connected" or "directly linked" to another component, it should be understood that another component does not exist in between.
[0041] The terms used in this application are only used to describe specific embodiments and are not intended to limit the present disclosure. Unless the context clearly indicates otherwise, singular expressions include plural expressions. In this application, it should be understood that terms such as "including" or "having" are intended to designate the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0042] The present disclosure relates to video / image coding. For example, the methods / embodiments disclosed herein may be applied to methods disclosed in the Versatile Video Coding (VVC) standard. Furthermore, the methods / embodiments disclosed herein may be applied to methods disclosed in the Essential Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second-generation Audio Video Coding standard (AVS2), or next-generation video / image coding standards (e.g., H.267 or H.268).
[0043] This specification proposes various embodiments of video / image coding, and unless otherwise stated, these embodiments may be performed in combination with each other.
[0044] Here, video can refer to a collection of images over time. A picture generally refers to a unit representing an image within a specific time period, and a slice / tile is a unit that forms part of a picture during coding. A slice / tile may include at least one coding tree unit (CTU). A picture may consist of at least one slice / tile. A tile is a rectangular area consisting of multiple CTUs within a specific tile column and tile row of a picture. A tile column is a rectangular area of CTUs with the same height as the picture and a width specified by the syntax requirements of the picture parameter set. A tile row is a rectangular area of CTUs with a height specified by the picture parameter set and the same width as the picture. CTUs within a tile can be arranged consecutively according to a CTU raster scan, while tiles within a picture can be arranged consecutively according to a tile raster scan. A slice may include an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture that can be exclusively included in a single NAL unit. A picture can also be divided into at least two sub-pictures. A sub-picture may be a rectangular area of at least one slice within a picture.
[0045] Pixel, pixel, or picture element can refer to the smallest unit that constitutes a picture (or image). In addition, "sample" can be used as a term corresponding to pixel. Sample can generally represent a pixel or pixel value, and can represent only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component.
[0046] A unit can represent the basic unit of image processing. A unit can include at least one of a specific region of a picture and information related to the corresponding region. A unit can include a luma block and two chroma (e.g., CB, CR) blocks. In some cases, the term "unit" can be used interchangeably with terms such as "block" or "region." In general, an MxN block can include a set (or array) of transform coefficients or samples (or sample arrays) consisting of M columns and N rows.
[0047] Here, "A or B" can mean "only A," "only B," or "both A and B." In other words, herein, "A or B" can be interpreted as "A and / or B." For example, herein, "A, B, or C" can mean "only A," "only B," "only C," or "any combination of A, B, and C."
[0048] As used herein, a slash mark ( / ) or a comma may mean "and / or." For example, "A / B" may mean "A and / or B." Thus, "A / B" may mean "only A," "only B," or "both A and B." For example, "A, B, C" may mean "A, B, or C."
[0049] Here, “at least one of A and B” may refer to “only A,” “only B,” or “both A and B.” In addition, herein, expressions such as “at least one of A or B” or “at least one of A and / or B” can be interpreted in the same manner as “at least one of A and B.”
[0050] In addition, herein, “at least one of A, B, and C” may refer to “only A,” “only B,” “only C,” or “any combination of A, B, and C.” In addition, “at least one of A, B, or C” or “at least one of A, B, and / or C” may refer to “at least one of A, B, and C.”
[0051] Additionally, parentheses used herein may refer to "for example." Specifically, when "prediction (intra-frame prediction)" is indicated, "intra-frame prediction" may be provided as an example of "prediction." In other words, "prediction" here is not limited to "intra-frame prediction," and "intra-frame prediction" may be provided as an example of "prediction." Furthermore, even when "prediction (i.e., intra-frame prediction)" is indicated, "intra-frame prediction" may be provided as an example of "prediction."
[0052] Here, technical features described individually in one drawing may be implemented individually or simultaneously.
[0053] Figure 1 A video / image coding system according to the present disclosure is shown.
[0054] refer to Figure 1 , a video / image coding system may include a first device (source device) and a second device (sink device).
[0055] The source device can transmit the encoded video / image information or data to the receiving device in the form of a file or stream transmission via a digital storage medium or a network. The source device may include a video source, an encoding device, and a transmitting unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be composed of a separate device or an external component.
[0056] A video source can obtain video / images through a process of capturing, synthesizing, or generating video / images. A video source can include both devices that capture video / images and devices that generate video / images. Devices that capture video / images may include at least one camera, a video / image archive containing previously captured video / images, and the like. Devices that generate video / images may include computers, tablets, smartphones, and the like, and can (electronically) generate video / images. For example, a virtual video / image can be generated by a computer, etc., and in this case, the process of capturing video / images can be replaced by a process of generating related data.
[0057] The encoding device can encode the input video / image. The encoding device can perform a series of processes such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0058] The transmitting unit can transmit the encoded video / image information or data output in the form of a bitstream to the receiving unit of the receiving device in the form of a file or stream transmission via a digital storage medium or network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include components for generating a media file in a predetermined file format and can also include components for transmitting via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to the decoding device.
[0059] The decoding device may decode the video / image by performing a series of processes such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.
[0060] The renderer may render the decoded video / image, and the rendered video / image may be displayed through a display unit.
[0061] Figure 2 A rough block diagram showing an encoding device to which an embodiment of the present disclosure can be applied and which performs encoding of a video / image signal is shown.
[0062] refer to Figure 2 The encoding device 200 may include an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, an inverse quantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the image segmenter 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or processor). Furthermore, the memory 270 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.
[0063] The image divider 210 may divide an input image (or picture, frame) input to the encoding apparatus 200 into at least one processing unit. For example, a processing unit may be referred to as a coding unit (CU). In this case, the coding unit may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) according to a quad-tree binary-tree ternary-tree (QTBTTT) structure.
[0064] For example, one coding unit may be split into a plurality of coding units having a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quadtree structure may be applied first, and the binary tree structure and / or the ternary structure may be applied later. Alternatively, the binary tree structure may be applied before the quadtree structure. The coding process according to this specification may be performed based on a final coding unit that is no longer split. In this case, based on coding efficiency according to image characteristics, etc., the maximum coding unit may be directly used as the final coding unit, or if necessary, the coding unit may be recursively split into coding units of a deeper depth, and the coding unit with the optimal size may be used as the final coding unit. Here, the coding process may include processes such as prediction, transformation, and reconstruction, which are described later.
[0065] As another example, a processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be divided or partitioned from the final coding unit. A prediction unit may be a unit for sample prediction, and a transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0066] In some cases, the term "unit" can be used interchangeably with terms such as "block" or "region." In general, an MxN block can represent a set of transform coefficients or samples consisting of M columns and N rows. A sample can generally represent a pixel or pixel value, and can represent only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component. A sample can be used as a term to refer to a picture (or image) corresponding to a pixel or picture element.
[0067] The encoding device 200 may subtract the prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 221 or the intra-frame predictor 222 from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual sample array), and the generated residual signal is sent to the transformer 232. In this case, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) within the encoding device 200 may be referred to as a subtractor 231.
[0068] The predictor 220 may perform prediction on a block to be processed (hereinafter referred to as a current block) and generate a predicted block including prediction samples for the current block. The predictor 220 may determine whether intra prediction or inter prediction is applied in units of the current block or CU. The predictor 220 may generate various information about the prediction, such as prediction mode information, and transmit it to the entropy encoder 240, as described later in the description of each prediction mode. The information about the prediction may be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0069] The intra-frame predictor 222 can predict the current block by referencing samples within the current picture. Depending on the prediction mode, the referenced samples can be located near the current block or can be located a certain distance away from the current block. In intra-frame prediction, the prediction mode may include at least one non-directional mode and multiple directional modes. The non-directional mode may include at least one of a DC mode or a planar mode. Depending on the level of detail of the prediction direction, the directional mode may include 33 directional modes or 65 directional modes. However, this is only an example, and more or fewer directional modes may be used depending on the configuration. The intra-frame predictor 222 may determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0070] The inter-frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information can include a motion vector and a reference picture index. The motion information can further include inter-frame prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in a reference picture. The reference picture containing the reference block and the reference picture containing the temporally neighboring block can be the same or different. Temporally neighboring blocks can be referred to as collocated reference blocks, collocated CUs (colCUs), etc., and the reference picture containing temporally neighboring blocks can be referred to as collocated pictures (colPics). For example, the inter-frame predictor 221 can configure a motion information candidate list based on the neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame predictor 221 can use the motion information of neighboring blocks as the motion information of the current block. In skip mode, unlike merge mode, a residual signal may not be transmitted. In motion vector prediction (MVP) mode, the motion vector of the neighboring block is used as the motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.
[0071] The predictor 220 can generate prediction signals based on various prediction methods described later. For example, the predictor can apply not only intra prediction or inter prediction to predict a block, but also both intra and inter prediction simultaneously. This is referred to as a combined inter and intra prediction (CIIP) mode. Alternatively, the predictor can be based on an intra block copy (IBC) prediction mode or a palette mode for block-specific prediction. The IBC prediction mode or palette mode can be used for content image / video coding, such as screen content coding (SCC), for gaming and the like. IBC essentially performs prediction within the current picture, but it can be performed similarly to inter prediction in that it derives reference blocks within the current picture. In other words, IBC can utilize at least one of the inter prediction techniques described herein. The palette mode can be considered an example of intra coding or intra prediction. When palette mode is applied, sample values within the picture can be signaled based on information about a palette table and a palette index. The prediction signal generated by the predictor 220 can be used to generate a reconstructed signal or a residual signal.
[0072] Transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, GBT refers to a transform obtained from a graph when relationship information between pixels is expressed as a graph. CNT refers to a transform obtained by generating a prediction signal using all previously reconstructed pixels. Furthermore, the transform process can be applied to square pixel blocks of the same size or to non-square blocks of variable size.
[0073] The quantizer 233 may quantize the transform coefficients and transmit them to the entropy encoder 240, and the entropy encoder 240 may encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantizer 233 may rearrange the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scanning order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0074] The entropy encoder 240 may perform various encoding methods such as Exponential Golomb, Context Adaptive Variable Length Coding (CAVLC), Context Adaptive Binary Arithmetic Coding (CABAC), etc. The entropy encoder 240 may encode information necessary for video / video image reconstruction (eg, values of syntax elements, etc.) in addition to transform coefficients quantized together or individually.
[0075] Encoded information (e.g., encoded video / image information) can be transmitted or stored in a bitstream in units of Network Abstraction Layer (NAL) units. The video / image information may further include information regarding various parameter sets, such as the Adaptation Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Furthermore, the video / image information may further include general constraint information. Information and / or syntax elements transmitted / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information may be encoded through the aforementioned encoding process and included in the bitstream. The bitstream may be transmitted over a network or stored in a digital storage medium. The network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit (not shown) for transmission and / or the storage unit (not shown) for storing the signal output from the entropy encoder 240 may be configured as internal or external components of the encoding device 200, or the transmission unit may also be included in the entropy encoder 240.
[0076] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, dequantization and inverse transform can be applied to the quantized transform coefficients via the inverse quantizer 234 and the inverse transformer 235 to reconstruct a residual signal (residual block or residual samples). The adder 250 can add the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, or reconstructed sample array). When there is no residual for the block to be processed, such as when skip mode is applied, the prediction block can be used as the reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed within the current picture and can also be used for inter-frame prediction of the next picture through filtering, which will be described later. Luma mapping with chroma scaling (LMCS) can also be applied during the picture encoding and / or reconstruction process.
[0077] The filter 260 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and the modified reconstructed picture can be stored in the memory 270, specifically in the DPB of the memory 270. Various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 260 can generate various information about filtering and send it to the entropy encoder 240. The information about filtering can be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0078] The modified reconstructed picture sent to the memory 270 may be used as a reference picture in the inter-frame predictor 221. When inter-frame prediction is applied therethrough, the encoding apparatus can avoid prediction mismatch in the encoding apparatus 200 and the decoding apparatus, and can also improve encoding efficiency.
[0079] The DPB of the memory 270 can store the modified reconstructed picture for use as a reference picture in the inter-frame predictor 221. The memory 270 can store motion information of the block from which the motion information in the current picture was derived (or encoded) and / or motion information of blocks in pre-reconstructed pictures. The stored motion information can be sent to the inter-frame predictor 221 to be used as motion information of spatially neighboring blocks or motion information of temporally neighboring blocks. The memory 270 can store reconstructed samples of the reconstructed blocks in the current picture and send them to the intra-frame predictor 222.
[0080] Figure 3 A rough block diagram showing a decoding device to which an embodiment of the present disclosure can be applied and which performs decoding of a video / image signal is shown.
[0081] refer to Figure 3 , the decoding apparatus 300 may be configured by including an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 332 and an intra-frame predictor 331. The residual processor 320 may include an inverse quantizer 321 and an inverse transformer 321.
[0082] According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 described above may be configured by a single hardware component (e.g., a decoder chipset or processor). Furthermore, the memory 360 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory 360 as an internal / external component.
[0083] When a bit stream including video / image information is input, the decoding apparatus 300 may generate a decoded image in response to the bit stream received in the decoded image. Figure 2 The image is reconstructed by processing video / image information in the encoding device of the decoding apparatus. For example, the decoding apparatus 300 can derive a unit / block based on relevant information of block segmentation obtained from the bitstream. The decoding apparatus 300 can perform decoding by using a processing unit applied in the encoding apparatus. Therefore, the processing unit of decoding can be a coding unit, and the coding unit can be divided from a coding tree unit or a maximum coding unit according to a quadtree structure, a binary tree structure and / or a ternary tree structure. At least one transform unit can be derived from the coding unit. And, the reconstructed image signal decoded and output by the decoding apparatus 300 can be played by a playback device.
[0084] The decoding device 300 may receive the data in the form of a bit stream from Figure 2 The received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information necessary for image reconstruction (or picture reconstruction) (e.g., video / image information). The video / image information may further include information about various parameter sets, such as the Adaptation Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Furthermore, the video / image information may further include general constraint information. The decoding device can further decode the picture based on the parameter set information and / or general constraint information. Signaled / received information and / or syntax elements described later herein can be decoded and obtained from the bitstream through a decoding process. For example, the entropy decoder 310 can decode the information in the bitstream based on coding methods such as Exponential Golomb coding, CAVLC, or CABAC, and output the values of the syntax elements necessary for image reconstruction and the quantized values of the transform coefficients of the residual. In more detail, the CABAC entropy decoding method receives bins corresponding to each syntax element from the bitstream, determines a context model using information about the syntax element to be decoded, decoded information about neighboring blocks and the block to be decoded, or information about symbols / bins decoded in a previous step, performs arithmetic decoding on the bins by predicting the probability of occurrence of the bins based on the determined context model, and generates symbols corresponding to the values of each syntax element. After determining the context model, the CABAC entropy decoding method can update the context model using information about the decoded symbols / bins for the context model for the next symbol / bin. Information decoded in the entropy decoder 310 regarding prediction is provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and residual values entropy-decoded in the entropy decoder 310, namely, quantized transform coefficients and related parameter information, are input to the residual processor 320. The residual processor 320 can derive a residual signal (residual block, residual sample, residual sample array). Furthermore, information regarding filtering in the information decoded in the entropy decoder 310 is provided to the filter 350. Meanwhile, a receiving unit (not shown) that receives a signal output from the encoding device may be further configured as an internal / external element of the decoding device 300 or the receiving unit may be a component of the entropy decoder 310 .
[0085] Meanwhile, the decoding device according to this specification may be referred to as a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoder 310, and the sample decoder may include at least one of an inverse quantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0086] The inverse quantizer 321 may inversely quantize the quantized transform coefficients and output the transform coefficients. The inverse quantizer 321 may rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding device. The inverse quantizer 321 may inversely quantize the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.
[0087] The inverse transformer 322 performs an inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0088] The predictor 320 may perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor 320 may determine whether to apply intra prediction or inter prediction to the current block based on the information on prediction output from the entropy decoder 310, and determine a specific intra / inter prediction mode.
[0089] The predictor 320 can generate a prediction signal based on various prediction methods described later. For example, the predictor 320 can not only apply intra prediction or inter prediction to predict a block, but can also apply both intra prediction and inter prediction simultaneously. This can be referred to as a combined inter and intra prediction (CIIP) mode. In addition, the predictor can be based on an intra block copy (IBC) prediction mode or a palette mode for block prediction. The IBC prediction mode or palette mode can be used for content image / video coding, such as screen content coding (SCC), for games and the like. IBC essentially performs prediction within the current picture, but it can be performed similarly to inter prediction in that it derives reference blocks within the current picture. In other words, IBC can use at least one of the inter prediction techniques described herein. The palette mode can be considered an example of intra coding or intra prediction. When the palette mode is applied, information about the palette table and palette index can be included in the video / image information and signaled.
[0090] The intra-frame predictor 331 can predict the current block by referencing samples within the current picture. Depending on the prediction mode, the referenced samples can be located near the current block or at a certain distance away from the current block. In intra-frame prediction, the prediction mode can include at least one non-directional mode and multiple directional modes. The intra-frame predictor 331 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0091] The inter-frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter-frame prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in the reference picture. For example, the inter-frame predictor 332 can configure a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index for the current block based on received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and prediction information can include information indicating the inter-frame prediction mode used for the current block.
[0092] The adder 340 may add the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including the inter-frame predictor 332 and / or the intra-frame predictor 331) to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). When there is no residual for the block to be processed, such as when skip mode is applied, the prediction block can be used as the reconstructed block.
[0093] Adder 340 may be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal may be used for intra-frame prediction of the next block to be processed in the current picture, may be output through filtering described later, or may be used for inter-frame prediction of the next picture. Luma mapping with chroma scaling (LMCS) may also be applied during picture decoding.
[0094] The filter 350 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and send the modified reconstructed picture to the memory 360, specifically the DPB of the memory 360. Various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0095] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter-frame predictor 332. The memory 360 can store the motion information of the block derived (or decoded) from the motion information in its current picture and / or the motion information of the block in the pre-reconstructed picture. The stored motion information can be sent to the inter-frame predictor 260 for use as the motion information of the spatially neighboring block or the motion information of the temporally neighboring block. The memory 360 can store the reconstructed samples of the reconstructed block in the current picture and send them to the intra-frame predictor 331.
[0096] Here, the embodiments described in the filter 260, the inter-frame predictor 221, and the intra-frame predictor 222 of the encoding device 200 may also be equally or correspondingly applied to the filter 350, the inter-frame predictor 332, and the intra-frame predictor 331 of the decoding device 300, respectively.
[0097] Figure 4 The diagram illustrates an image decoding method performed in an image decoding apparatus according to the present disclosure.
[0098] refer to Figure 4 , a prediction sample of the current block may be generated based on a predetermined inter prediction mode ( S400 ).
[0099] When coding a current block in merge mode, any one of a plurality of predefined inter prediction modes in an image decoding apparatus may be set as an inter prediction mode of the current block. In this case, the plurality of inter prediction modes may include at least one of a subblock merge mode, a merge mode with motion vector difference (MMVD), a normal merge mode, a combined inter and intra prediction (CIIP) mode, or a geometric partition merge mode (GPM mode).
[0100] Subblock merge mode divides the current block into multiple subblocks and derives motion information from spatially and temporally neighboring blocks on a per-subblock basis. Prediction samples for the current block are generated based on the derived motion information. Subblock merge mode can be adaptively used based on a merge_subblock_flag flag that indicates whether subblock-based motion information is derived from neighboring blocks of the current block.
[0101] MMVD mode derives motion information based on the regular merge mode, but modifies the motion vectors of the motion information based on the signaled motion vector difference (MVD). Prediction samples for the current block are generated based on the motion information. MMVD mode can be used adaptively based on an MMVD flag (mmvd_merge_flag) that indicates whether MMVD is used for the current block.
[0102] The CIIP mode can derive prediction samples based on a weighted sum of inter-frame prediction samples and intra-frame prediction samples. The CIIP mode can be adaptively used based on a CIIP flag (ciip_flag) indicating whether the CIIP mode is applied to the current block. The CIIP flag can be signaled via the bitstream. The CIIP flag can be derived as 1 when a predetermined condition is met, and can be derived as 0 otherwise. Here, the predetermined condition can be at least one of the following conditions:
[0103] - sps_ciip_enabled_flag is equal to 1. That is, CIIP mode is available.
[0104] - general_merge_flag is equal to 1. That is, the current block is a block compiled in merge mode.
[0105] - merge_subblock_flag is equal to 0. That is, the inter prediction mode of the current block is not subblock merge mode.
[0106] - regular_merge_flag is equal to 1. That is, the motion information of the current block is derived based on the regular merge mode or MMVD mode.
[0107] - cu_skip_flag is equal to 0. That is, the current block is not a block coded in skip mode.
[0108] - The width and height of the current block are less than 128.
[0109] - The product of the width and height of the current block is greater than or equal to 64.
[0110] GPM mode can be a geometric partition-based prediction mode and can be divided into GPM-INTRA-based merge mode and GPM-based merge mode. GPM-INTRA-based merge mode partitions the current block into two partitions (P0, P1) according to the geometric partition mode (GPM), derives inter-prediction blocks and intra-prediction blocks corresponding to the two partitions, and generates a prediction block for the current block based on a weighted sum of the inter-prediction blocks and the intra-prediction blocks. Here, the inter-prediction block can be derived based on the conventional merge mode. The intra-prediction block can be derived based on a predetermined intra-prediction mode, which can be set to one of one or more candidate modes in a candidate list. To this end, a mode index specifying one of the multiple candidate modes in the candidate list can be signaled. The one or more candidate modes can include at least one of a parallel mode parallel to the boundary between partitions in the current block, a perpendicular mode perpendicular to the boundary between the partitions, or a planar mode. GPM-INTRA-based merge mode can be adaptively used based on a flag (gpm_intra_flag) indicating whether to apply the GPM-INTRA-based merge mode to the current block.
[0111] The GPM-based merge mode partitions the current block into two partitions (P0, P1) according to the geometric partition mode (GPM), derives a first inter-frame prediction block and a second inter-frame prediction block corresponding to the two partitions, and generates a prediction block for the current block based on a weighted sum of the first and second inter-frame prediction blocks. The first and second inter-frame prediction blocks can be derived based on the conventional merge mode. The GPM-based merge mode can be adaptively used based on a flag (gpm_intra_flag) indicating whether the GPM-INTRA-based merge mode is applied to the current block. The GPM-based merge mode can be adaptively used based on a flag (MergeGpmFlag) indicating whether the GPM-based merge mode is applied to the current block. The MergeGpmFlag can be derived as 1 when a predetermined condition is met and as 0 otherwise. The predetermined condition can include at least one of the following conditions.
[0112] - sps_gpm_enabled_flag is equal to 1. That is, GPM-based merge mode is available.
[0113] - The slice type of the current block is B slice.
[0114] - general_merge_flag is equal to 1. That is, the current block is a block compiled in merge mode.
[0115] - The width and height of the current block is greater than or equal to 8.
[0116] - cbWidth is less than (8*cbHeight). Here, cbWidth and cbHeight represent the width and height of the current block respectively.
[0117] - cbHeight is less than (8*cbWidth).
[0118] - regular_merge_flag is equal to 0. That is, the motion information of the current block is not derived based on the regular merge mode or MMVD mode.
[0119] - merge_subblock_flag is equal to 0. That is, the inter prediction mode of the current block is not subblock merge mode.
[0120] - ciip_flag is equal to 0. That is, CIIP mode is not applied to the current block.
[0121] When the current block is coded in merge mode, it is possible to sequentially check which of the inter prediction modes corresponds to the inter prediction mode of the current block according to the priority order between the inter prediction modes. For example, first, it is possible to check whether the inter prediction mode of the current block is the sub-block merge mode. When merge_subblock_flag is equal to 1, the inter prediction mode of the current block can be determined as the sub-block merge mode. When merge_subblock_flag is equal to 0, it is possible to check whether the inter prediction mode of the current block is the regular merge mode based on regular_merge_flag. When regular_merge_flag is equal to 1, it is possible to check whether the inter prediction mode of the current block is the MMVD mode based on mmvd_merge_flag. When mmvd_merge_flag is equal to 1, the inter prediction mode of the current block can be determined as the MMVD mode. When mmvd_merge_flag is equal to 0, the inter prediction mode of the current block can be determined as the regular merge mode. When regular_merge_flag is equal to 0, it is possible to check whether the inter-frame prediction mode of the current block is CIIP mode based on ciip_flag. When ciip_flag is equal to 1, the inter-frame prediction mode of the current block can be determined as CIIP mode. When ciip_flag is equal to 0, it is possible to check whether the inter-frame prediction mode of the current block is GPM-INTRA-based merge mode based on gpm_intra_flag. When gpm_intra_flag is equal to 1, the inter-frame prediction mode of the current block is determined to be GPM-INTRA-based merge mode; and when gpm_intra_flag is equal to 0, the inter-frame prediction mode of the current block can be determined to be GPM-based merge mode.
[0122] When the current block is coded in skip mode, prediction samples for the current block may be generated based on any of the aforementioned inter prediction modes, and the generated prediction samples may be set as reconstructed samples. However, when the current block is coded in skip mode, application of the CIIP mode may be limited. That is, when the current block is coded in skip mode and the regular_merge_flag for the current block is equal to 0, the ciip_flag is not signaled and is derived as 0, and the inter prediction mode of the current block may be determined as either the GPM-based merge mode or the GPM-INTRA-based merge mode.
[0123] Alternatively, the restriction on applying the CIIP mode to the current block coded in skip mode is based on the assumption that the intra-frame predicted block according to the CIIP mode has a sufficient residual signal. If the compression efficiency of the CIIP mode is improved, forcing the unnecessary residual signal to be always transmitted may actually lead to a decrease in compression performance. Therefore, similar to the merge mode described above, the CIIP mode can be applied to the current block even if the current block is coded in skip mode.
[0124] For example, the regular_merge_flag may be signaled when a predetermined condition is met, and may not be signaled otherwise. Here, the predetermined condition may include at least one of the following conditions:
[0125] - The width and height of the current block are less than 128.
[0126] - sps_ciip_enabled_flag is equal to 1, regardless of cu_skip_flag, and the product of the width and height of the current block is greater than or equal to 64.
[0127] - sps_gpm_enabled_flag is equal to 1, the slice type of the current block is a B slice, the width (cbWidth) and height (cbHeight) of the current block are greater than or equal to 8, cbWidth is less than (8*cbHeight), and cbHeight is less than (8*cbWidth).
[0128] Regardless of whether the current block is coded in skip mode (i.e., cu_skip_flag), ciip_flag can be signaled when a predetermined condition is met. In this way, by skipping the signaling of residual information in CIIP mode, compression efficiency can be increased, and complexity can be reduced by reducing the number of condition checks in the signaling / parsing process of ciip_flag. The predetermined condition can be at least one of the following conditions:
[0129] - sps_ciip_enabled_flag is equal to 1.
[0130] - sps_gpm_enabled_flag is equal to 1.
[0131] - The slice type of the current block is B slice.
[0132] - The width (cbWidth) and height (cbHeight) of the current block are greater than or equal to 8.
[0133] - cbWidth is less than (8*cbHeight).
[0134] - cbHeight is less than (8*cbWidth).
[0135] - The width and height of the current block are less than 128.
[0136] When the predetermined condition is not met, ciip_flag may not be signaled. In this case, ciip_flag may be derived as 0 or 1 based on the above-mentioned derivation conditions of ciip_flag. However, the conditions related to cu_skip_flag may be excluded from the above-mentioned derivation conditions of ciip_flag.
[0137] The intra prediction sample according to the CIIP mode may be derived based on a predetermined intra prediction mode. The predetermined intra prediction mode may be at least one of the intra prediction modes predefined for the encoding apparatus and the decoding device. Figure 5 The predefined intra prediction modes and their prediction directions according to the present disclosure are exemplarily shown. Figure 5 , the predefined intra-frame prediction modes may include planar mode (mode 0), DC mode (mode 1), directional mode (modes 2 to 66) and wide-angle mode (modes -1 to -14, modes 67 to 80). One or more intra-frame prediction modes among the predefined intra-frame prediction modes can be used as candidate modes to construct a candidate list for the current block, and the intra-frame prediction mode of the current block is derived based on any one of the one or more candidate modes belonging to the candidate list. To this end, a mode index (ciip_mpm_idx) specifying any one of the multiple candidate modes belonging to the candidate list can be sent by signaling. The candidate list may include at least one of the planar mode, DC mode or directional mode. When the current block is a non-square block, the candidate list may further include at least one of the wide-angle modes.
[0138] Alternatively, the predefined intra-frame prediction modes may be divided into multiple groups. Based on a group index (ciip_group_idx) specifying one of the multiple groups, the group to which the intra-frame prediction mode of the current block belongs may be determined from among the multiple groups. One or more intra-frame prediction modes belonging to the determined group may be used as candidate modes to construct a candidate list. Based on a mode index specifying any one of the one or more candidate modes belonging to the candidate list, the intra-frame prediction mode of the current block may be determined.
[0139] For example, the predefined intra prediction modes may be divided into: a first group including at least one of a planar mode, a DC mode, or a wide-angle mode; a second group including at least one of directional modes with horizontal directivity (modes 2 to 34); and a third group including at least one of directional modes with vertical directivity (modes 35 to 66). A group index of 0 may specify the first group, a group index of 1 may specify the second group, and a group index of 2 may specify the third group.
[0140] When the CIIP mode is applied to the current block, a prediction sample of the current block may be generated based on a weighted sum of inter-prediction samples and intra-prediction samples. Here, the weight used for the weighted sum may be determined based on at least one of the intra-prediction mode of the current block, the position of the subregion to which the prediction sample belongs, or a weight index.
[0141] Weights can be applied equally to the entire inter-prediction block and intra-prediction block of the current block. Alternatively, the current block can be divided into multiple sub-regions. For example, the current block can be divided into multiple sub-regions of equal size based on symmetric partitioning. Alternatively, the current block can be divided into multiple sub-regions of different sizes based on asymmetric partitioning. The partition direction of the current block can be determined based on signaled information indicating the partition direction, or based on predetermined coding parameters. Coding parameters can include at least one of the intra-prediction mode of the current block (e.g., mode value, whether it is non-directional mode, directionality, angle, etc.) or the position of a neighboring block coded using intra prediction. The weight applied to any one of the multiple sub-regions can be different from the weight applied to at least one other sub-region. To this end, at least two weights can be determined for the current block. A weight can be applied to any one of the multiple sub-regions, while no weight is applied to at least one other sub-region. Hereinafter, weights may be expressed as (w0, w1), where w0 may refer to a first weight applied to intra-prediction samples and w1 may refer to a second weight applied to inter-prediction samples.
[0142] Specifically, the weight used for the weighted sum may be determined based on the intra prediction mode of the current block (eg, the value of the intra prediction mode, the directionality of the intra prediction mode, or the range to which the intra prediction mode belongs).
[0143] For example, the weight may be determined as at least one of a plurality of predefined weight candidates. Here, the plurality of weight candidates may include at least one of (3, 1), (2, 2), or (1, 3). Alternatively, the plurality of weight candidates may include at least one of (7, 1), (6, 2), (5, 3), (4, 4), (3, 5), (6, 2), or (1, 7).
[0144] When the intra prediction mode of the current block is a non-directional mode (eg, planar mode or DC mode), the same weight may be applied to the entire region of the current block. In this case, the weight may be determined as one of the above-mentioned multiple weight candidates.
[0145] When the intra-frame prediction mode of the current block is a horizontal directivity mode, the left sub-region and the right sub-region within the current block may have different weights. Here, the horizontal directivity mode may be defined as a mode within the range of modes 2 to 34. Alternatively, the horizontal directivity mode may be defined as a mode within the range of modes 13 to 23. When the intra-frame prediction mode of the current block is a vertical directivity mode, the upper sub-region and the lower sub-region within the current block may have different weights. Here, the vertical directivity mode may be defined as a mode within the range of modes 35 to 66. Alternatively, the vertical directivity mode may be defined as a mode within the range of modes 45 to 55.
[0146] When the intra prediction mode of the current block is the wide-angle mode, the same weight may be applied to the entire area of the current block. In this case, the weight may be determined as one of the above-mentioned multiple weight candidates.
[0147] The weight may be determined based on a signaled weight index. For example, when CIIP mode is applicable (AvailableCiip = 1), ciip_flag may be signaled. When ciip_flag is 1, at least one of the merge index (merge_idx) for inter-frame prediction, the mode index (ciip_mpm_idx) for intra-frame prediction, or the weight index (weight_idx) for the weighted sum may be signaled. When the signaling conditions for ciip_flag are met, the variable AvailableCiip may be set to 1, and otherwise set to 0.
[0148] For example, a weight index of 0 may indicate (2, 2), a weight index of 1 may indicate (3, 1), and a weight index of 2 may indicate (1, 3).
[0149] Alternatively, the weight may be a weight applied to a specific area within the current block, as shown in the following Table 1. That is, depending on the intra prediction mode of the current block, the area in which the weighted sum between inter prediction samples and intra prediction samples is performed may vary.
[0150] [Table 1]
[0151]
[0152] In Table 1, left mode may refer to a horizontal mode (i.e., mode 18) or a horizontal directivity mode. The horizontal directivity mode may be defined as a mode within the range of modes 2 to 34, or may be defined as a mode within the range of modes 13 to 23. Up mode may refer to a vertical mode (i.e., mode 50) or a vertical directivity mode. The vertical directivity mode may be defined as a mode within the range of modes 35 to 66, or may be defined as a mode within the range of modes 45 to 55. (Left mode! & Up mode!) may represent modes other than the left mode and the up mode. (Left mode! & Up mode!) may include non-directional modes such as a planar mode and / or a DC mode. (Left mode! & Up mode!) may include a wide-angle mode as a directional mode.
[0153] In Table 1, the entire area may indicate that the weight is applied to the entire area of the current block. The left area may indicate that the weight is applied to the left sub-area within the current block. The upper area may indicate that the weight is applied to the upper sub-area within the current block. For example, the current block may be divided into two sub-areas in the vertical direction, and in this case, the left sub-area of the two sub-areas may correspond to the left area. Similarly, the current block may be divided into two sub-areas in the horizontal direction, and in this case, the upper sub-area of the two sub-areas may correspond to the upper area. Hereinafter, the sub-area to which the weight is applied (i.e., the left area or the upper area) of the two sub-areas is referred to as sub-area 0 (R0), and the other sub-area is referred to as sub-area 1 (R1).
[0154] For sub-region 0 (R0), weighted summation can be performed based on the weight indicated by the weight index (weight_idx). For sub-region 1 (R1), predefined prediction samples can be set as the prediction samples of the current block as is. Here, the predefined prediction samples can be inter-frame prediction samples or intra-frame prediction samples. For example, the prediction samples for each sub-region can be derived as shown in the following equation 1.
[0155] [Formula 1]
[0156] If (x, y) ∈ R0, then P(x, y) = (w0 * P Intra (x, y) + w1 * P Inter(x, y)) / (w0+w1)
[0157] Otherwise (if (x, y) ∈ R1), then P(x, y) = P Inter (x, y)
[0158] Alternatively, a weighted summation may be performed on sub-region 0 (R0) based on a weight indicated by a weight index (weight_idx). A weighted summation may be performed on sub-region 1 (R1) based on a predetermined weight. Here, the predetermined weight may be a weight derived based on the weight indicated by the weight index (weight_idx). The predetermined weight may be derived by adding or subtracting a predetermined offset (α) from the weight indicated by the weight index (weight_idx). α may be an integer greater than or equal to -w0 and less than or equal to w1. For example, the predicted sample for each sub-region may be derived as shown in the following Equation 2.
[0159] [Formula 2]
[0160] If (x, y) ∈ R0, then P(x, y) = (w0 * P Intra (x, y) + w1 * P Inter (x, y)) / (w0+w1)
[0161] Otherwise (if (x, y) ∈ R1), then P(x, y) = ((w0 + α) * P Intra (x, y) + (w1 - α) * P Inter (x, y)) / (w0 + w1)
[0162] Alternatively, the current block can be divided into four sub-regions. These four sub-regions can be represented as sub-regions 0 to 3, i.e., R0 to R3. The current block can be divided horizontally into four sub-regions based on three horizontal lines. The current block can be divided vertically into four sub-regions based on four vertical lines. The current block can be divided into four sub-regions by one horizontal line and one vertical line passing through the center of the current block.
[0163] For sub-region 0 (R0), a weighted sum can be performed based on the weight indicated by the weight index (weight_idx). For sub-regions 1 to 3 (R1 to R3), a weighted sum can be performed based on predetermined weights. Here, the weight for sub-region 1 can be derived by adding / subtracting a predetermined offset (α) to the weight indicated by the weight index (weight_idx). The weight for sub-region 2 can be derived by adding / subtracting a predetermined offset (β) to the weight indicated by the weight index (weight_idx). The weight for sub-region 3 can be derived by adding / subtracting a predetermined offset (δ) to the weight indicated by the weight index (weight_idx). The offsets (α, β, δ) are integers greater than or equal to -w0 and less than or equal to w1, and cannot be 0. α, β, and δ can be different values. For example, the predicted samples for each sub-region can be derived as in the following equation 3.
[0164] [Formula 3]
[0165] If (x, y) ∈ R0, then P(x, y) = (w0 * P Intra (x, y) + w1 * PInter (x, y)) / (w0+ w1)
[0166] Otherwise (if (x, y) ∈ R1), then P(x, y) = ((w0 + α) * P Intra (x, y) + (w1 - α) * P Inter (x, y)) / (w0 + w1)
[0167] Otherwise (if (x, y) ∈ R2), then P(x, y) = ((w0 + β) * P Intra (x, y) + (w1 - β) *P Inter (x, y)) / (w0 + w1)
[0168] Otherwise (if (x, y) ∈ R3), then P(x, y) = ((w0 + δ) * P Intra (x, y) + (w1 – δ) *P Inter (x, y)) / (w0 + w1)
[0169] Alternatively, because directionality-based intra prediction utilizes the spatial proximity of samples, weights can be increased uniformly as the distance from the reference sample increases. Specifically, for sub-region 0 (R0), weighted summation can be performed based on the weight indicated by the weight index (weight_idx). For sub-regions 1 to 3 (R1 to R3), weighted summation can be performed based on predetermined weights. Here, the weight for sub-region 1 can be derived by adding / subtracting a predetermined offset (α) from the weight indicated by the weight index (weight_idx). The weight for sub-region 2 can be derived by adding / subtracting a predetermined offset (2α) from the weight indicated by the weight index (weight_idx). The weight for sub-region 3 can be derived by adding / subtracting a predetermined offset (3α) from the weight indicated by the weight index (weight_idx). α is an integer greater than or equal to -w0 and less than or equal to w1, and cannot be 0. For example, the predicted samples for each sub-region can be derived as shown in the following equation 4.
[0170] [Formula 4]
[0171] [If (x, y) ∈ R0, then P(x, y) = (w0 * P Intra (x, y) + w1 * P Inter (x, y)) / (w0+w1)
[0172] Otherwise (if (x, y) ∈ R1), then P(x, y) = ((w0 + α) * P Intra (x, y) + (w1 – α) *P Inter (x, y)) / (w0 + w1)
[0173] Otherwise (if (x, y) ∈ R2), then P(x, y) = ((w0 + 2α) * P Intra (x, y) + (w1 – 2α)*P Inter (x, y)) / (w0 + w1)
[0174] Otherwise (if (x, y) ∈ R3), then P(x, y) = ((w0 + 3α) * P Intra (x, y) + (w1 – 3α)*P Inter (x, y)) / (w0 + w1)
[0175] The information about the GPM according to the present disclosure may include at least one of information for a GPM-INTRA-based merge mode or information for a GPM-based merge mode. The information about the GPM may include at least one of a partition index (merge_gpm_partition_idx) specifying a partition type or partition shape of the GPM, a flag (gpm_intra_flag0) indicating whether a first partition (P0) of a current block is predicted based on an intra mode, a mode index (gpm_intra_mode_idx0) specifying an intra prediction mode for the first partition from a candidate list, a merge index (merge_gpm_idx0) specifying a merge candidate for the first partition of the current block from a merge candidate list, a flag (gpm_intra_flag1) indicating whether a second partition (P1) of the current block is predicted based on an intra mode, a mode index (gpm_intra_mode_idx1) specifying an intra prediction mode for the second partition from a candidate list, or a merge index (merge_gpm_idx1) specifying a merge candidate for the second partition of the current block from a merge candidate list.
[0176] For example, gpm_intra_flag0 being 1 may indicate that the first partition is predicted based on intra mode. gpm_intra_flag0 being 0 may indicate that the first partition is not predicted based on intra mode. Alternatively, gpm_intra_flag0 being 0 may indicate that the motion information of the first partition is derived based on merge mode. Therefore, when gpm_intra_flag0 is equal to 1, gpm_intra_mode_idx0 may be signaled. An intra prediction mode for the first partition may be derived based on gpm_intra_mode_idx0, and an intra prediction block corresponding to the first partition may be generated based on the derived intra prediction mode. On the other hand, when gpm_intra_flag0 is equal to 0, merge_gpm_idx0 may be signaled. A merge candidate for the first partition may be determined from the merge candidate list of the current block based on merge_gpm_idx0, and motion information for the first partition may be derived based on the motion information of the determined merge candidate. An inter prediction block corresponding to the first partition may be generated based on the derived motion information.
[0177] gpm_intra_flag1 being 1 may indicate that the second partition is predicted based on intra mode. gpm_intra_flag1 being 0 may indicate that the second partition is not predicted based on intra mode. Alternatively, gpm_intra_flag1 being 0 may indicate that the motion information of the second partition is derived based on merge mode. When gpm_intra_flag0 is equal to 0, gpm_intra_flag1 may be signaled.
[0178] When gpm_intra_flag1 is equal to 1, gpm_intra_mode_idx1 may be signaled. An intra prediction mode for the second partition may be derived based on gpm_intra_mode_idx1, and an intra-prediction block corresponding to the second partition may be generated based on the derived intra prediction mode. On the other hand, when gpm_intra_flag1 is equal to 0, merge_gpm_idx1 may be signaled. A merge candidate for the second partition may be determined from a merge candidate list of the current block based on merge_gpm_idx1, and motion information for the second partition may be derived based on the motion information of the determined merge candidate. An inter-prediction block corresponding to the second partition may be generated based on the derived motion information.
[0179] CIIP mode generates a final prediction block by weighting the intra-prediction block and the inter-prediction block. GPM-INTRA-based merge mode divides the current block into two partitions and generates a prediction block for one of them based on intra mode. They both generate intra-prediction blocks. The present disclosure relates to a method for signaling residual information for CIIP mode and GPM-INTRA-based merge mode, as well as a method for coordinating the two.
[0180] According to the above example, when the current block is not encoded in skip mode, CIIP mode may be allowed. That is, when CIIP mode is applied to the current block, the signaling of residual information for the current block is not skipped. Similarly, when the current block is not encoded in skip mode, the GPM-INTRA-based merge mode may be restricted. Conversely, when the current block is encoded in skip mode, the GPM-INTRA-based merge mode may not be allowed.
[0181] To this end, at least one of the aforementioned information about the GPM may be adaptively signaled based on a flag (cu_skip_flag) indicating whether the current block is a block coded in skip mode. For example, when cu_skip_flag is equal to 0, gpm_intra_flag0 may be signaled; and when cu_skip_flag is equal to 1, gpm_intra_flag0 may not be signaled. When both gpm_intra_flag0 and cu_skip_flag are equal to 0, gpm_intra_flag1 may be signaled; and when at least one of gpm_intra_flag0 or cu_skip_flag is equal to 1, gpm_intra_flag1 may not be signaled.
[0182] refer to Figure 4 , residual samples of the current block may be derived based on information about sub-block transform (SBT) ( S410 ).
[0183] Sub-block transform can divide the current block into multiple sub-blocks, perform residual coding on some of the multiple sub-blocks, and skip residual coding on the remaining sub-blocks. Here, residual coding can include at least one of the following: 1) deriving transform coefficients based on residual information; 2) dequantizing the transform coefficients; 3) inverse transforming the transform coefficients.
[0184] For example, when a sub-block transform is used for a current block, the current block may be divided into two sub-blocks. Residual coding may be performed on one of the two sub-blocks to derive residual samples. Residual coding may be skipped for the other of the two sub-blocks, and the residual samples of the corresponding sub-block may be derived as 0.
[0185] The information about sub-block transformation according to the present disclosure may include at least one of an SBT flag (cu_sbt_flag), an SBT size flag (cu_sbt_quad_flag), an SBT direction flag (cu_sbt_horizontal_flag), or an SBT position flag (cu_sbt_pos_flag).
[0186] Specifically, the SBT flag (cu_sbt_flag) may indicate whether sub-block transform is used for the current block. For example, an SBT flag of 1 may indicate that sub-block transform is used for the current block, and an SBT flag of 0 may indicate that sub-block transform is not used for the current block.
[0187] When a predetermined first condition is met, the SBT flag may be signaled and otherwise set to 0. The predetermined first condition may include at least one of the following conditions:
[0188] - sps_sbt_enabled_flag is equal to 1. That is, sub-block transform is available.
[0189] - cbWidth is less than or equal to MaxTbSizeY. That is, the width of the current block (cbWidth) is less than or equal to the maximum transform block size (MaxTbSizeY).
[0190] - cbHeight is less than or equal to MaxTbSizeY. That is, the height of the current block (cbHeight) is less than or equal to the maximum transform block size (MaxTbSizeY).
[0191] - The width and height of the current block is greater than or equal to 8.
[0192] According to the above conditions, the SBT flag can be signaled regardless of ciip_flag.By allowing sub-block transform in CIIP mode, signaling of residual information for some sub-blocks in the current block can be omitted, and compression performance can be improved.
[0193] Alternatively, the SBT flag may be signaled when a predetermined second condition is met, otherwise it may be derived as 0. The predetermined second condition may include at least one of the following conditions:
[0194] - sps_sbt_enabled_flag is equal to 1. That is, sub-block transform is available.
[0195] - ciip_flag is equal to 0. That is, CIIP mode is not applied to the current block.
[0196] - cbWidth is less than or equal to MaxTbSizeY. That is, the width of the current block (cbWidth) is less than or equal to the maximum transform block size (MaxTbSizeY).
[0197] - cbHeight is less than or equal to MaxTbSizeY. That is, the height of the current block (cbHeight) is less than or equal to the maximum transform block size (MaxTbSizeY).
[0198] - The width and height of the current block is greater than or equal to 8.
[0199] According to the above conditions, regardless of the flag (gpm_intra_flag) indicating whether the GPM-Intra based merge mode is applied to the current block, the SBT flag can be signaled. By allowing sub-block transforms in the GPM-Intra based merge mode, the signaling of residual information for some sub-blocks in the current block can be omitted, thereby improving compression performance.
[0200] The SBT size flag (cu_sbt_quad_flag) may indicate the size of the sub-block on which residual coding is performed. For example, an SBT size flag of 1 may indicate that residual coding is performed on a sub-block of 1 / 4 the current block size, and an SBT size flag of 0 may indicate that residual coding is performed on a sub-block of 1 / 2 the current block size.
[0201] When a predetermined condition is met, the SBT size flag may be signaled, otherwise it may be set to 0. The predetermined condition may include at least one of the following conditions:
[0202] - The width or height of the current block is greater than or equal to 8.
[0203] - The width or height of the current block is greater than or equal to 16.
[0204] The SBT direction flag (cu_sbt_horizontal_flag) may indicate whether the current block is split horizontally. For example, an SBT direction flag of 1 may indicate that the current block is split into two sub-blocks in the horizontal direction, and an SBT direction flag of 0 may indicate that the current block is split into two sub-blocks in the vertical direction.
[0205] When a predetermined condition is met, the SBT direction flag may be signaled. The predetermined condition may be any one of the following:
[0206] - cu_sbt_quad_flag is equal to 1, and the width and height of the current block are greater than or equal to 16.
[0207] - cu_sbt_quad_flag is equal to 0, and the width and height of the current block are greater than or equal to 8.
[0208] If the above conditions are not met, the SBT direction flag may not be signaled. However, the SBT direction flag may be derived based on the SBT size flag and / or the width of the current block. For example, if cu_sbt_quad_flag is equal to 1, cu_sbt_horizontal_flag may be derived as 1 when the width of the current block is greater than or equal to 16, otherwise (i.e., when the width of the current block is less than 16), cu_sbt_horizontal_flag may be derived as 0. If cu_sbt_quad_flag is equal to 0, cu_sbt_horizontal_flag may be derived as 1 when the width of the current block is greater than or equal to 8, otherwise (i.e., when the width of the current block is less than 8), cu_sbt_horizontal_flag may be derived as 0.
[0209] The SBT position flag (cu_sbt_pos_flag) may indicate the position of the sub-block on which residual coding is performed. For example, an SBT position flag of 1 may indicate that residual coding is performed on the second sub-block of the two sub-blocks in the current block, and an SBT position flag of 0 may indicate that residual coding is performed on the first sub-block of the two sub-blocks. Depending on the partition direction of the current block, the second sub-block may refer to the right sub-block or the lower sub-block, and the first sub-block may refer to the left sub-block or the upper sub-block.
[0210] By using the characteristics of the CIIP mode, all or part of the information about the sub-block transform can be derived, thereby reducing the number of bits required to signal the information about the sub-block transform. Intra-frame prediction has the characteristic that the prediction performance improves as the distance from the reference sample approaches. Therefore, all or part of the information about the sub-block transform can be derived by considering the directionality of the intra-frame prediction mode used for the current block. Depending on the intra-frame prediction mode of the current block, cu_sbt_flag can be derived as a specific value and not be signaled. Depending on the intra-frame prediction mode of the current block, cu_sbt_quad_flag can be derived as a specific value and not be signaled. Depending on the intra-frame prediction mode of the current block, cu_sbt_horizontal_flag can be derived as a specific value and not be signaled. Depending on the intra-frame prediction mode of the current block, cu_sbt_pos_flag can be derived as a specific value and not be signaled.
[0211] As an example, Table 2 shows an example in which characteristics cu_sbt_flag, cu_sbt_quad_flag, cu_sbt_horizontal_flag, and cu_sbt_pos_flag are signaled or mapped to specific values depending on the intra prediction mode for the CIIP mode.
[0212] [Table 2]
[0213]
[0214] Referring to Table 2, when the intra prediction mode of the current block corresponds to the non-directional mode or the wide-angle mode, this may mean that the similarity between the current block and the adjacent samples is relatively low. In this case, the cu_sbt_flag for the current block may not be sent with a signal and may be derived as 0. When the intra prediction mode of the current block belongs to the first range, this may mean that the similarity between the current block and the adjacent samples is relatively low. For example, the first range may include at least one of the range of modes 2 to 12, the range of modes 24 to 44, or the range of modes 56 to 80. In this case, the cu_sbt_flag for the current block may not be sent with a signal and may be derived as 0. When the intra prediction mode of the current block belongs to the second range, the cu_sbt_flag for the current block may not be sent with a signal and may be derived as 1. For example, the second range may include at least one of the range of modes 13 to 23 or the range of modes 45 to 55. Additionally, when the intra prediction mode of the current block belongs to the second range, cu_sbt_quad_flag for the current block may not be signaled and may be derived as 0. When the intra prediction mode of the current block corresponds to a horizontal directionality mode (e.g., when the intra prediction mode of the current block belongs to the range of modes 13 to 23), cu_sbt_horizontal_flag for the current block may be derived as 0. When the intra prediction mode of the current block corresponds to a vertical directionality mode (e.g., when the intra prediction mode of the current block belongs to the range of modes 45 to 55), cu_sbt_horizontal_flag for the current block may be derived as 1. When the intra prediction mode of the current block belongs to the second range, cu_sbt_pos_flag for the current block may not be signaled and may be derived as 1.
[0215] In other words, when the current block has a horizontal or vertical directivity mode, it may be determined that a sub-block transform is applied to the current block. When the current block has a horizontal or vertical directivity mode, it may be determined that the current block is split into two sub-blocks having the same size. When the current block has a horizontal directivity mode, it may be determined that the current block is split in the vertical direction taking into account distances of adjacent samples of the current block. When the current block has a vertical directivity mode, it may be determined that the current block is split in the horizontal direction taking into account distances of adjacent samples of the current block. In this case, it may be determined that residual information of a first sub-block having a short distance to a reference sample among the two sub-blocks is not signaled, and residual information of a second sub-block having a long distance to the reference sample is signaled.
[0216] Table 3 shows how to signal / derive information about subblock transform based on the range to which the intra prediction mode of the current block belongs.
[0217] [Table 3]
[0218]
[0219] Referring to Table 3, the predefined intra-frame prediction modes can be divided into three ranges. Specifically, the predefined intra-frame prediction modes can be divided into: a first range, which includes modes 2 to 34; a second range, which includes modes 34 to 66; and a third range, which includes at least one of a non-directional mode or a wide-angle mode. When the intra-frame prediction mode of the current block corresponds to a directional mode (for example, the intra-frame prediction mode of the current block belongs to the first range or the second range), the cu_sbt_flag for the current block can be derived as 1 and not signaled. When the intra-frame prediction mode of the current block corresponds to a directional mode (for example, the intra-frame prediction mode of the current block belongs to the first range or the second range), the cu_sbt_quad_flag for the current block can be derived as 1 and not signaled. When the intra-frame prediction mode of the current block corresponds to a horizontal directionality mode (for example, when the intra-frame prediction mode of the current block belongs to the first range), the cu_sbt_horizontal_flag for the current block can be derived as 0. When the intra prediction mode of the current block corresponds to the vertical directionality mode (for example, when the intra prediction mode of the current block belongs to the second range), cu_sbt_horizontal_flag for the current block may be derived as 1. When the intra prediction mode of the current block corresponds to the directional mode (for example, when the intra prediction mode of the current block belongs to the first or second range), cu_sbt_pos_flag for the current block may be derived as 1 and not signaled. On the other hand, when the intra prediction mode of the current block corresponds to the non-directional mode or the wide-angle mode (for example, when the intra prediction mode of the current block belongs to the third range), information about the sub-block transform for the current block may be signaled via the bitstream.
[0220] In other words, when the current block has a horizontal or vertical directivity mode, it may be determined that a sub-block transform is applied to the current block. When the current block has a horizontal or vertical directivity mode, it may be determined that the current block is split into two sub-blocks having sizes of 1 / 4 and 3 / 4 of the current block. When the current block has a horizontal directivity mode, it may be determined that the current block is split in the vertical direction taking into account distances of adjacent samples of the current block. When the current block has a vertical directivity mode, it may be determined that the current block is split in the horizontal direction taking into account distances of adjacent samples of the current block. In this case, it may be determined that residual information of a first sub-block having a short distance to a reference sample among the two sub-blocks is not signaled, and residual information of a second sub-block having a long distance to the reference sample is signaled.
[0221] Table 4 shows how to signal some information about sub-block transforms and derive other information based on the range to which the intra prediction mode of the current block belongs.
[0222] [Table 4]
[0223]
[0224] Referring to Table 4, the predefined intra prediction modes can be divided into three ranges. Specifically, the predefined intra prediction modes can be divided into: a first range including modes 2 to 34; a second range including modes 34 to 66; and a third range including at least one of a non-directional mode or a wide-angle mode. When the intra prediction mode of the current block corresponds to a directional mode (e.g., the intra prediction mode of the current block belongs to the first range or the second range), cu_sbt_flag for the current block can be derived as 1 and not signaled. On the other hand, when the intra prediction mode of the current block does not correspond to a directional mode (e.g., when the intra prediction mode of the current block belongs to the third range), cu_sbt_flag for the current block may not be signaled and may be derived as 0. cu_sbt_quad_flag can be signaled regardless of the intra prediction mode of the current block (or the range to which the intra prediction mode of the current block belongs). When the intra prediction mode of the current block corresponds to the horizontal directionality mode (e.g., the intra prediction mode of the current block belongs to the first range), cu_sbt_horizontal_flag for the current block may be derived as 0. When the intra prediction mode of the current block corresponds to the vertical directionality mode (e.g., the intra prediction mode of the current block belongs to the second range), cu_sbt_horizontal_flag for the current block may be derived as 1. cu_sbt_pos_flag may be signaled regardless of the intra prediction mode of the current block (or the range to which the intra prediction mode of the current block belongs).
[0225] In other words, when the current block has a directional mode, it can be determined that the sub-block transform is applied to the current block, and otherwise, it can be determined that the sub-block transform is not applied to the current block. When the current block has a horizontal directional mode, the current block can be divided in the vertical direction in consideration of the distance of the adjacent samples of the current block. When the current block has a vertical directional mode, the current block can be divided in the horizontal direction in consideration of the distance of the adjacent samples of the current block.
[0226] The above-described method for signaling / deriving information about subblock transformation can be applied when the intra prediction mode of the current block is derived before signaling information about subblock transformation. Alternatively, even when the intra prediction mode of the current block is not derived before signaling information about subblock transformation, the above-described method can be applied when the range to which the intra prediction mode of the current block belongs can be derived using a simple prediction method.
[0227] The simple prediction method may specify the range to which the intra prediction mode of the current block belongs, the directionality of the intra prediction mode of the current block, etc. based on the gradient (or gradient change, histogram) between reference samples in the neighboring area (or template area) of the current block. Alternatively, the simple prediction method may specify the range to which the intra prediction mode of the current block belongs, the directionality of the intra prediction mode of the current block, etc. based on statistical information (e.g., mode, minimum value, maximum value, average value, median value, etc.) about the intra prediction modes in the neighboring area of the current block.
[0228] The above method is merely an example, and the type of information about the derived sub-block transform, the derived specific value, etc. may be changed. In addition, the range of dividing the predefined intra prediction mode may be defined in a more detailed manner, and to this end, the size and / or shape of the current block may be considered.
[0229] When the CIIP mode is applied to the current block, at least one of the information about the sub-block transformation can be derived as a specific value based on at least one of the weights (or weight indices) in the aforementioned CIIP mode, the intra-frame prediction mode, the area range of applying the weights, or the position of the sub-area to which the weights are applied, and is not sent using a signal.
[0230] Specifically, depending on the weights for the weighted sum in CIIP mode, cu_sbt_flag may be derived as a specific value and not signaled. Depending on the weights for the weighted sum in CIIP mode, cu_sbt_quad_flag may be derived as a specific value and not signaled. Depending on the weights for the weighted sum in CIIP mode, cu_sbt_horizontal_flag may be derived as a specific value and not signaled. Depending on the weights for the weighted sum in CIIP mode, cu_sbt_pos_flag may be derived as a specific value and not signaled. Depending on the scope of application of the signaled / derived weights for the weighted sum in intra and inter modes, cu_sbt_flag may be derived as a specific value and not signaled. Depending on the scope of the area to which the weights for the weighted sum in CIIP mode are applied, cu_sbt_quad_flag may be derived as a specific value and not signaled. Depending on the extent of the area to which the weights for the weighted sum in CIIP mode are applied, cu_sbt_horizontal_flag may be derived as a specific value and not signaled. Depending on the extent of the area to which the weights for the weighted sum in CIIP mode are applied, cu_sbt_pos_flag may be derived as a specific value and not signaled. Depending on the intra-prediction mode used for CIIP mode, cu_sbt_flag may be derived as a specific value and not signaled. Depending on the intra-prediction mode used for CIIP mode, cu_sbt_quad_flag may be derived as a specific value and not signaled. Depending on the intra-prediction mode used for CIIP mode, cu_sbt_horizontal_flag may be derived as a specific value and not signaled. Depending on the intra-prediction mode used for CIIP mode, cu_sbt_pos_flag may be derived as a specific value and not signaled.
[0231] As an example, Table 5 is an example in which all or part of the information about the subblock transform is derived as a specific value based on at least one of the weights (or weight indices) for the weighted sum in the CIIP mode or the intra prediction mode.
[0232] [Table 5]
[0233]
[0234] In Table 5, the left mode, the up mode, (left mode! and up mode!) are the same as those in Table 1, and repeated explanation will be omitted here.
[0235] Based on the weight index (weight_idx), it can be determined whether a sub-block transform is applied to the current block. For example, when the weight index is equal to 0 or 2, it can be determined that a sub-block transform is applied to the current block. In this case, cu_sbt_flag can be derived as 1 and not signaled. When the weight index is equal to 1, it can be determined that a sub-block transform is not applied to the current block. In this case, cu_sbt_flag can be derived as 0 and not signaled. Alternatively, when the first weight (w0) applied to intra-predicted samples is less than or equal to the second weight (w1) applied to inter-predicted samples, it can be determined that a sub-block transform is applied to the current block. In this case, cu_sbt_flag can be derived as 1 and not signaled. When the first weight (w0) applied to intra-predicted samples is greater than the second weight (w1) applied to inter-predicted samples, it can be determined that a sub-block transform is not applied to the current block. In this case, cu_sbt_flag can be derived as 0 and not signaled.
[0236] When it is determined that a sub-block transform is applied to the current block, at least one of cu_sbt_quad_flag, cu_sbt_horizontal_flag, or cu_sbt_pos_flag may be derived as a specific value based on the intra prediction mode of the current block. The SBT information may be set to a specific value according to the classification of each intra mode. For example, when the intra prediction mode of the current block does not correspond to left mode or top mode, this may mean that the similarity between the current block and adjacent samples is relatively low. In this case, it may be determined that the current block is divided into two sub-blocks having sizes of 1 / 4 and 3 / 4 of the current block, and cu_sbt_quad_flag may not be signaled and derived as 1. On the other hand, when the intra prediction mode of the current block corresponds to left mode or top mode, it may be determined that the current block is divided into two sub-blocks having the same size, and cu_sbt_quad_flag may be derived as 0 without being signaled. Here, the case where the intra prediction mode of the current block does not correspond to the left mode and the up mode may mean that the intra prediction mode of the current block corresponds to a non-directional mode, such as a planar mode or a DC mode. The case where the intra prediction mode of the current block does not correspond to the left mode and the up mode may mean that the intra prediction mode of the current block belongs to a wide-angle mode (modes -14 to -1, modes 67 to 80). The case where the intra prediction mode of the current block does not correspond to the left mode and the up mode may mean that it belongs to a predetermined directional mode (for example, modes 2 to 12, modes 24 to 44, modes 56 to 66).
[0237] The partition direction of the current block can be determined based on the distance to the reference samples used for intra prediction of the current block (or the position of the reference samples). When the intra prediction mode of the current block corresponds to left mode, this may mean that the reference samples of the current block are located to the left of the current block. In this case, the current block can be determined to be partitioned in the vertical direction, and cu_sbt_horizontal_flag may not be signaled and may be derived as 0. When the intra prediction mode of the current block corresponds to top mode, this may mean that the reference samples of the current block are located at the top of the current block. In this case, the current block can be determined to be partitioned in the horizontal direction, and cu_sbt_horizontal_flag may be derived as 1 without being signaled. When the intra prediction mode of the current block does not correspond to left mode or top mode, the current block can be determined to be partitioned in a predefined direction (e.g., horizontally or vertically), and cu_sbt_horizontal_flag may be derived as 1 or 0 without being signaled. When the intra prediction mode of the current block does not correspond to the left mode and the top mode, cu_sbt_horizontal_flag for the current block may also be signaled.
[0238] It may be determined that the residual information of the first subblock close to the reference sample among the subblocks of the current block is not signaled, and the residual information of the second subblock far from the reference sample is signaled. That is, cu_sbt_pos_flag may be derived as 1 and not signaled.
[0239] Alternatively, as shown in Table 6, whether to signal the information on the sub-block transform and / or how to derive the information on the sub-block transform may be determined based on a weight index.
[0240] [Table 6]
[0241]
[0242] In Table 6, the left mode, the upper mode (left mode! and upper mode!) are the same as those in Table 1, and repeated explanations will be omitted here.
[0243] Based on the weight index (weight_idx), it can be determined whether to signal information about the sub-block transform of the current block. For example, when the weight index is equal to 1, information about the sub-block transform of the current block can be signaled. When the weight index is equal to 0 or 2, at least one of the information about the sub-block transform of the current block can be derived as a specific value. Alternatively, when the first weight (w0) applied to intra-frame predicted samples is greater than the second weight (w1) applied to inter-frame predicted samples, information about the sub-block transform of the current block can be signaled. When the first weight (w0) applied to intra-frame predicted samples is less than or equal to the second weight (w1) applied to inter-frame predicted samples, at least one of the information about the sub-block transform of the current block can be derived.
[0244] Meanwhile, based on at least one of a weight index or an intra prediction mode, at least one of cu_sbt_flag, cu_sbt_quad_flag, cu_sbt_horizontal_flag, or cu_sbt_pos_flag may be derived as a specific value, as seen in Reference Table 5.
[0245] The above-described method for signaling / deriving information about subblock transformation can be applied when the intra prediction mode of the current block is derived before signaling information about subblock transformation. Alternatively, even when the intra prediction mode of the current block is not derived before signaling information about subblock transformation, the above-described method can be applied when the range to which the intra prediction mode of the current block belongs can be derived using the above-described simple prediction method.
[0246] When the GPM mode is applied to the current block, sub-block transformation may be enabled for the current block. Based on the information about the GPM for the current block, at least one of the information about the sub-block transformation may be derived as a specific value without being signaled.
[0247] At least one of the information about the sub-block transform may not be signaled and may be derived as a specific value based on the partition index (merge_gpm_partition_idx) that specifies the geometric partition type. Depending on merge_gpm_partition_idx, cu_sbt_flag may be derived as a specific value and not signaled. Depending on merge_gpm_partition_idx, cu_sbt_quad_flag may be derived as a specific value and not signaled. Depending on merge_gpm_partition_idx, cu_sbt_horizontal_flag may be derived as a specific value and not signaled. Depending on merge_gpm_partition_idx, cu_sbt_pos_flag may be derived as a specific value and not signaled.
[0248] Table 7 below lists angleIdx and distanceIdx mapped to merge_gpm_partition_idx. angleIdx represents the angle of the boundary line according to the geometric partition, and the angle of the boundary line corresponding to each angleIdx is as follows: Figure 6 As shown in . distanceIdx can represent the distance from the center position of the current block to the boundary line.
[0249] [Table 7]
[0250]
[0251] Referring to Table 7, depending on angleIdx, cu_sbt_flag may not be signaled and may be derived as a specific value. Specifically, when angleIdx indicates an angle corresponding to the vertical direction (e.g., angleIdx = 0, 16), it can be determined that the sub-block transform is applied to the current block, and cu_sbt_flag may be derived as 1. When angleIdx indicates an angle corresponding to the horizontal direction (e.g., angleIdx = 8, 24), it can be determined that the sub-block transform is applied to the current block, and cu_sbt_flag may be derived as 1. When angleIdx indicates an angle adjacent to the vertical direction (e.g., angleIdx = 2, 14, 18, 30), it can be determined that the sub-block transform is applied to the current block, and cu_sbt_flag may be derived as 1. When angleIdx indicates any other angle, it can be determined that the sub-block transform is not applied to the current block, and cu_sbt_flag may be derived as 0.
[0252] The angles adjacent to the vertical direction described above are only examples. Even when angleIdx corresponds to at least one of 3, 5, 11, 13, 19, 21, 27, or 29, it can be determined that the sub-block transform is applied to the current block, and cu_sbt_flag can be derived as 1. When angleIdx represents an angle corresponding to the vertical and horizontal directions (i.e., angleIdx = 0, 16, 8, 24), cu_sbt_flag can be derived as 1; otherwise, cu_sbt_flag can be derived as 0.
[0253] The predefined angleIdx (or angles according to angleIdx) in the encoding and decoding devices can be divided into multiple groups. For example, the predefined angleIdx can be divided into at least two of the following: a first group consisting of angleIdx with vertical directivity; a second group consisting of angleIdx with horizontal directivity; or a third group consisting of the remaining angleIdx. Here, angleIdx with vertical directivity can refer to vertical angleIdx (i.e., 0, 16). Alternatively, angleIdx with vertical directivity can refer to a vertical angleIdx and one or more adjacent angleIdx (i.e., at least one of 2, 3, 4, 12, 13, 14, 18, 19, 20, 28, 29, or 30). An angleIdx with horizontal directivity can refer to a horizontal angleIdx (i.e., 8, 24). Alternatively, an angleldx having horizontal directivity may mean a horizontal angleldx and one or more angleldxs adjacent thereto (ie, at least one of 4, 5, 11, 12, 20, 21, 27, or 28).
[0254] Based on the directionality of angleIdx according to the partition index of the current block, information about subblock transformation can be derived as a specific value. As an example, Table 8 shows how to derive information about subblock transformation based on the group to which angleIdx according to the partition index of the current block belongs.
[0255] [Table 8]
[0256]
[0257] In Table 8, VER may correspond to the first group consisting of angleIdx with vertical directivity. HOR may correspond to the second group consisting of angleIdx with horizontal directivity. (!VER && !HOR) may correspond to the third group consisting of the remaining angleIdx.
[0258] When angleIdx for the current block belongs to the first group or the second group (ie, angleIdx corresponds to an angle having a vertical or horizontal orientation), it may be determined that the sub-block transform is applied to the current block, and cu_sbt_flag may be derived as 1. Otherwise, it may be determined that the sub-block transform is not applied to the current block, and cu_sbt_flag may be derived as 0.
[0259] When it is determined that the subblock transform is applied to the current block, the current block may be determined to be split into two subblocks having the same size, and cu_sbt_quad_flag may be derived as 0. However, this is merely an example, and the size of the subblock may be determined based on the signaled cu_sbt_quad_flag.
[0260] When it is determined that the sub-block transform is applied to the current block, the partition direction of the current block can be determined based on angleIdx of the current block (or the group to which angleIdx belongs). For example, when angleIdx of the current block belongs to the first group, the current block can be determined to be split in the vertical direction, and cu_sbt_horizontal_flag can be derived as 0. When angleIdx of the current block belongs to the second group, the current block can be determined to be split in the horizontal direction, and cu_sbt_horizontal_flag can be derived as 1. However, this is only an example, and the partition direction can also be determined based on cu_sbt_horizontal_flag sent by signal.
[0261] It may be determined that the residual information of the first subblock among the two subblocks in the current block is not signaled, and the residual information of the second subblock is signaled. The cu_sbt_pos_flag of the current block may not be signaled and may be derived as 1. However, this is merely an example, and the position of the subblock whose residual information is signaled may be determined based on the signaled cu_sbt_pos_flag.
[0262] In GPM mode, at least one of the sub-block transform information may not be signaled and may be derived as a specific value based on at least one of angleIdx or distanceIdx. As an example, Table 9 shows how to derive the sub-block transform information based on the group to which angleIdx belongs and distanceIdx according to the partition index of the current block.
[0263] [Table 9]
[0264]
[0265] When angleIdx for the current block belongs to the first group or the second group (ie, angleIdx corresponds to an angle having a vertical or horizontal orientation), it may be determined that the sub-block transform is applied to the current block, and cu_sbt_flag may be derived as 1. Otherwise, it may be determined that the sub-block transform is not applied to the current block, and cu_sbt_flag may be derived as 0.
[0266] When determining that a sub-block transform is applied to the current block, the size of the sub-block may be determined based on distanceIdx according to the partition index of the current block. For example, when distanceIdx is 0 or 1, the current block may be determined to be split into two sub-blocks of equal size, and cu_sbt_quad_flag may be derived as 0. When distanceIdx is 2 or 3, the current block may be determined to be split into two sub-blocks of 1 / 4 and 3 / 4 the size of the current block, and cu_sbt_quad_flag may be derived as 1. When distanceIdx is 0 or 1, this may mean that the mixed region at the partition boundary is approximately half of the current block. In this case, the current block may be determined to be split into sub-blocks of equal size. On the other hand, when distanceIdx is 2 or 3, this may mean that the mixed region at the partition boundary is approximately 1 / 4 of the current block. In this case, the current block may be determined to be split into two sub-blocks of 1 / 4 and 3 / 4 the size of the current block.
[0267] The partition direction for sub-block transformation can be determined in a direction similar to the GPM mode. For example, when angleIdx of the current block belongs to the first group, the current block can be determined to be split in the vertical direction, and cu_sbt_horizontal_flag can be derived as 0. When angleIdx of the current block belongs to the second group, the current block can be determined to be split in the horizontal direction, and cu_sbt_horizontal_flag can be derived as 1. However, this is only an example, and the partition direction can be determined based on the signaled cu_sbt_horizontal_flag.
[0268] It may be determined that the residual information of the first subblock among the two subblocks in the current block is not signaled, and the residual information of the second subblock is signaled. The cu_sbt_pos_flag of the current block may not be signaled and may be derived as 1. However, this is merely an example, and the position of the subblock whose residual information is signaled may be determined based on the signaled cu_sbt_pos_flag.
[0269] Alternatively, signaling of residual information for subblocks having statistically less residual information may be skipped. When a GPM-INTRA-based merge mode is applied to a current block, a subblock including relatively more partitions for deriving motion information based on the merge mode may be determined, and signaling of residual information for the corresponding subblock may be skipped.
[0270] As described above, when the GPM mode is applied to the current block, at least one of the information on the sub-block transform may be derived based on at least one of angleIdx and distanceIdx according to the partition index. Furthermore, when the GPM-INTRA-based merge mode is applied to the current block, at least one of the information on the sub-block transform may not be signaled and may be derived as a specific value based on the intra prediction mode for the current block as in the above-described method, and a detailed description thereof is omitted herein.
[0271] The above-mentioned step S410 may be performed when the current block is coded in the merge mode, and may be omitted when the current block is coded in the skip mode.
[0272] refer to Figure 4 , the current block may be reconstructed based on the prediction samples and the residual samples of the current block ( S420 ).
[0273] Specifically, when the current block is coded in merge mode, reconstructed samples of the current block may be generated based on the prediction samples and residual samples of the current block. When the current block is coded in skip mode, since the residual samples for the current block are not encoded and signaled, the prediction samples of the current block may be set as the reconstructed samples of the current block.
[0274] Figure 7 The diagram illustrates a schematic configuration of a decoding device (300) for performing an image decoding method according to the present disclosure.
[0275] refer to Figure 7 , the decoding apparatus ( 300 ) may include a prediction sample generator ( 700 ), a residual sample deriver ( 710 ), and a reconstructor ( 720 ).
[0276] The prediction sample generator (700) can be configured in Figure 3 The prediction sample generator (700) may generate a prediction sample of the current block based on a predetermined inter-frame prediction mode.
[0277] Specifically, when the current block is encoded in the merge mode, the prediction sample generator (700) may set any one of a plurality of inter-frame prediction modes predefined in the image decoding apparatus as the inter-frame prediction mode of the current block. In this case, the plurality of inter-frame prediction modes may include at least one of a sub-block merge mode, an MMVD (merge mode with motion vector difference), a normal merge mode, a CIIP (combined inter-frame and intra-frame prediction) mode, or a GPM mode. This has been referred to Figure 4 Described.
[0278] Alternatively, when the current block is coded in skip mode, the prediction sample generator (700) may generate prediction samples for the current block based on any one of the aforementioned inter prediction modes. The generated prediction samples may be set as reconstructed samples. However, when the current block is coded in skip mode, application of the CIIP mode may be restricted. Alternatively, similar to the aforementioned merge mode, application of the CIIP mode to the current block may be allowed even when the current block is coded in skip mode. When the current block is coded in skip mode, the GPM-INTRA-based merge mode may not be allowed.
[0279] refer to Figure 4 A method for generating prediction samples according to a predetermined inter prediction mode is described, and a detailed description thereof will be omitted here.
[0280] The residual sample deriver (710) can be configured in Figure 3 The residual sample deriver (710) may derive the residual sample of the current block based on the information about the sub-block transform. Here, the information about the sub-block transform may include at least one of an SBT flag (cu_sbt_flag), an SBT size flag (cu_sbt_quad_flag), an SBT direction flag (cu_sbt_horizontal_flag), or an SBT position flag (cu_sbt_pos_flag), as described in reference to Figure 4 As stated.
[0281] In addition, as referenced Figure 4When the CIIP mode is applied to the current block, all or part of the information regarding sub-block transformations may be derived based on the characteristics of the CIIP mode. When the CIIP mode is applied to the current block, at least one of the sub-block transformation information may not be signaled and may be derived as a specific value based on at least one of the weights (or weight indices) in the CIIP mode, the intra-prediction mode, the range of the region to which the weights are applied, or the position of the sub-region to which the weights are applied. When the GPM mode is applied to the current block, sub-block transformations may be permitted for the current block, and at least one of the sub-block transformation information may not be signaled and may be derived as a specific value based on information regarding the GPM used for the current block.
[0282] The reconstructor (720) may be configured in Figure 3 The reconstructor (720) may reconstruct the current block based on the prediction sample and the residual sample of the current block.
[0283] Figure 8 The diagram illustrates an image encoding method performed in an encoding device according to the present disclosure.
[0284] refer to Figure 8 , a prediction sample of the current block may be generated based on a predetermined inter prediction mode ( S800 ).
[0285] When encoding the current block in the merge mode, any one of the multiple inter-frame prediction modes predefined in the image encoding device can be set as the inter-frame prediction mode of the current block. In this case, the multiple inter-frame prediction modes may include at least one of the sub-block merge mode, MMVD mode, normal merge mode, CIIP mode, or GPM mode. This is as described in reference Figure 4 As stated.
[0286] When encoding the current block in merge mode, it is possible to sequentially check which of the aforementioned inter prediction modes corresponds to the inter prediction mode of the current block according to the priority order between the inter prediction modes. For example, first, it is possible to determine whether the inter prediction mode of the current block is the subblock merge mode. When the inter prediction mode of the current block is determined to be the subblock merge mode, merge_subblock_flag may be encoded as 1, and when it is determined that the inter prediction mode of the current block is not determined to be the subblock merge mode, merge_subblock_flag may be encoded as 0.
[0287] When the inter prediction mode of the current block is not determined as the sub-block merge mode, it can be determined whether the inter prediction mode of the current block is the regular merge mode. When the inter prediction mode of the current block is determined as the regular merge mode, regular_merge_flag can be encoded as 1, and it can be additionally determined whether the inter prediction mode of the current block is the MMVD mode. When the inter prediction mode of the current block is determined as the MMVD mode, mmvd_merge_flag can be encoded as 1. When the inter prediction mode of the current block is not determined as the MMVD mode, mmvd_merge_flag can be encoded as 0, and the inter prediction mode of the current block can be determined as the regular merge mode.
[0288] When the inter-frame prediction mode of the current block is not determined as the normal merge mode, it can be determined whether the inter-frame prediction mode of the current block is the CIIP mode. When the inter-frame prediction mode of the current block is determined as the CIIP mode, ciip_flag can be encoded as 1. When the inter-frame prediction mode of the current block is not determined as the CIIP mode, ciip_flag can be encoded as 0, and it can be additionally determined whether the inter-frame prediction mode of the current block is the GPM-INTRA-based merge mode. When the inter-frame prediction mode of the current block is determined as the GPM-INTRA-based merge mode, gpm_intra_flag can be encoded as 1. When the inter-frame prediction mode of the current block is not determined as the GPM-INTRA-based merge mode, gpm_intra_flag can be encoded as 0, and the inter-frame prediction mode of the current block can be determined as the GPM-based merge mode.
[0289] When the current block is encoded in the skip mode, the prediction samples of the current block can be generated based on any one of the aforementioned inter-frame prediction modes, and the generated prediction samples can be set as reconstructed samples. However, when the current block is encoded in the skip mode, the application of the CIIP mode can be restricted. Alternatively, similar to the aforementioned merge mode, the CIIP mode can be allowed to be applied to the current block even when the current block is encoded in the skip mode. This is as described in reference to Figure 4 As stated.
[0290] The intra prediction sample according to the CIIP mode can be derived based on any one of the predefined intra prediction modes. A candidate list can be constructed to derive the intra prediction mode of the current block, and at least one of the mode index (ciip_mpm_idx) or the group index (ciip_group_idx) for specifying the intra prediction mode of the current block can be encoded. This is as described in reference Figure 4 As stated.
[0291] When the CIIP mode is applied to the current block, the weight of the weighted sum between the inter-frame prediction sample and the intra-frame prediction sample for the current block can be determined. The weight can be determined based on at least one of the intra-frame prediction mode of the current block or the position of the sub-region to which the prediction sample belongs. Alternatively, the image encoding device can determine the best weight among predefined weight candidates and encode a weight index indicating the best weight. The method for determining the weight and the method for applying the same are as described in reference to Figure 4 As stated.
[0292] The information about the GPM according to the present disclosure may include at least one of information for a GPM-INTRA-based merging mode or information for a GPM-based merging mode. Figure 4 The information about the GPM may include at least one of merge_gpm_partition_idx, gpm_intra_flag0, gpm_intra_mode_idx0, merge_gpm_idx0, gpm_intra_flag1, gpm_intra_mode_idx1, or merge_gpm_idx1.
[0293] It may be determined whether the first partition in the current block is predicted based on the intra mode, and gpm_intra_flag0 may be encoded based on the determination. When gpm_intra_flag0 is encoded as 1, the intra prediction mode for the first partition may be determined, and gpm_intra_mode_idx0 indicating the intra prediction mode may be encoded. On the other hand, when gpm_intra_flag0 is encoded as 0, a merge candidate for the first partition may be determined from the merge candidate list of the current block, and merge_gpm_idx0 indicating the merge candidate may be encoded.
[0294] Similarly, whether the second partition in the current block is predicted based on the intra prediction mode may be determined, and gpm_intra_flag1 may be encoded based on the determination. When gpm_intra_flag1 is encoded as 1, the intra prediction mode for the second partition may be determined, and gpm_intra_mode_idx1 indicating the intra prediction mode may be encoded. On the other hand, when gpm_intra_flag1 is encoded as 0, a merge candidate for the second partition may be determined from the merge candidate list of the current block, and merge_gpm_idx1 indicating the merge candidate may be encoded.
[0295] When the current block is not encoded in the skip mode, the CIIP mode may be allowed. Similarly, when the current block is not encoded in the skip mode, the GPM-INTRA-based merge mode may be allowed to be restricted. That is, when the current block is encoded in the skip mode, the GPM-INTRA-based merge mode may not be allowed for the current block. To this end, at least one of the above-mentioned GPM information may be adaptively signaled based on a flag (cu_skip_flag) indicating whether the current block is a block encoded in the skip mode, as described in reference to Figure 4 As stated.
[0296] refer to Figure 8 , residual samples of the current block may be derived based on the prediction samples of the current block ( S810 ).
[0297] The residual samples of the current block may be derived by subtracting the predicted samples of the current block from the original samples of the current block.
[0298] refer to Figure 8 , information on sub-block transformation for encoding the current block residual sample may be determined ( S820 ).
[0299] The information about the sub-block transformation according to the present disclosure may include at least one of an SBT flag (cu_sbt_flag), an SBT size flag (cu_sbt_quad_flag), an SBT direction flag (cu_sbt_horizontal_flag), or an SBT position flag (cu_sbt_pos_flag). Figure 4 As stated.
[0300] When the CIIP mode is applied to the current block, all or part of the information on the sub-block transformation can be derived by using the characteristics of the CIIP mode, thereby reducing the number of bits used to signal the information on the sub-block transformation. Figure 4 As stated.
[0301] When the CIIP mode is applied to the current block, at least one of the information on the sub-block transformation may not be signaled and may be derived as a specific value based on at least one of the weight (or weight index) in the CIIP mode, the intra prediction mode, the range of the area where the weight is applied, or the position of the sub-area where the weight is applied. This is as described in reference Figure 4 As stated.
[0302] When the GPM mode is applied to the current block, sub-block transformation may be allowed for the current block. At least one of the information about the sub-block transformation may not be signaled and may be derived as a specific value based on the information about the GPM for the current block. Figure 4 As stated.
[0303] refer to Figure 8 , the residual samples of the current block may be encoded to generate residual information (S830). Here, the encoding of the residual samples may include at least one of the following: 1) transforming the residual samples; 2) quantizing the transform coefficients; or 3) determining residual information about the quantized transform coefficients.
[0304] refer to Figure 8 , the residual information of the current block may be encoded to generate a bitstream (S840). The bitstream may further include information specifying an inter-frame prediction mode of the current block. The bitstream may further include all or part of information about sub-block transformation of the current block.
[0305] Figure 9 The diagram illustrates a schematic configuration of an encoding device (200) that performs an image encoding method according to the present disclosure.
[0306] refer to Figure 9 , the encoding device (200) may include a prediction sample generator (900), a residual sample deriver (910), an SBT information determiner (920), a residual information generator (930) and a residual information encoder (940).
[0307] The prediction sample generator (900) can be configured in Figure 2 The prediction sample generator (900) may generate the prediction sample of the current block based on a predetermined inter-frame prediction mode. The method for determining / signaling the inter-frame prediction mode of the current block and the method for generating the prediction sample according to the method are as described in reference Figure 4 and Figure 8 described.
[0308] exist Figure 2 The residual processor (230) may be configured with a residual sample deriver (910), an SBT information determiner (920), and a residual information generator (930). The residual sample deriver (910) may derive the residual sample of the current block based on the prediction sample of the current block. The SBT information determiner (920) may determine information about sub-block transformation for encoding the residual sample of the current block. Figure 4 and Figure 8 Methods for signaling / deriving information about sub-block transforms have been described. The residual information generator (930) may generate residual information by encoding residual samples of the current block.
[0309] exist Figure 2 The residual information encoder (940) may be configured in the entropy encoder (240). The residual information encoder (940) may encode the residual information of the current block to generate a bitstream.
[0310] In the above embodiments, the methods are described as a series of steps or blocks based on the flowcharts, but the corresponding embodiments are not limited to the order of the steps, and some steps may occur simultaneously or in a different order than other steps described above. In addition, those skilled in the art will understand that the steps shown in the flowcharts are not exclusive, and other steps may be included or one or more steps in the flowcharts may be deleted without affecting the scope of the embodiments of the present disclosure.
[0311] The above-mentioned method according to an embodiment of the present disclosure can be implemented in the form of software, and the encoding device and / or decoding device according to the present disclosure can be included in a device that performs image processing, such as a TV, a computer, a smart phone, a set-top box, a display device, etc.
[0312] In the present disclosure, when embodiments are implemented as software, the above methods can be implemented as modules (processes, functions, etc.) that perform the above functions. The modules can be stored in memory and executed by a processor. The memory can be located internally or externally to the processor and can be connected to the processor via various well-known means. The processor may include an application-specific integrated circuit (ASIC), another chipset, logic circuits, and / or a data processing device. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. In other words, the embodiments described herein can be implemented on a processor, microprocessor, controller, or chip. For example, the functional units shown in each figure can be implemented on a computer, processor, microprocessor, controller, or chip. In this case, information (e.g., information regarding instructions) or algorithms used for implementation can be stored on a digital storage medium.
[0313] In addition, the decoding device and encoding device using the embodiments of the present disclosure can be included in multimedia broadcast transmission and reception devices, mobile communication terminals, home theater video equipment, digital theater video equipment, surveillance cameras, video conversation equipment, real-time communication equipment such as video communication, mobile streaming equipment, storage media, cameras, equipment for providing video on demand (VoD) services, over-the-top (OTT) equipment, equipment for providing Internet streaming services, three-dimensional (3D) video equipment, virtual reality (VR) equipment, augmented reality (AR) equipment, videophone video equipment, transportation terminals (e.g., vehicle (including autonomous vehicle) terminals, aircraft terminals, ship terminals, etc.), and medical video equipment, and can be used to process video signals or data signals. For example, over-the-top (OTT) devices may include game consoles, Blu-ray players, networked TVs, home theater systems, smartphones, tablet computers, digital video recorders (DVRs), and the like.
[0314] In addition, the processing method of the embodiment of the present disclosure can be generated in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiment of the present disclosure can also be stored in a computer-readable recording medium. Computer-readable recording media include all types of storage devices and distributed storage devices that store computer-readable data. Computer-readable recording media may include, for example, Blu-ray discs (BDs), universal serial buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical media storage devices. In addition, computer-readable recording media include media implemented in the form of carrier waves (for example, transmitted via the Internet). In addition, the bit stream generated by the encoding method can be stored in a computer-readable recording medium or can be sent via a wired or wireless communication network.
[0315] In addition, the embodiments of the present disclosure may be implemented by a computer program product through program code, and the program code may be executed on a computer by the embodiments of the present disclosure. The program code may be stored on a computer-readable carrier.
[0316] Figure 10 An example of a content streaming system to which embodiments of the present disclosure can be applied is shown.
[0317] refer to Figure 10 The content streaming transmission system to which the embodiments of the present disclosure are applied may mainly include an encoding server, a streaming transmission server, a web server, a media storage, a user device, and a multimedia input device.
[0318] The encoding server compresses content input from a multimedia input device such as a smartphone, a camera, or a camcorder into digital data to generate a bitstream and transmits it to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, or a camcorder directly generates a bitstream, the encoding server may be omitted.
[0319] A bitstream may be generated by applying the encoding method or the bitstream generating method of the embodiment of the present disclosure, and the streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0320] The streaming server transmits multimedia data to the user device via a web server based on the user's request, and the web server serves as a medium for notifying the user of available services. When the user requests the desired service from the web server, the web server delivers it to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, and in this case, the control server controls the commands and responses between each device in the content streaming system.
[0321] The streaming server can receive content from media storage and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0322] Examples of user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation, tablet PCs, tablet computers, ultrabooks, wearable devices (e.g., smart watches, smart glasses, head-mounted displays (HMDs), digital TVs, desktop computers, digital signage, etc.).
[0323] Each server in the content streaming system may be operated as a distributed server, and in this case, data received from each server may be distributed and processed.
[0324] The claims set forth herein can be combined in various ways. For example, the technical features of the method claims of the present disclosure can be combined and implemented as a device, and the technical features of the device claims of the present disclosure can be combined and implemented as a method. In addition, the technical features of the method claims of the present disclosure and the technical features of the device claims of the present disclosure can be combined and implemented as a device, and the technical features of the method claims of the present disclosure and the technical features of the device claims of the present disclosure can be combined and implemented as a method.
Claims
1. An image decoding method, comprising: Generate a prediction sample of the current block based on a predetermined inter-frame prediction mode; deriving residual samples of the current block based on information about a sub-block transform (SBT); as well as reconstructing the current block based on the prediction samples and the residual samples of the current block, wherein at least one of the information about the SBT is derived based on the inter prediction mode, and The information about the SBT includes at least one of an SBT flag, an SBT size flag, an SBT direction flag, or an SBT position flag.
2. The method according to claim 1, wherein When the inter prediction mode of the current block is the CIIP mode, the SBT is allowed for the current block.
3. The method according to claim 2, wherein: When the inter prediction mode of the current block is the CIIP mode, at least one of the information about the SBT is not signaled and is derived based on the intra prediction mode for the CIIP mode.
4. The method according to claim 3, wherein: The intra prediction mode of the current block is derived as any one of one or more candidate modes belonging to a candidate list.
5. The method according to claim 2, wherein: When the inter prediction mode of the current block is the CIIP mode, at least one of the information about the SBT is not signaled and is derived based on a weight for the CIIP mode.
6. The method according to claim 5, wherein: The weight is determined based on at least one of an intra prediction mode of the current block, a position of a subregion to which a prediction sample of the current block belongs, or a weight index.
7. The method according to claim 5, wherein: The CIIP mode is allowed regardless of a flag indicating whether the current block is a block coded in skip mode.
8. The method according to claim 1, wherein When the inter prediction mode of the current block is a geometric partition merging mode, the sub-block transform is allowed for the current block.
9. The method according to claim 8, wherein When the inter prediction mode of the current block is the geometric partition merge mode, at least one of the information about the SBT is not sent using a signal and is derived based on at least one of an angle of a boundary line according to a geometric partition of the current block or a distance from a center position of the current block to the boundary line.
10. An image encoding method, comprising: Generate a prediction sample of the current block based on a predetermined inter-frame prediction mode; deriving residual samples of the current block based on the prediction samples of the current block; determining information about a sub-block transform (SBT) used to encode the residual samples of the current block; Encoding the residual samples of the current block to generate residual information; as well as encoding the residual information of the current block to generate a bitstream, wherein at least one of the information about the SBT is derived based on the inter prediction mode, and The information about the SBT includes at least one of an SBT flag, an SBT size flag, an SBT direction flag, or an SBT position flag. 11 . A computer-readable storage medium storing a bit stream generated by the image encoding method according to claim 10 .
12. A method for transmitting data, comprising: Obtaining a bitstream for image information, wherein the bitstream is generated by: generating prediction samples of a current block based on a predetermined inter-frame prediction mode, deriving residual samples of the current block based on the prediction samples of the current block, determining information about a sub-block transform (SBT) for encoding the residual samples of the current block, encoding the residual samples of the current block to generate residual information, and encoding the residual information of the current block; and sending said data comprising said bitstream, wherein at least one of the information about the SBT is derived based on the inter prediction mode, and The information about the SBT includes at least one of an SBT flag, an SBT size flag, an SBT direction flag, or an SBT position flag.