Image Compilation Method and Its Device
By adopting an intra prediction-based method in image encoding, using intra MIP syntax elements and variable MIP flags, the problem of low compression efficiency of high resolution and high-quality image/video data is solved, and more efficient image/video encoding is achieved.
Patent Information
- Application Number
- CN202180033890.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-05-11
- Filing Date
- 2021-05-11
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-05-11
AI Technical Summary
The prior art is difficult to effectively compress and transmit or store high-resolution and high-quality image/video data, especially when meeting the needs of applications such as ultra-high definition, virtual reality and holograms.
Using an intra prediction-based image encoding method, the intra prediction mode is derived and the reconstruction block is generated through intra MIP syntax elements and variable MIP flags, and the image encoding efficiency is improved.
The image/video compression efficiency is improved, especially in intra prediction, matrix-based prediction and ISP prediction scenarios, high resolution and high-quality image/video data can be processed more efficiently.
Smart Images

Figure CN115516855B_ABST
Abstract
Description
Technical Field
[0001] This document relates to image coding technology, and more specifically, to an image coding method and apparatus based on intra prediction in an image coding system. Background Art
[0002] Nowadays, the demand for high-resolution and high-quality images / videos such as 4K, 8K, or higher ultra-high definition (UHD) images / videos has been increasing in various fields. As the image / video data becomes higher resolution and higher quality, the amount of information or bits transmitted increases compared to traditional image data. Therefore, when transmitting image data using a medium such as a traditional wired / wireless broadband line or storing image / video data using an existing storage medium, the transmission cost and storage cost increase.
[0003] In addition, nowadays, the interest and demand for immersive media such as virtual reality (VR) and artificial reality (AR) content or holograms are increasing, and the broadcasting of images / videos with image characteristics different from those of real images such as game images is increasing.
[0004] Therefore, there is a need for an efficient image / video compression technology that can effectively compress, transmit, store, and reproduce the information of high-resolution and high-quality images / videos with various characteristics as described above. Summary of the Invention
[0005] Technical Problem
[0006] A technical aspect of the present disclosure is to provide a method and device for improving image coding efficiency.
[0007] Another technical aspect of this document is to provide an efficient intra prediction method and its device.
[0008] Another technical aspect of this document is to provide an image coding method and apparatus for intra prediction based on a matrix.
[0009] Another technical aspect of this document is to provide an image coding method and apparatus for intra prediction based on IPS.
[0010] Technical Solution
[0011] According to an embodiment of this document, an image decoding method performed by a decoding device is provided. The method may include: receiving, from a bitstream, image information including intra prediction type information, the intra prediction type information including an intra MIP syntax element for a first target block; deriving a value of the intra MIP syntax element; based on the value of the intra MIP syntax element, setting a variable MIP flag for a preset specific region that is the same as the region of the first target block; deriving an intra prediction mode for a second target block; deriving prediction samples for the second target block based on the intra prediction mode for the second target block; and generating a reconstructed block based on the prediction samples, wherein the intra prediction mode for the second target block may be derived based on the variable MIP flag for the first target block.
[0012] The first target block may be the left neighboring block of the second target block. Deriving the intra prediction mode for the second target block may include: deriving candidate intra prediction modes based on the variable MIP flag; and deriving the intra prediction mode for the second target block based on the candidate intra prediction modes. The specific region may include the sample position of (xCb - 1, yCb + cbHeight - 1), (xCb, yCb) may be the top-left sample position of the second target block, and cbHeight may indicate the height of the second target block.
[0013] The first target block may be the upper neighboring block of the second target block. Deriving the intra prediction mode for the second target block may include: deriving candidate intra prediction modes based on the variable MIP flag; and deriving the intra prediction mode for the second target block based on the candidate intra prediction modes. The specific region may include the sample position of (xCb + cbWidth - 1, yCb - 1), (xCb, yCb) may be the position of the top-left sample of the second target block, and cbWidth may indicate the width of the second target block.
[0014] The second target block may include a chrominance block. The first target block may be the luminance block related to the chrominance block. Deriving the intra prediction mode for the second target block may include: deriving a corresponding luminance intra prediction mode based on the variable MIP flag; and deriving the intra prediction mode for the second target block based on the corresponding luminance intra prediction mode.
[0015] Here, based on that the tree type of the second target block is not a single tree or its chrominance array type is not 3, the specific region may include the sample position of (xCb + cbWidth / 2, yCb + cbHeight / 2), (xCb, yCb) may indicate the top-left position of the chrominance block in the luminance sample unit, cbWidth may indicate the width of the corresponding luminance block corresponding to the chrominance block, and cbHeight may indicate the height of the corresponding luminance block.
[0016] A variable MIP flag can be set for a specific region based on the fact that the tree type of the first target block is not a dual-tree chrominance.
[0017] According to another embodiment of this document, an image encoding method performed by an encoding device is provided. The method may include: deriving a value of an intra MIP flag for a first target block when applying an intra MIP mode to the first target block; setting a variable MIP flag for a preset specific region that is the same as the region of the first target block based on the value of the intra MIP flag; deriving an intra prediction mode for a second target block; deriving prediction samples for the second target block based on the intra prediction mode for the second target block; deriving residual samples for the second target block based on the prediction samples; and encoding and outputting transform coefficient information generated based on the intra MIP flag and the residual samples, where the intra prediction mode for the second target block can be derived based on the variable MIP flag for the first target block.
[0018] According to still another embodiment of this document, a digital storage medium can be provided, which stores image data including encoded image information and / or a bitstream generated according to an image encoding method performed by an encoding device.
[0019] According to yet another embodiment of this document, a digital storage medium can be provided, which stores image data including encoded image information and / or a bitstream to cause a decoding device to perform an image decoding method.
[0020] Beneficial effects
[0021] According to this document, the overall image / video compression efficiency can be increased.
[0022] According to this document, the efficiency in intra prediction can be improved.
[0023] According to this document, the image compilation efficiency of matrix-based intra prediction can be improved.
[0024] According to this document, the image compilation efficiency of ISP-based intra prediction can be improved.
[0025] The effects that can be obtained through specific examples of the present disclosure are not limited to the effects listed above. For example, there may be various technical effects that can be understood by those of ordinary skill in the relevant art or derived from the present disclosure. Therefore, the specific effects of the present disclosure are not limited to the effects explicitly described in the present disclosure, and can include various effects that can be understood or derived from the technical features of the present disclosure. Description of the drawings
[0026] Figure 1 Schematically illustrate an example of a video / image compilation system to which the present disclosure is applicable.
[0027] Figure 2 FIG. is a diagram schematically illustrating the configuration of a video / image encoding device to which the present disclosure is applicable.
[0028] Figure 3 FIG. is a diagram schematically illustrating the configuration of a video / image decoding device to which the present disclosure is applicable.
[0029] Figure 4 FIG. illustrates the structure of a content streaming system to which the present disclosure is applied.
[0030] Figure 5 FIG. illustrates an intra prediction mode of 65 prediction directions.
[0031] Figure 6 FIG. illustrates a process of generating prediction samples based on MIP according to an example.
[0032] Figure 7 FIG. illustrates an example of sub-blocks into which a compilation block is divided.
[0033] Figure 8 FIG. illustrates another example of sub-blocks into which a compilation block is divided.
[0034] Figure 9 FIG. schematically illustrates a multi-transformation scheme according to an embodiment of this document.
[0035] Figure 10 FIG. illustrates RST according to an embodiment of this document.
[0036] Figure 11 FIG. is a flowchart illustrating the operation of a video decoding device according to an embodiment of this document.
[0037] Figure 12a and Figure 12b FIG. shows that the variable MIP flag for a first target block is used to derive an intra prediction mode for a second target block according to an embodiment of this document.
[0038] Figure 13 FIG. illustrates the configuration of samples according to a chroma format.
[0039] Figures 14a to 14c FIG. shows that the variable MIP flag for a corresponding luma block as a first target block is used to derive an intra prediction mode for a chroma block as a second target block according to an embodiment of this document.
[0040] Figure 15 FIG. illustrates the operation of a video encoding device according to an embodiment of this document. DETAILED DESCRIPTION
[0041] Although the present disclosure may be susceptible to various modifications and include various embodiments, specific embodiments thereof have been shown by way of example in the drawings and will now be described in detail. However, this is not intended to limit the present disclosure to the specific embodiments disclosed herein. The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the technical concept of the present disclosure. Unless the context clearly indicates otherwise, the singular forms may include the plural forms. Terms such as "including" and "having" are intended to indicate the presence of the features, numbers, steps, operations, elements, components, or combinations thereof used in the following description, and thus should not be construed as precluding the possibility of the presence or addition of one or more different features, numbers, steps, operations, elements, components, or combinations thereof.
[0042] Meanwhile, for the convenience of describing different characteristic functions from each other, each component in the drawings described herein is illustrated independently. However, it is not meant that each component is implemented by separate hardware or software. For example, any two or more of these components can be combined to form a single component, and any single component can be divided into multiple components. Embodiments in which components are combined and / or divided will fall within the scope of the patent right of the present disclosure as long as they do not depart from the essence of the present disclosure.
[0043] Hereinafter, preferred embodiments of the present disclosure will be described in more detail with reference to the drawings. Additionally, in the drawings, the same reference numerals are used for the same components, and repeated descriptions of the same components will be omitted.
[0044] This document relates to video / image compilation. For example, the methods / examples disclosed in this document may relate to the VVC (Versatile Video Coding) standard (ITU-T Rec.H.266), the next-generation video / image compilation standard after VVC, or other video compilation-related standards (e.g., the HEVC (High Efficiency Video Coding) standard (ITU-T Rec.H.265), the EVC (Essential Video Coding) standard, the AVS2 standard, etc.).
[0045] In this document, various embodiments related to video / image compilation may be provided, and unless otherwise specified, these embodiments may be combined with each other and executed.
[0046] In this document, video may refer to a collection of a series of images over a period of time. Generally, a picture refers to a unit representing an image of a specific time region, and a slice / tile is a unit that constitutes a part of a picture. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / tiles. A picture may be composed of one or more tile groups. A tile group may include one or more tiles.
[0047] A pixel or pel may refer to the smallest unit that makes up a picture (or image). Additionally, the term "sample" can be used as a term corresponding to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component. Alternatively, a sample can mean a pixel value in the spatial domain, or when the pixel value is transformed into the frequency domain, it can mean a transform coefficient in the frequency domain.
[0048] A unit can represent the basic unit of image processing. A unit can include at least one of a specific region and information related to that region. A unit can include one luminance block and two chrominance (e.g., cb, cr) blocks. Depending on the situation, terms such as unit and terms like block, region, etc. can be used interchangeably. Usually, an M×N block can include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.
[0049] In this document, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" can mean "A and / or B". Additionally, "A, B" can mean "A and / or B". Additionally, "A / B / C" can mean "at least one of A, B, and / or C". Additionally, "A / B / C" can mean "at least one of A, B, and / or C".
[0050] Additionally, in this document, the term "or" should be interpreted as indicating "and / or". For example, the expression "A or B" can include 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted as indicating "additionally or alternatively".
[0051] In this disclosure, "at least one of A and B" can mean "only A", "only B", or "both A and B". Furthermore, in this disclosure, the expression "at least one of A or B" or "at least one of A and / or B" can be interpreted as "at least one of A and B".
[0052] Furthermore, in this disclosure, "at least one of A, B, and C" can mean "only A", "only B", "only C", or "any combination of A, B, and C". Additionally, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C".
[0053] In addition, the parentheses used in the present disclosure may indicate "for example". Specifically, when indicated as "prediction (intra prediction)", it may mean that "intra prediction" is presented as an example of "prediction". In other words, "prediction" in the present disclosure is not limited to "intra prediction", and "intra prediction" is presented as an example of "prediction". In addition, when indicated as "prediction (i.e., intra prediction)", this may also mean that "intra prediction" is presented as an example of "prediction".
[0054] The technical features separately described in one of the drawings in the present disclosure may be implemented separately or may be implemented simultaneously.
[0055] Figure 1 Schematically illustrate an example of a video / image compilation system to which the present disclosure is applicable.
[0056] Reference Figure 1 , the video / image encoding system may include a first device (source device) and a second device (receiving device). The source device may deliver encoded video / image information or data to the receiving device in the form of a file or a stream via a digital storage medium or a network.
[0057] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0058] The video source may obtain video / images through processes such as capturing, synthesizing, or generating video / images. The video source may include a video / image capturing device and / or a video / image generating device. The video / image capturing device may include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generating device may include, for example, a computer, a tablet computer, and a smart phone, and may (electronically) generate video / images. For example, virtual video / images may be generated by a computer or the like. In this case, the video / image capturing process may be replaced by a process of generating relevant data.
[0059] The encoding device may encode the input video / images. The encoding device may perform a series of processes such as prediction, transformation, and quantization for compression and compilation efficiency. The encoded data (encoded video / image information) may be output in the form of a bitstream.
[0060] The transmitter can send the encoded video / image information or data output in the form of a bitstream to the receiver of the receiving device in the form of a file or a stream via a digital storage medium or a network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include elements for generating a media file in a predetermined file format and can include elements for transmitting via a broadcast / communication network. The receiver can receive / extract the bitstream and send the received / extracted bitstream to the decoding device.
[0061] The decoding device can decode the video / image by performing a series of processes such as dequantization, inverse transformation, prediction, etc. corresponding to the operations of the encoding device.
[0062] The renderer can render the decoded video / image. The rendered video / image can be displayed via a display.
[0063] Figure 2 FIG. is a schematic diagram illustrating the configuration of a video / image encoding device to which the present disclosure can be applied. Hereinafter, the term referred to as a video encoding device may include an image encoding device.
[0064] Reference Figure 2 , the encoding device 200 can include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 can further include a subtractor 231. The adder 250 can be referred to as a reconstructor or a reconstruction block generator. The above-described image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 can be configured by one or more hardware components (e.g., an encoder chipset or a processor) according to an embodiment. In addition, the memory 270 can include a decoded picture buffer (DPB) and can be configured by a digital storage medium. The hardware components can also include the memory 270 as an internal / external component.
[0065] The image partitioner 210 may partition an input image (or picture or frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit may be referred to as a coding unit (CU). In this case, starting from a coding tree unit (CTU) or a largest coding unit (LCU), the coding unit may be recursively divided according to a quadtree binary tree ternary tree (QTBTTT) structure. For example, a coding unit may be divided into multiple coding units with a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, the quadtree structure may be applied first and the binary tree structure and / or the ternary tree structure may be applied later. Alternatively, the binary tree structure may be applied first. The coding process according to the present disclosure may be performed based on the final coding unit that is not further divided. In this case, the largest coding unit may be directly used as the final coding unit based on the coding efficiency according to the image characteristics. Alternatively, the coding unit may be recursively divided into further deeper coding units as needed, such that the coding unit with the optimal size may be used as the final coding unit. Here, the coding process may include processes such as prediction, transformation, and reconstruction, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be split or partitioned from the above-mentioned final coding unit. The prediction unit may be a unit for sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0066] Depending on the situation, the terms unit and terms such as block, region, etc. may be used in place of each other. In general, an M×N block may represent a set of samples or transform coefficients composed of M columns and N rows. Samples generally may represent pixels or pixel values, and may represent only the pixels / pixel values of the luminance component, or only the pixels / pixel values of the chrominance component. Samples may be used as a term corresponding to the pixels or pels of a picture (or image).
[0067] The subtractor 231 subtracts the prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 221 or the intra-frame predictor 222 from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to the transformer 232. In this case, as shown, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) in the encoder 200 may be referred to as the subtractor 231. The predictor may perform prediction on the processing target block (hereinafter referred to as the "current block") and may generate a prediction block including prediction samples for the current block. The predictor may determine whether to apply intra-frame prediction or inter-frame prediction on a per current block or CU basis. As discussed later, in the description of each prediction mode, the predictor may generate various information related to prediction, such as prediction mode information, and send the generated information to the entropy encoder 240. The information about the prediction may be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0068] The intra-frame predictor 222 may predict the current block by referring to samples in the current picture. Depending on the prediction mode, the reference samples may be located in the neighboring region of the current block or in a distant region far from the current block. In intra-frame prediction, the prediction mode may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, the DC mode and the planar mode. Depending on the level of detail of the prediction direction, the directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the configuration. The intra-frame predictor 222 may determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring block.
[0069] The inter-frame predictor 221 may derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter-frame prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, the adjacent blocks may include spatial adjacent blocks present in the current picture and temporal adjacent blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal adjacent block may be the same as or different from each other. The temporal adjacent block may be referred to as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal adjacent block may be referred to as a collocated picture (colPic). For example, the inter-frame predictor 221 may configure a motion information candidate list based on adjacent blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. The inter-frame prediction may be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter-frame predictor 221 may use the motion information of adjacent blocks as the motion information of the current block. In the skip mode, different from the merge mode, the residual signal may not be transmitted. In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vector of an adjacent block may be used as a motion vector predictor, and the motion vector of the current block may be indicated by signaling a motion vector difference.
[0070] The predictor 220 may generate a prediction signal based on various prediction methods. For example, the predictor may apply intra-frame prediction or inter-frame prediction for predicting a block and may also apply intra-frame prediction and inter-frame prediction simultaneously. This may be referred to as combined inter-frame and intra-frame prediction (CIIP). In addition, the predictor may perform prediction on a block based on the intra-block copy (IBC) prediction mode or the palette mode. The IBC prediction mode or the palette mode may be used for content image / video compilation in games, etc., such as screen content compilation (SCC). Although IBC basically performs prediction in the current block, the aspect of deriving a reference block within the current block may be performed similarly to inter-frame prediction. That is, IBC can use at least one of the inter-frame prediction techniques described in the present disclosure.
[0071] The prediction signal generated by the inter-frame predictor 221 and / or the intra-frame predictor 222 can be used to generate a reconstructed signal or to generate a residual signal. The transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional non-linear transform (CNT). Here, GBT means a transform obtained from a graph when the relationship information between pixels is represented by the graph. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. Additionally, the transform process can be applied to square pixel blocks of the same size or can be applied to blocks of variable size other than square blocks.
[0072] Quantizer 233 may quantize the transform coefficients and send the quantized transform coefficients to entropy encoder 240, and entropy encoder 240 may encode the quantized signal (information regarding the quantized transform coefficients) and may output the encoded signal in a bitstream. The information regarding the quantized transform coefficients may be referred to as residual information. Quantizer 233 may rearrange the block type quantized transform coefficients into a one-dimensional vector form based on the coefficient scan order, and generate information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. Entropy encoder 240 may perform various encoding methods, such as, for example, exponential Golomb, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), etc. Entropy encoder 240 may encode, together or separately, information necessary for video / image reconstruction other than the quantized transform coefficients (e.g., values of syntax elements, etc.). The encoded information (e.g., encoded video / image information) may be sent or stored in the form of a bitstream on a unit basis of a network abstraction layer (NAL). The video / image information may also include information regarding various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), a video parameter set (VPS), etc. In addition, the video / image information may also include general constraint information. In the present disclosure, information sent / signaled from an encoding device to a decoding device and / or syntax elements may be included in the video / image information. The video / image information may be encoded through the above encoding process and included in the bitstream. The bitstream may be sent through a network or stored in a digital storage medium. Here, the network may include a broadcast network, a communication network, and / or the like, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that sends the signal output from entropy encoder 240 and / or a storage device (not shown) that stores the signal may be configured as an internal / external element of encoding device 200, or the transmitter may be included in entropy encoder 240.
[0073] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, by applying dequantization and inverse transformation to the quantized transform coefficients via the dequantizer 234 and the inverse transformer 235, a residual signal (residual block or residual samples) can be reconstructed. The adder 255 adds the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222, enabling the generation of a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). When there is no residual for the target block being processed, as in the case of applying the skip mode, the predicted block can be used as the reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next target block in the current block, and as will be described later, the generated reconstructed signal can be used for inter-frame prediction of the next picture through filtering.
[0074] Meanwhile, during picture encoding and / or reconstruction, luminance mapping and chrominance scaling (LMCS) can be applied.
[0075] The filter 260 can improve the subjective / objective video quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and the filter 260 can store the modified reconstructed picture in the memory 270, more specifically, in the DPB of the memory 270. Various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. As will be discussed later in the description of each filtering method, the filter 260 can generate various information related to the filtering and send the generated information to the entropy encoder 240. The information about the filtering can be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0076] The modified reconstructed picture sent to the memory 270 can be used as a reference picture in the inter-frame predictor 221. By doing so, the encoding device can avoid prediction mismatches in the encoding device 200 and the decoding device when applying inter-frame prediction, and can also improve the coding efficiency.
[0077] The memory 270 DPB can store the modified reconstructed picture to use it as a reference picture in the inter-frame predictor 221. The memory 270 can store the motion information of the blocks in the current picture from which the motion information (or for which it is encoded) is derived and / or the motion information of the blocks in the picture that has been (or previously) reconstructed. The stored motion information can be sent to the inter-frame predictor 221 to be used as the motion information of neighboring blocks or temporally neighboring blocks. The memory 270 can store the reconstructed samples of the reconstructed blocks in the current picture and can send the reconstructed samples to the intra-frame predictor 222.
[0078] Figure 3FIG. is a diagram schematically illustrating a configuration of a video / image decoding apparatus to which the present disclosure can be applied.
[0079] Reference Figure 3 , the video decoding apparatus 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-predictor 331 and an intra-predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. The entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 described above may be configured by one or more hardware components (e.g., a decoder chipset or a processor) according to an embodiment. In addition, the memory 360 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.
[0080] When receiving a bitstream including video / image information, the decoding apparatus 300 may reconstruct an image corresponding to a process of processing the video / image information in an Figure 2 encoding apparatus that has already been processed. For example, the decoding apparatus 300 may derive units / blocks based on information related to block partitioning obtained from the bitstream. The decoding apparatus 300 may perform decoding by using a processing unit applied in the encoding apparatus. Thus, the decoding processing unit may be, for example, a coding unit that may be partitioned from a coding tree unit or a largest coding unit along a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. And, a reconstructed image signal decoded and output by the decoding apparatus 300 may be reproduced by a reproducer.
[0081] The decoding apparatus 300 is capable of receiving, in the form of a bitstream, from Figure 2The signal output by the encoding device, and the received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information can also include information about various parameter sets such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), video parameter sets (VPS), etc. In addition, the video / image information can also include general constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or the general constraint information. In the present disclosure, the information and / or syntax elements sent / received by signal to be described later can be decoded through the decoding process and obtained from the bitstream. For example, the entropy decoder 310 can decode the information in the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, CABAC, etc., and can output the values of the syntax elements necessary for image reconstruction and the quantization values of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive the bins corresponding to each syntax element in the bitstream, use the decoding target syntax element information and the decoding information of the neighboring and decoding target blocks or the information of the symbols / bins decoded in the previous step to determine the context model, predict the bin generation probability according to the determined context model, and perform arithmetic decoding of the bins to generate symbols corresponding to each syntax element value. Here, the CABAC entropy decoding method can update the context model using the information of the symbols / bins decoded by the context model for the next symbol / bin after determining the context model. The information about prediction among the information decoded in the entropy decoder 310 can be provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values, i.e., the quantized transform coefficients and the related parameter information, for which entropy decoding has been performed in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive a residual signal (residual block, residual sample, residual sample array). In addition, the information about filtering among the information decoded in the entropy decoder 310 can be provided to the filter 350. Meanwhile, a receiver (not shown) that receives the signal output by the encoding device can further configure the decoding device 300 as internal / external components, and the receiver can be a component of the entropy decoder 310. Meanwhile, the decoding device according to the present disclosure can be referred to as a video / image / picture compilation device, and the decoding device can be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoder 310, and the sample decoder can include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0082] The dequantizer 321 can output transform coefficients by dequantizing the quantized transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement process can be performed based on the order of coefficient scanning executed in the encoding device. The dequantizer 321 can use quantization parameters (e.g., quantization step information) to perform dequantization on the quantized transform coefficients and obtain the transform coefficients.
[0083] The inverse transformer 322 obtains a residual signal (residual block, residual sample array) by performing an inverse transform on the transform coefficients.
[0084] The predictor can perform prediction on the current block and generate a prediction block including prediction samples of the current block. The predictor can determine whether to apply intra prediction or inter prediction to the current block based on the information about prediction output from the entropy decoder 310, and more specifically, the predictor can determine the intra / inter prediction mode.
[0085] The predictor can generate a prediction signal based on various prediction methods. For example, the predictor can apply intra prediction or inter prediction for predicting a block, and can also apply intra prediction and inter prediction simultaneously. This can be referred to as combined inter and intra prediction (CIIP). Additionally, the predictor can perform intra block copy (IBC) for predicting a block. Intra block copy can be used for content image / video compilation in games, etc., such as screen content compilation (SCC). Although IBC basically performs prediction in the current block, the aspect of deriving a reference block within the current block can be performed similarly to inter prediction. That is, IBC can use at least one of the inter prediction techniques described in the present disclosure.
[0086] The intra predictor 331 can predict the current block by referring to samples in the current picture. Depending on the prediction mode, the reference samples can be located in the neighboring region of the current block or in a distant region from the current block. In intra prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra predictor 331 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring block.
[0087] The inter-frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter-frame prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, the neighboring blocks can include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter-frame predictor 332 can configure a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction can include information indicating the inter-frame prediction mode for the current block.
[0088] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor 330. When there is no residual for the target block being processed, as in the case of applying the skip mode, the prediction block can be used as the reconstructed block.
[0089] The adder 340 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next target block in the current block, and as will be described later, the generated reconstructed signal can be output through filtering or used for inter-frame prediction of the next picture.
[0090] Meanwhile, during picture decoding, luminance mapping and chrominance scaling (LMCS) can be applied.
[0091] The filter 350 can improve the subjective / objective video quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and can send the modified reconstructed picture to the memory 360, more specifically, to the DPB of the memory 360. Various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0092] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter-frame predictor 332. The memory 360 can store the motion information of the blocks in the current picture from which the motion information has been derived (or decoded) and / or the motion information of the blocks in the pictures that have been (or previously) reconstructed. The stored motion information can be sent to the inter-frame predictor 260 for use as the motion information of neighboring blocks or temporally neighboring blocks. The memory 360 can store the reconstructed samples of the reconstructed blocks in the current picture and can send the reconstructed samples to the intra-frame predictor 331.
[0093] In this specification, the examples described in the predictor 330, dequantizer 321, inverse transformer 322, and filter 350 of the decoding device 300 can be similarly or correspondingly applied to the predictor 220, dequantizer 234, inverter 235, and filter 260 of the encoding device 200, respectively.
[0094] As described above, prediction is performed to improve the compression efficiency during video compilation. In doing so, a prediction block including prediction samples for the current block that is the target block for compilation can be generated. Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block can be derived in the same way in both the encoding device and the decoding device, and the encoding device can improve the image compilation efficiency by signaling to the decoding device information about the residual between the original block and the prediction block (residual information) instead of the original sample values of the original block itself. The decoding device can derive a residual block including residual samples based on the residual information, generate a reconstructed block including reconstructed samples by adding the residual block to the prediction block, and generate a reconstructed picture including the reconstructed block.
[0095] The residual information can be generated through a transformation and quantization process. For example, the encoding device can derive a residual block between the original block and the prediction block, derive transform coefficients by performing a transformation process on the residual samples (residual sample array) included in the residual block, and derive quantized transform coefficients by performing a quantization process on the transform coefficients so that it can signal the relevant residual information to the decoding device (through the bitstream). Here, the residual information can include value information, position information, transformation technique, transformation kernel, quantization parameter, etc. of the quantized transform coefficients. The decoding device can perform a quantization / dequantization process and derive residual samples (or residual sample blocks) based on the residual information. The decoding device can generate a reconstructed block based on the prediction block and the residual block. The encoding device can dequantize / inverse-transform the quantized transform coefficients to derive a residual block for reference in the inter-frame prediction of the next picture, and can generate a reconstructed picture based on the derived residual block.
[0096] Figure 4 The figure illustrates the structure of a content streaming system applying the present disclosure.
[0097] In addition, the content streaming system applying the present disclosure may generally include an encoding server, a streaming server, a web server, a media storage, user devices, and a multimedia input device.
[0098] The encoding server is used to compress the content input from a multimedia input device such as a smart phone, a camera, a video camera, etc. into digital data to generate a bitstream, and send it to the streaming server. As another example, in the case where a multimedia input device such as a smart phone, a camera, a video camera, etc. directly generates a bitstream, the encoding server may be omitted. The bitstream can be generated by applying the encoding method or the bitstream generation method of the present disclosure. And the streaming server may temporarily store the bitstream during the process of sending or receiving the bitstream.
[0099] The streaming server sends multimedia data to the user device through the web server based on the user's request, and the web server serves as a means for notifying the user of what services exist. When the user requests a service that the user wants, the web server transmits the request to the streaming server, and the streaming server sends the multimedia data to the user. In this regard, the content streaming system may include a separate control server, and in this case, the control server is used to control the commands / responses between the corresponding devices in the content streaming system.
[0100] The streaming server may receive content from a media storage device and / or an encoding server. For example, in the case of receiving content from the encoding server, the content may be received in real time. In this case, in order to smoothly provide the streaming service, the streaming server may store the bitstream for a predetermined time.
[0101] For example, the user device may include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigator, a slate PC, a tablet PC, a superbook, a wearable device (e.g., a watch-type terminal (smart watch), a glasses-type terminal (smart glasses), a head-mounted display (HMD)), a digital TV, a desktop computer, a digital signage, etc. Each server in the content streaming system may operate as a distributed server, and in this case, the data received by each server may be processed in a distributed manner.
[0102] When performing intra prediction, the correlation between samples can be used, and the difference between the original block and the predicted block, i.e., the residual, can be obtained. The above-mentioned transformation and quantization can be applied to the residual, so the spatial redundancy can be reduced. Hereinafter, the encoding method and the decoding method using intra prediction are specifically described.
[0103] Intra prediction refers to prediction for generating prediction samples for a current block based on reference samples outside the current block within a picture (hereinafter, the current picture) including the current block. In this case, the reference samples outside the current block may refer to samples adjacent to the current block. When intra prediction is applied to the current block, neighboring reference samples to be used for intra prediction of the current block can be derived.
[0104] For example, when the size (width × height) of the current block is nW × nH, the neighboring reference samples of the current block may include: a total of 2 × nH samples including samples adjacent to the left boundary of the current block and lower-left neighboring samples, a total of 2 × nW samples including samples adjacent to the upper boundary of the current block and upper-right neighboring samples, and one upper-left neighboring sample of the current block. Alternatively, the neighboring reference samples of the current block may include multiple rows of upper neighboring samples and multiple columns of left neighboring samples. In addition, the neighboring reference samples of the current block may include a total of nH samples adjacent to the right boundary of the current block of size nW × nH, a total of nW samples adjacent to the lower boundary of the current block, and one lower-right neighboring sample of the current block.
[0105] Here, some of the neighboring reference samples of the current block may not have been decoded or may be unavailable. In this case, the decoding device can configure the neighboring reference samples to be used for prediction by replacing the unavailable samples with available samples. Alternatively, the decoding device can configure the neighboring reference samples to be used for prediction by interpolation of available samples.
[0106] When the neighboring reference samples are derived, (i) the prediction samples can be derived based on the average or interpolation of the neighboring reference samples of the current block, or (ii) the prediction samples can be derived based on the reference samples among the neighboring reference samples of the current block that exist in a specific (prediction) direction for the prediction samples. (i) can be applied when the intra prediction mode is a non-directional mode or a non-angle mode, while (ii) can be applied when the intra prediction mode is a directional mode or an angle mode.
[0107] The intra prediction mode can include a non-directional (or non-angle) intra prediction mode and a directional (or angle) intra prediction mode. For example, in HEVC, an intra prediction mode including two non-directional intra prediction modes and 33 directional intra prediction modes is used. The non-directional intra prediction modes can include a planar intra prediction mode as intra prediction mode 0 and a DC intra prediction mode as intra prediction mode 1, while the directional intra prediction modes can include intra prediction modes 2 to 34. The planar intra prediction mode can be referred to as the planar mode, and the DC intra prediction mode can be referred to as the DC mode.
[0108] Alternatively, in order to capture any edge direction presented in natural videos, as described below in Figure 10As shown in [reference], the number of directional intra prediction modes is extended from 33 to 65. In this case, the intra prediction modes may include two non-directional intra prediction modes and 65 directional intra prediction modes. The non-directional intra prediction modes may include a planar intra prediction mode as intra prediction mode 0 and a DC intra prediction mode as intra prediction mode 1, while the directional intra prediction modes may include intra prediction modes 2 to 66. The extended directional intra prediction modes may be applied to blocks of any size and may be applied to both the luminance component and the chrominance component. However, the above example is for illustration, and the embodiments of this document may also be applied when the number of intra prediction modes is different. Intra prediction mode 67 may be further used depending on the situation, and intra prediction mode 67 may refer to a linear model (LM) mode.
[0109] Figure 5 An example of an intra prediction mode to which an embodiment of this document is applicable is illustrated.
[0110] Reference Figure 5 , the intra prediction modes may be divided into intra prediction modes with horizontal directivity and intra prediction modes with vertical directivity based on intra prediction mode 34 having an upper left diagonal prediction direction. In Figure 5 , H and V respectively indicate horizontal directivity and vertical directivity, and each of the numbers -32 to 32 indicates a displacement in 1 / 32 units at the sample grid position. Intra prediction modes 2 to 33 have horizontal directivity, while intra prediction modes 34 to 66 have vertical directivity. Intra prediction mode 18 and intra prediction mode 50 respectively refer to a horizontal intra prediction mode and a vertical intra prediction mode. Intra prediction mode 2 may be referred to as a lower left diagonal intra prediction mode, intra prediction mode 34 may be referred to as an upper left diagonal intra prediction mode, and intra prediction mode 66 may be referred to as an upper right diagonal intra prediction mode.
[0111] Matrix-based intra prediction (hereinafter, MIP) may be used as a method for intra prediction. MIP may be referred to as affine linear weighted intra prediction (ALWIP) or matrix weighted intra prediction (MWIP).
[0112] When MIP is applied to a current block, the predicted samples for the current block may be derived by: i) using neighboring reference samples that have undergone an averaging process to ii) perform a matrix-vector multiplication process, and (iii) further perform a horizontal / vertical interpolation process if necessary. The intra prediction modes for MIP may be configured differently from the intra prediction modes used in the foregoing LIP, PDPC, MRL, or ISP intra prediction or the intra prediction modes used in normal intra prediction.
[0113] The intra prediction mode for MIP may be referred to as an affine linear weighted intra prediction mode or a matrix-based intra prediction mode. For example, the matrix and offset used in matrix-vector multiplication may be configured differently depending on the intra prediction mode for MIP. Here, the matrix may be referred to as an (affine) weighting matrix, and the offset may be referred to as an (affine) offset vector or an (affine) bias vector. In the present disclosure, the intra prediction mode for MIP may be referred to as the MIP intra prediction mode, the affine linear weighted intra prediction (ALWIP) mode, the matrix weighted intra prediction (MWI) mode, or the matrix-based intra prediction mode.
[0114] To predict samples of a rectangular block having a width (W) and a height (H), MIP uses one H-line among the reconstructed samples adjacent to the left boundary of the block and one W-line among the reconstructed samples adjacent to the upper boundary of the block as input values. When no reconstructed samples are available, reference samples may be generated by an interpolation method applied to general intra prediction.
[0115] Figure 6 The figure illustrates a process of generating prediction samples based on MIP according to an example. Refer to Figure 6 The MIP process is described as follows.
[0116] 1. Averaging process
[0117] Among the boundary samples, four samples for the case of W = H = 4 and eight samples for any other case are extracted by an averaging process.
[0118] 2. Matrix-vector multiplication process
[0119] The matrix-vector multiplication is performed with the averaged samples as input and then adding the offset. Through this operation, reduced prediction samples for the subsampled sample set in the original block can be derived.
[0120] 3. (Linear) interpolation process
[0121] The prediction samples at the remaining positions are generated by linear interpolation from the prediction samples of the subsampled sample set, and the linear interpolation is a single-step linear interpolation in each direction.
[0122] The matrix and offset vector necessary for generating the prediction block or prediction samples may be selected from three sets S 0 、S 1 and S 2 .
[0123] Set S 0 may include: 16 matrices A 0 i , i ∈ {0, …, 15}, and each matrix may include 16 rows and four columns; and 16 offset vectors b0 i , where \(i\in\{0,\ldots,15\}\). The matrices and offset vectors of set \(S\) 0 can be used for \(4\times4\) blocks. In another example, set \(S\) 0 can include 18 matrices.
[0124] Set \(S\) 1 can include: eight matrices \(A\) 1 i , where \(i\in\{0,\ldots,7\}\), and each matrix can include 16 rows and eight columns; and eight offset vectors \(b\) 1 i , where \(i\in\{0,\ldots,7\}\). In another example, set \(S\) 1 can include six matrices. Set \(S\) 1 's matrices and offset vectors can be used for \(4\times8\), \(8\times4\), and \(8\times8\) blocks. Alternatively, the matrices and offset vectors of set \(S\) 1 can be used for \(4\times H\) or \(W\times4\) blocks.
[0125] Finally, set \(S\) 2 can include: six matrices \(A\) 2 i , where \(i\in\{0,\ldots,5\}\), and each matrix can include 64 rows and eight columns; and six offset vectors \(b\) 2 i , where \(i\in\{0,\ldots,5\}\). The matrices and offset vectors of set \(S\) 2 or a part of them can be used for any blocks with different sizes that sets \(S\) 0 and set \(S\) 1 are not applicable to. For example, the matrices and offset vectors of set \(S\) 2 can be used for the operation of blocks with a height and width of 8 or greater.
[0126] The total number of multiplications required for the calculation of the matrix - vector product is always less than or equal to \(4\times W\times H\). That is, in the MIP mode, up to four multiplications are required per sample.
[0127] In addition, the current block can be divided into vertical or horizontal sub - partitions, and intra - frame prediction can be performed based on the same intra - frame prediction mode, where neighboring reference samples can be derived and used in units of sub - partitions. That is, in this case, the intra - frame prediction mode for the current block can be equally applied to the sub - partitions, and neighboring reference samples can be derived and used in units of sub - partitions, thereby improving the performance of intra - frame prediction in some cases. This prediction method can be called intra - frame sub - partition (ISP) or ISP - based intra - frame prediction.
[0128] Intra Sub - Partition (ISP) compilation refers to performing intra - prediction compilation by splitting the block to be currently compiled in the horizontal or vertical direction. In this case, the reconstructed block can be generated by performing encoding / decoding in units of the split blocks, and the reconstructed block can be used as a reference block for the next split block. According to an example, in ISP compilation, a compilation block can be split into two or four sub - blocks and then can be compiled, and in ISP, the reconstructed pixel values of a sub - block referring to the left - adjacent sub - block or the upper - adjacent sub - block are subject to intra - prediction. As used herein, "compilation" can be used as a concept including both the compilation performed by an encoding device and the decoding performed by a decoding device.
[0129] Table 1 shows the number of sub - blocks split according to the size of the block when ISP is applied, and the sub - partitions split according to ISP can be referred to as transform units (TUs).
[0130] [Table 1]
[0131] Block size (CU) Number of partitions 4×4 Not available 4×8、8×4 2 Any other case 4
[0132] ISP splits the block through luminance intra - prediction into two or four sub - partitions in the vertical or horizontal direction according to the size of the block. For example, the minimum size of the block to which ISP is applicable is 4×8 or 8×4. When the size of the block is greater than 4×8 or 8×4, the block is split into four sub - partitions.
[0133] Figure 7 and Figure 8 illustrates an example of a compilation block being split into sub - blocks. Specifically, Figure 7 illustrates an example of splitting a compilation block (width (W)×height (H)) that is a 4×8 block or an 8×4 block, and Figure 8 illustrates an example of splitting a compilation block that is not a 4×8 block, an 8×4 block, or a 4×4 block.
[0134] When ISP is applied, the sub - blocks can be compiled sequentially according to the split type, such as horizontally, vertically, from left to right, or from top to bottom. One sub - block can undergo inverse transformation, intra - prediction, and reconstruction, and then the next sub - block can be compiled. For the left - most or top - most sub - block, the reconstructed pixels of the already - compiled compilation block are used as a reference as in the conventional intra - prediction method. In addition, when each side of a subsequent internal sub - block is not adjacent to the previous sub - block, the reconstructed pixels of the already - compiled adjacent compilation blocks are used as a reference to derive the reference pixels adjacent to that side, as in the conventional intra - prediction method.
[0135] In the ISP compilation mode, all sub - blocks can be compiled with the same intra - prediction mode, and a flag indicating whether ISP compilation is used and a flag indicating the direction of the split (horizontal or vertical) can be signaled. As Figure 7 andFigure 8 As shown, the number of sub - blocks can be adjusted to 2 or 4 depending on the block shape, and when the size (width × height) of a sub - block is less than 16, splitting into the corresponding number of sub - blocks may not be allowed or a limit not to apply ISP compilation may be set.
[0136] In the ISP prediction mode, a compilation unit is split into two or four partitioned blocks, i.e., sub - blocks, to be predicted, and the same intra - prediction mode is applied to these two or four split blocks.
[0137] Figure 9 Schematically illustrate the multi - transform technology according to an embodiment of the present disclosure.
[0138] Refer to Figure 9 , the transformer can correspond to the transformer in the aforementioned Figure 2 encoding device, and the inverse transformer can correspond to the inverse transformer in the aforementioned Figure 2 encoding device, or correspond to the inverse transformer in Figure 3 the decoding device.
[0139] The transformer can derive (primary) transform coefficients (S910) by performing a primary transform based on the residual samples (residual sample array) in the residual block. This primary transform can be referred to as the core transform. Herein, the primary transform can be based on multi - transform selection (MTS), and when multi - transform is applied as the primary transform, it can be referred to as multi - core transform.
[0140] The multi - core transform can represent a method of additionally using Discrete Cosine Transform (DCT) type 2 and Discrete Sine Transform (DST) type 7, DCT type 8, and / or DST type 1 for transformation. That is, the multi - core transform can represent a transform method of transforming a residual signal (or residual block) in the spatial domain into transform coefficients (or primary transform coefficients) in the frequency domain based on multiple transform kernels selected from DCT type 2, DST type 7, DCT type 8, and DST type 1. Herein, from the perspective of the transformer, the primary transform coefficients can be referred to as temporary transform coefficients.
[0141] In other words, when applying the conventional transform method, transform coefficients can be generated by applying a transform from the spatial domain to the frequency domain to the residual signal (or residual block) based on DCT type 2. In contrast, when applying the multi - core transform, transform coefficients (or primary transform coefficients) can be generated by applying a transform from the spatial domain to the frequency domain to the residual signal (or residual block) based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1. Herein, DCT type 2, DST type 7, DCT type 8, and DST type 1 can be referred to as transform types, transform kernels, or transform cores. These DCT / DST transform types can be defined based on basis functions.
[0142] When performing multi-core transformation, a vertical transformation kernel and a horizontal transformation kernel for a target block can be selected from among the transformation kernels, a vertical transformation can be performed on the target block based on the vertical transformation kernel, and a horizontal transformation can be performed on the target block based on the horizontal transformation kernel. Here, the horizontal transformation can indicate a transformation of the horizontal component of the target block, and the vertical transformation can indicate a transformation of the vertical component of the target block. The vertical transformation kernel / horizontal transformation kernel can be adaptively determined based on the prediction mode and / or transformation index of the target (CU or sub-block) including the residual block.
[0143] In addition, according to an example, if the primary transformation is performed by applying MTS, the mapping relationship of the transformation kernel can be set by setting a specific basis function to a predetermined value and combining the basis functions to be applied in the vertical transformation or the horizontal transformation. For example, when the horizontal transformation kernel is represented as trTypeHor and the vertical transformation kernel is represented as trTypeVer, trTypeHor or trTypeVer with a value of 0 can be set to DCT2, trTypeHor or trTypeVer with a value of 1 can be set to DST-7, and trTypeHor or trTypeVer with a value of 2 can be set to DCT-8.
[0144] In this case, the MTS index information can be encoded and signaled to the decoding device to indicate any one of the multiple transformation kernel sets. For example, MTS index 0 can indicate that both trTypeHor and trTypeVer values are 0, MTS index 1 can indicate that both trTypeHor and trTypeVer values are 1, MTS index 2 can indicate that the trTypeHor value is 2 and the trTypeVer value is 1, MTS index 3 can indicate that the trTypeHor value is 1 and the trTypeVer value is 2, and MTS index 4 can indicate that both trTypeHor and trTypeVer values are 2.
[0145] In one example, the transformation kernel sets according to the MTS index information are shown in the following table.
[0146] [Table 2]
[0147] tu_mts_idx[x0][y0] 0 1 2 3 4 trTypeHor 0 1 2 1 2 trTypeVer 0 1 1 2 2
[0148] The transformer may perform a secondary transform based on a (primary) transform coefficient to derive a modified (secondary) transform coefficient (S920). The primary transform is a transform from the spatial domain to the frequency domain, and the secondary transform refers to a transform into a more compressed expression using the correlation existing between the (primary) transform coefficients. The secondary transform may include a non-separable transform. In this case, the secondary transform may be referred to as a non-separable secondary transform (NSST) or a mode-dependent non-separable secondary transform (MDNSST). The non-separable secondary transform may represent a transform of (primary) transform coefficients derived by a primary transform based on a non-separable transform matrix to generate modified transform coefficients (or secondary transform coefficients) for a residual signal. Here, instead of applying the vertical transform and the horizontal transform separately (or independently) to the (primary) transform coefficients, the transform may be applied at once based on the non-separable transform matrix. In other words, the non-separable secondary transform may represent a transform method in which the (primary) transform coefficients in the vertical and horizontal directions are not separately applied, and a two-dimensional signal (transform coefficients) may be rearranged into a one-dimensional signal, for example, in a specific predetermined direction (e.g., row-first direction or column-first direction), and then modified transform coefficients (or secondary transform coefficients) may be generated based on the non-separable transform matrix. For example, according to the row-first order, the MxN blocks are arranged in a line in the order of the first row, the second row, …, the Nth row. According to the column-first order, the MxN blocks are arranged in a line in the order of the first column, the second column, …, the Mth column. The NSST may be applied to the upper left region of a block configured with (primary) transform coefficients (hereinafter, referred to as a transform coefficient block). For example, when both the width W and the height H of the transform coefficient block are 8 or more, the 8×8 NSST may be applied to the upper left 8×8 region of the transform coefficient block. In addition, while both the width (W) and the height (H) of the transform coefficient block are 4 or more, when the width (W) or the height (H) of the transform coefficient block is less than 8, the 4×4 NSST may be applied to the upper left min(8,W)×min(8,H) region of the transform coefficient block. However, the embodiments are not limited thereto, and for example, even if only the width W or the height H of the transform coefficient block satisfies the condition of being 4 or more, the 4×4 NSST may be applied to the upper left min(8,W)×min(8,H) region of the transform coefficient block.
[0149] Here, in order to select the transform kernels, for both the 8×8 transform and the 4×4 transform, two non-separable secondary transform kernels can be configured for each transform set for the non-separable secondary transform, and there can be four transform sets. That is, four transform sets can be configured for the 8×8 transform, and four transform sets can be configured for the 4×4 transform. In this case, each of the four transform sets for the 8×8 transform can include two 8×8 transform kernels, and each of the four transform sets for the 4×4 transform can include two 4×4 transform kernels.
[0150] However, as the size of the transform (i.e., the size of the region to which the transform is applied) can be, for example, a size other than 8×8 or 4×4, the number of sets can be n, and the number of transform kernels in each set can be k.
[0151] The transform sets can be referred to as NSST sets or LFNST sets. A specific set among the transform sets can be selected, for example, based on the intra prediction mode of the current block (CU or sub-block). The low-frequency non-separable transform (LFNST) can be an example of a reduced non-separable transform, which will be described later, and represents a non-separable transform for the low-frequency components.
[0152] According to an example, four transform sets can be mapped according to the intra prediction mode, for example, as shown in the following table.
[0153] [Table 3]
[0154] predModeIntra lfnstTrSetIdx predModeIntra < 0 1 0 <= predModeIntra <= 1 0 2 <= predModeIntra <= 12 1 13 <= predModeIntra <= 23 2 24 <= predModeIntra <= 44 3 45 <= predModeIntra <= 55 2 56 <= predModeIntra <= 80 1
[0155] As shown in Table 3, according to the intra prediction mode, any one of the four transform sets, i.e., lfnstTrSetIdx, can be mapped to any one of the four indices (i.e., 0 to 3).
[0156] When a specific set is determined for the non-separable transform, one of the k transform kernels in the specific set can be selected by the non-separable secondary transform index. The encoding device can derive the non-separable secondary transform index indicating the specific transform kernel based on rate-distortion (RD) checking, and can signal the non-separable secondary transform index to the decoding device. The decoding device can select one of the k transform kernels in the specific set based on the non-separable secondary transform index. For example, the lfnst index value 0 can refer to the first non-separable secondary transform kernel, the lfnst index value 1 can refer to the second non-separable secondary transform kernel, and the lfnst index value 2 can refer to the third non-separable secondary transform kernel. Alternatively, the lfnst index value 0 can indicate that the first non-separable secondary transform is not applied to the target block, and the lfnst index values 1 to 3 can indicate three transform kernels.
[0157] The transformer can perform an inseparable secondary transform based on the selected transform kernel and can obtain modified (secondary) transform coefficients. As described above, the modified transform coefficients can be derived as the transform coefficients quantized by the quantizer, and can be encoded and signaled to the decoding device, and transmitted to the dequantizer / inverse transformer in the encoding device.
[0158] Meanwhile, as described above, if the secondary transform is omitted, the (primary) transform coefficients that are the output of the primary (separable) transform can be derived as the transform coefficients quantized by the quantizer as described above, and can be encoded and signaled to the decoding device, and transmitted to the dequantizer / inverse transformer in the encoding device.
[0159] The inverse transformer is capable of performing a series of processes in an order opposite to the order in which the series of processes are performed in the above-mentioned transformer. The inverse transformer can receive the (dequantized) transformer coefficients, and derive the (primary) transform coefficients by performing the secondary (inverse) transform (S950), and can obtain the residual block (residual samples) by performing the primary (inverse) transform on the (primary) transform coefficients (S960). In this regard, from the perspective of the inverse transformer, the primary transform coefficients can be referred to as modified transform coefficients. As described above, the encoding device and the decoding device can generate a reconstructed block based on the residual block and the prediction block, and can generate a reconstructed picture based on the reconstructed block.
[0160] The decoding device may further include a secondary inverse transform application determiner (or an element for determining whether to apply the secondary inverse transform) and a secondary inverse transform determiner (or an element for determining the secondary inverse transform). The secondary inverse transform application determiner can determine whether to apply the secondary inverse transform. For example, the secondary inverse transform can be NSST, RST, or LFNST, and the secondary inverse transform application determiner can determine whether to apply the secondary inverse transform based on the secondary transform flag obtained by parsing the bitstream. In another example, the secondary inverse transform application determiner can determine whether to apply the secondary inverse transform based on the transform coefficients of the residual block.
[0161] The secondary inverse transform determiner can determine the secondary inverse transform. In this case, the secondary inverse transform determiner can determine the secondary inverse transform to be applied to the current block based on the LFNST (NSST or RST) transform set specified according to the intra prediction mode. In an embodiment, the secondary transform determination method can be determined depending on the primary transform determination method. Various combinations of the primary transform and the secondary transform can be determined according to the intra prediction mode. In addition, in an example, the secondary inverse transform determiner can determine the region to which the secondary inverse transform is applied based on the size of the current block.
[0162] Meanwhile, as described above, if the secondary (inverse) transform is omitted, the (dequantized) transform coefficients can be received, the primary (separable) inverse transform can be performed, and the residual block (residual samples) can be obtained. As described above, the encoding device and the decoding device can generate a reconstructed block based on the residual block and the prediction block, and can generate a reconstructed picture based on the reconstructed block.
[0163] In this document, a reduced secondary transform (RST) in which the size of the transform matrix (kernel) is reduced can be applied in the concept of NSST in order to reduce the amount of calculation and memory required for the non-separable secondary transform. The RST is generally performed in the low-frequency region including non-zero coefficients in the transform block, and thus the RST can be referred to as a low-frequency non-separable transform (LFNST). The transform index can be referred to as the LFNST index.
[0164] In this specification, the LFNST can refer to a transform performed on the residual samples for a target block based on a transform matrix having a reduced size. In the case of performing the reduced transform, due to the reduction in the size of the transform matrix, the amount of calculation required for the transform can be reduced. That is, the LFNST can be used to solve the computational complexity problem that occurs under non-separable transforms or transforms of very large blocks.
[0165] When performing the secondary inverse transform based on the LFNST, the inverse transformers 235 of the encoding device 200 and 322 of the decoding device 300 can include: an inverse reduced secondary transformer that derives modified transform coefficients based on the inverse RST of the transform coefficients; and an inverse primary transformer that derives the residual samples for the target block based on the inverse primary transform of the modified transform coefficients. The inverse primary transform refers to the inverse transform of the primary transform applied to the residuals. In this document, deriving the transform coefficients based on a transform can refer to deriving the transform coefficients by applying the transform.
[0166] Figure 10 FIG. is a diagram illustrating RST according to an embodiment of the present disclosure.
[0167] In the present disclosure, a "target block" can refer to the current block to be encoded, a residual block, or a transform block.
[0168] In the RST according to the example, an N-dimensional vector can be mapped to an R-dimensional vector located in another space, and thus a reduced transform matrix can be determined, where R is less than N. N can refer to the square of the length of the side of the block to which the transform is applied, or the total number of transform coefficients corresponding to the block to which the transform is applied, and the reduction factor can refer to the R / N value. The reduction factor can be referred to by various terms such as reduction factor, shrinking factor, simplification factor, simple factor, or other various terms. In addition, R can be referred to as the reduction coefficient, but depending on the situation, the reduction factor can refer to R. In addition, depending on the situation, the reduction factor can refer to the N / R value.
[0169] In the example, the reduction factor or reduction coefficient may be signaled by a bitstream, but the example is not limited thereto. For example, predefined values for the reduction factor or reduction coefficient may be stored in each of the encoding device 200 and the decoding device 300, and in this case, the reduction factor or reduction coefficient may not be signaled separately.
[0170] The size of the reduction transform matrix according to the example may be R×N which is smaller than N×N (the size of the conventional transform matrix), and may be defined as in Equation 1 below.
[0171] [Equation 1]
[0172]
[0173] Figure 10 The matrix T in the reduction transform block shown in (a) of may mean the matrix T of Equation 1 RxN . As Figure 10 shown in (a) of, when multiplying the reduction transform matrix T RxN by the residual samples of the target block, the transform coefficients for the target block may be derived.
[0174] In the example, if the size of the block to which the transform is applied is 8x8 and R = 16 (i.e., R / N = 16 / 64 = 1 / 4), then the RST according to Figure 10 (a) can be expressed as a matrix operation as shown in Equation 2 below. In this case, the memory and multiplication calculations can be reduced to approximately 1 / 4 by the reduction factor.
[0175] In the present disclosure, a matrix operation may be understood as an operation of obtaining a column vector by multiplying a column vector by a matrix set on the left side of the column vector.
[0176] [Equation 2]
[0177]
[0178] In Equation 2, r 1 to r 64 may represent the residual samples of the target block, and specifically may be the transform coefficients generated by applying a primary transform. As a result of the calculation of Equation 2, the transform coefficients c i of the target block can be derived, and the process of deriving c i can be as shown in Equation 3.
[0179] [Equation 3]
[0180]
[0181] As a result of the calculation of Equation 3, the transform coefficients c for the target block can be derived. 1 through c R . That is, when R = 16, the transform coefficients c for the target block can be derived. 1 through c 16 . Although 64 (N) transform coefficients are derived for the target block, if a regular transform is applied instead of RST and a transform matrix of size 64x64 (NxN) is multiplied by the residual samples of size 64x1 (Nx1), only 16 (R) transform coefficients are derived for the target block because RST is applied. Since the total number of transform coefficients for the target block is reduced from N to R, the amount of data transmitted from the encoding device 200 to the decoding device 300 is reduced, thus improving the efficiency of the transmission between the encoding device 200 and the decoding device 300.
[0182] When considered from the perspective of the size of the transform matrix, the size of the regular transform matrix is 64×64 (N×N), but the size of the reduced transform matrix is reduced to 16×64 (R×N). Therefore, the storage utilization rate in the case of performing LFNST can be reduced by the ratio of R / N compared to the case of performing a regular transform. In addition, when compared with the number of multiplication calculations N×N in the case of using a regular transform matrix, using the reduced transform matrix can reduce the number of multiplication calculations (R×N) by the ratio of R / N.
[0183] In the example, the transformer 232 of the encoding device 200 can derive the transform coefficients for the target block by performing a primary transform and an RST-based secondary transform on the residual samples for the target block. These transform coefficients can be passed to the inverse transformer of the decoding device 300, and the inverse transformer 322 of the decoding device 300 can derive the modified transform coefficients based on the inverse reduced secondary transform (RST) for the transform coefficients, and can derive the residual samples for the target block based on the inverse primary transform for the modified transform coefficients.
[0184] The size of the inverse RST matrix T NxR according to the example is NxR, which is smaller than the size NxN of the regular inverse transform matrix, and is in a transposed relationship with the reduced transform matrix T RxN shown in Equation 4.
[0185] Figure 10 The matrix T t in the reduced inverse transform block shown in (b) of RxN T can mean the inverse RST matrix T Figure 10 shown in (b) of RxN TWhen multiplying with the transform coefficients for a target block, modified transform coefficients for the target block or residual samples for the current block can be derived. The inverse RST matrix T RxN T can be expressed as (T RxN )T NxR .
[0186] More specifically, when the inverse RST is used as a secondary inverse transform, when the inverse RST matrix T N×R T is multiplied by the transform coefficients of the target block, modified transform coefficients of the target block can be derived. In addition, the inverse RST can be used as an inverse primary transform, and in this case, when the inverse RST matrix T N×R T is multiplied by the transform coefficients of the target block, residual samples of the target block can be derived.
[0187] In an example, if the size of the block to which the inverse transform is applied is 8x8 and R = 16 (i.e., R / N = 16 / 64 = 1 / 4), then the RST according to Figure 10 (b) can be expressed as a matrix operation as shown in Equation 4 below.
[0188] [Equation 4]
[0189]
[0190] In Equation 4, c 1 to c 16 can represent the transform coefficients of the target block. As a result of the calculation of Equation 4, r i , which represents modified transform coefficients of the target block or residual samples of the target block, can be derived, and the process of deriving r i can be as shown in Equation 5.
[0191] [Equation 5]
[0192]
[0193] As a result of the calculation of Equation 5, r 1 to r N , which represent modified transform coefficients of the target block or residual samples of the target block, can be derived. Considering from the perspective of the size of the inverse transform matrix, the size of the conventional inverse transform matrix is 64×64 (N×N), but the size of the inverse reduced transform matrix is reduced to 64×16 (R×N). Therefore, compared with the case of performing a conventional inverse transform, the storage utilization rate in the case of performing an inverse RST can be reduced by the R / N ratio. In addition, when compared with the number of multiplication calculations N×N in the case of using a conventional inverse transform matrix, using the inverse reduced transform matrix can reduce the number of multiplication calculations (N×R) by the R / N ratio.
[0194] According to an embodiment of the present disclosure, for the transform in the encoding process, only 48 pieces of data can be selected, and a maximum 16×48 transform kernel matrix can be applied thereto, instead of applying a 16×64 transform kernel matrix to 64 pieces of data forming an 8×8 region. Here, "maximum" means that m has a maximum value of 16 in the m×48 transform kernel matrix for generating m coefficients. That is to say, when performing RST by applying an m×48 transform kernel matrix (m≤16) to an 8×8 region, 48 pieces of data are input, and m coefficients are generated. When m is 16, 48 pieces of data are input and 16 coefficients are generated. That is to say, assuming that 48 pieces of data form a 48×1 vector, the 16×48 matrix and the 48×1 vector are multiplied in sequence, thereby generating a 16×1 vector. Here, the 48 pieces of data forming the 8×8 region can be appropriately arranged to form a 48×1 vector. For example, a 48×1 vector can be constructed based on 48 pieces of data constituting a region other than the lower-right 4×4 region among the 8×8 regions. Here, when performing matrix operations by applying a maximum 16×48 transform kernel matrix, 16 modified transform coefficients are generated, and the 16 modified transform coefficients can be arranged in the upper-left 4×4 region according to the scanning order, and the upper-right 4×4 region and the lower-left 4×4 region can be filled with zeros.
[0195] For the inverse transform in the decoding process, the transpose matrix of the aforementioned transform kernel matrix can be used. That is to say, when performing inverse RST or LFNST in the inverse transform process executed by a decoding device, the input coefficient data for applying inverse RST is configured in a one-dimensional vector according to a predetermined arrangement order, and the modified coefficient vector obtained by multiplying the one-dimensional vector by the corresponding inverse RST matrix on the left side of the one-dimensional vector can be arranged into a two-dimensional block according to a predetermined arrangement order.
[0196] In summary, in the transform process, when RST or LFNST is applied to an 8×8 region, matrix operations are performed on 48 transform coefficients in the upper-left region, upper-right region, and lower-left region of the 8×8 region except for the lower-right region with a 16×48 transform kernel matrix. For the matrix operations, 48 transform coefficients are input in a one-dimensional array. When the matrix operations are performed, 16 modified transform coefficients are derived, and the modified transform coefficients can be arranged in the upper-left region of the 8×8 region.
[0197] In contrast, in the inverse transformation process, when applying the inverse RST or LFNST to an 8×8 region, the 16 transform coefficients corresponding to the upper left region of the 8×8 region among the transform coefficients in the 8×8 region can be input in a one-dimensional array according to the scanning order, and can undergo matrix operations with a 48×16 transform kernel matrix. That is to say, the matrix operation can be expressed as (48×16 matrix) * (16×1 transform coefficient vector) = (48×1 modified transform coefficient vector). Here, an n×1 vector can be interpreted as having the same meaning as an n×1 matrix, and thus can be represented as an n×1 column vector. In addition, * represents matrix multiplication. When performing the matrix operation, 48 modified transform coefficients can be derived, and the 48 modified transform coefficients can be arranged in the upper left region, upper right region, and lower left region of the 8×8 region except for the lower right region.
[0198] According to the example, syntax elements for MIP can be signaled in the compilation unit syntax table as in Table 4.
[0199] [Table 4]
[0200]
[0201]
[0202]
[0203] The intra_mip_flag[x0][y0] signaled in the compilation unit syntax table of Table 4 is flag information indicating whether the MIP intra mode is applied to the compilation unit.
[0204] In addition to Table 4, as in Table 5, intra_mip_flag[x0][y0] is referenced multiple times in the specification text, and in particular, the value of intra_mip_flag for positions other than (x0,y0), such as intra_mip_flag[xCb+cbWidth / 2][yCb+cbHeight / 2], is also referenced.
[0205] Here, x0 and y0 respectively indicate the x coordinate and y coordinate based on the luma picture. When the leftmost position of the luma picture is defined as 0, the x coordinate increases from left to right in luma sample units, and when the uppermost position of the luma picture is defined as 0, the y coordinate increases from top to bottom in luma sample units. The x coordinate and y coordinate can be expressed in the form of two-dimensional coordinates such as (x,y).
[0206] However, intra_mip_flag[x0][y0] can be considered valid only for the (x0,y0) position, and (x0,y0) corresponds to the upper left position of the coding unit (CU) that signals intra_mip_flag[x0][y0]. That is, for each coding unit, it can be considered that intra_mip_flag[x0][y0] is signaled only for (x0,y0) which is the upper left position as a representative position in the coding unit.
[0207] In Table 5, since the intra_mip_flag values for positions other than (x0,y0) which is the upper left position in the coding unit are referenced (underlined), it is necessary to fill in the information of the intra_mip_flag values for positions other than the upper left position.
[0208] [Table 5]
[0209]
[0210]
[0211]
[0212]
[0213]
[0214] As in Tables 4 and 5, intra_mip_flag is defined as a syntax element, and as in the semantics of intra_mip_flag illustrated in Table 5 (7.4.11.5 Coding Unit Semantics), intra_mip_flag is inferred to be 0 when it does not exist.
[0215] The value of intra_mip_flag (8.4.1 and 8.4.2) can be used when deriving the intra-frame prediction mode. For example, as described in 8.4.2, the value of intra_mip_flag for a neighboring block or a specific position in a neighboring block (intra_mip_flag[xNbX][yNbX] equal to 1) can be used.
[0216] In addition, as described in 8.3.4, when deriving the intra-frame prediction mode for a chrominance block, the value of intra_mip_flag at the position of the corresponding luma block can be used (- the chrominance intra-frame prediction mode IntraPredModeC[xCb][yCb] is set to be equal to IntraPredModeY[xCb+cbWidth / 2][yCb+cbHeight / 2]).
[0217] In addition, the value of intra_mip_flag can be used not only in intra prediction but also in the transformation process (8.7.4.1 - when intra_mip_flag[xTbY][yTbY] is equal to 1 and cIdx is equal to 0, predModeIntra is set to be equal to INTRA_PLANAR), and the value of intra_mip_flag for neighboring blocks can be used in the process of deriving the context index of the target block to be compiled (Table 132: ctxInc specification using left and upper syntax elements).
[0218] Therefore, considering intra_mip_flag as a two-dimensional array variable, ambiguity may occur in the method of filling the intra_mip_flag array variable values at positions other than (x0, y0). That is, since the semantics stipulate that intra_mip_flag[x0][y0] is inferred to be 0 when it does not exist, filling any non-zero value in the unfilled positions in the intra_mip_flag array may be ambiguous, i.e., uncertain.
[0219] Therefore, according to the example, it is possible to propose defining a separate two-dimensional array, such as IntraLumaMipFlag[x][y] as shown in Table 6, and utilizing this two-dimensional array in the subsequent image compilation process.
[0220] [Table 6]
[0221]
[0222] In Table 6, (x0, y0) indicates the upper left position of the current compiled coding unit (CU), and cbWidth and cbHeight indicate the width and height of the coding unit, respectively. In addition, "x = x0..x0 + cbWidth - 1" means that the x-coordinate value changes from x0 to x0 + cbWidth - 1, while "y = y0..y0 + cbHeight - 1" means that the y-coordinate value changes from y0 to y0 + cbHeight - 1. Therefore, in the IntraLumaMipFlag[x][y] array in Table 6, the area corresponding to the coding unit is filled with the intra_mip_flag[x0][y0] value.
[0223] According to the example, all parts (underlined) in the current VVC specification text that reference the intra_mip_flag information can be replaced or modified with the IntraLumaMipFlag variable as shown in Table 7.
[0224] [Table 7]
[0225]
[0226]
[0227]
[0228]
[0229]
[0230] In Table 7, the existing intra_mip_flag can be used as it is to describe the case of determining the IntraLumaMipFlag information that references the upper-left position in the compilation unit. Table 8 only shows the part of the existing intra_mip_flag that can be used as extracted from Table 7.
[0231] [Table 8]
[0232]
[0233] In Table 4, the same problem may occur for intra_subpartitions_mode_flag other than intra_mip_flag in the current VVC specification text. That is, the intra_subpartitions_mode_flag information at positions other than the upper-left position (x0, y0) in the compilation unit is referenced, which is shown in Table 9 (indicated by the underline).
[0234] Therefore, a two-dimensional array variable of IntraSubPartitionsModeFlag[x][y] can be configured to be defined in a manner similar to the IntraLumaMipFlag variable, to fill all positions in the compilation unit with the necessary information and then reference the appropriate value for any position. That is, all positions in the compilation unit can be filled with the corresponding intra_subpartitions_mode_flag[x0][y0] value.
[0235] Table 10 shows using the IntraSubPartitionsModeFlag variable to replace the part (underline) that references intra_subpartitions_mode_flag with the IntraSubPartitionsModeFlag variable. In the case of IntraSubPartitionsModeFlag, when referencing the information about the upper-left position in the compilation unit, intra_subpartitions_mode_flag can be configured to be referenced as it is.
[0236] According to another example, the IntraLumaMipFlag variable or IntraSubPartitionModeFlag variable presented in this embodiment may have different variable names. That is to say, these variables can be expressed or referred to as different variables. For example, the IntraLumaMipFlag variable can be referred to as the MipFlag variable.
[0237] [Table 9]
[0238]
[0239]
[0240] [Table 10]
[0241]
[0242] The following drawings are provided to describe specific examples of the present disclosure. Since specific terms for the devices shown in the drawings are provided for illustration, or specific terms for signals / messages / fields, the technical features of the present disclosure are not limited to the specific terms used in the following drawings.
[0243] Figure 11 is a flowchart illustrating the operation of a video decoding device according to an embodiment of the present disclosure.
[0244] Figure 11 Each process disclosed in is based on some details described with reference to Figures 5 to 10 described. Therefore, the description of the specific details overlapping with those described with reference to Figure 3 and Figures 5 to 10 described will be omitted or will be presented schematically.
[0245] According to an embodiment, the decoding device 300 may receive information about an intra prediction mode, residual information, etc. from a bitstream, and may receive, for example, image information including intra prediction type information that includes an intra MIP syntax element (intra_mip_flag) for a first target block (S1110).
[0246] Specifically, the decoding device 300 may decode information on quantization transform coefficients for a current block from a bitstream and may derive quantization transform coefficients for a target block based on the information on quantization transform coefficients for the current block. The information on quantization transform coefficients for the target block may be included in a sequence parameter set (SPS) or a slice header and may include at least one of information on whether to apply RST, information on a reduction factor, information on a minimum transform size for applying RST, information on a maximum transform size for applying RST, an inverse RST size, and information on a transform index indicating any one of transform kernel matrices included in a transform set.
[0247] The decoding device may further receive information on an intra prediction mode for a current block and information on whether to apply ISP to the current block. The decoding device may receive and parse flag information indicating whether to apply ISP compilation or an ISP mode, thereby deriving whether the current block is divided into a predetermined number of sub-partition transform blocks. Here, the current block may be a compilation block. In addition, the decoding device may derive the size and number of sub-partitioned blocks of the division through flag information indicating a direction in which the current block is divided.
[0248] The decoding device 300 may decode an intra MIP syntax element for a first target block, thereby deriving a value of the intra MIP syntax element (S1120).
[0249] The decoding device may set a variable MIP flag for a preset specific region that is the same as a region of the first target block based on the value of the intra MIP syntax element (S1130).
[0250] As described in reference table 6, the variable MIP flag (IntraLumaMipFlag[x][y]) may be set in the semantics of the intra MIP syntax element and may be utilized in subsequent processes for decoding an image.
[0251] The specific region may be set to be the same as a region in which samples are located in the first target block (x = x0..x0 + cbWidth – 1 and y = y0..y0 + cbHeight – 1). Here, cbWidth and cbHeight indicate the width and height of the first target block.
[0252] The variable MIP flag may be set for the specific region based on the tree type of the first target block not being a dual-tree chrominance. That is, the variable MIP flag may be set to the received value of the intra MIP syntax element for the specific region only when the tree type of the first target block is a single-tree or dual-tree luminance rather than a dual-tree chrominance.
[0253] The intra_mip_flag syntax element for intra MIP is signaled only when the tree type of the corresponding coding block is single-tree or dual-tree luminance, and not signaled when the tree type is dual-tree chrominance and is thus inferred as 0. Therefore, when the tree type of the first target block is dual-tree chrominance, the value of intra_mip_flag is 0. In intra prediction or transformation, when the intra prediction mode for a chrominance block is derived, the intra prediction mode for the luminance block is adopted, and in this case, the newly set variable MIP flag can be used. When the condition for setting the variable MIP flag does not have the condition of setting the variable MIP flag only in the case of "single-tree and dual-tree luminance", the variable MIP flag can also be set in dual-tree chrominance. In this case, since the value of intra_mip_flag is 0, the variable MIP flag does not include information about the luminance component, and thus luminance information cannot be used when deriving the intra prediction mode for the chrominance block. To prevent this problem from occurring, the variable MIP flag can be set only when the corresponding coding block (i.e., the first target block) is not dual-tree chrominance.
[0254] The decoding device can derive the intra prediction mode for the second target block based on the variable MIP flag of the first target block (S1140).
[0255] Figure 12a and Figure 12b shows that the variable MIP flag for the first target block is used to derive the intra prediction mode for the second target block.
[0256] As Figure 12a shown, the first target block can be the left neighboring block of the second target block, and the specific region can include the sample position of (xCb - 1, yCb + cbHeight - 1). In this case, (xCb, yCb) is the position of the upper-left sample of the second target block, and cbHeight indicates the height of the second target block.
[0257] That is, the candidate intra prediction mode for the second target block can be derived based on the variable MIP flag of the sample position (i.e., the specific region) of (xCb - 1, yCb + cbHeight - 1) included in the first target block, and the intra prediction mode for the second target block can be derived based on this candidate intra prediction mode.
[0258] In another example, as Figure 12b shown, the first target block can be the upper neighboring block of the second target block, and the specific region can include the sample position of (xCb + cbWidth - 1, yCb - 1). In this case, (xCb, yCb) is the position of the upper-left sample of the second target block, and cbWidth indicates the width of the second target block.
[0259] That is, a candidate intra prediction mode for a second target block can be derived based on a variable MIP flag of a sample position (i.e., a specific region) of (xCb+cbWidth-1, yCb-1) included in a first target block, and an intra prediction mode for the second target block can be derived based on the candidate intra prediction mode.
[0260] In yet another example, the second target block may include a chroma block, and the first target block may be a luma block related to the chroma block. As described above, an intra prediction mode for the chroma block can be derived based on a variable MIP flag value for the luma block.
[0261] In this document, a picture / image may include an array of luma components, and in some cases may further include two arrays of chroma components (cb, cr). That is, one pixel of a picture / image may include a luma sample and chroma samples (cb, cr).
[0262] A color format may indicate a configuration format of luma components and chroma components (cb, cr), and may be referred to as a chroma format or a chroma array type. The color format (or chroma format) may be determined in advance or may be signaled adaptively. For example, the chroma format may be signaled based on at least one of chroma_format_idc and separate_colour_plane_flag as shown in Table 11.
[0263] [Table 11]
[0264] chroma_format_idc separate_colour_plane_flag Chroma format SubWidthC SubHeightC 0 0 Monochrome 1 1 1 0 4:2:0 2 2 2 0 4:2:2 2 1 3 0 4:4:4 1 1 3 1 4:4:4 1 1
[0265] Figure 13 Shows the configuration of samples for the chroma format according to Table 11.
[0266] A chroma format index of 1, i.e., 4:2:0 sampling of chroma array type 1, indicates that the height and width of the two chroma arrays are each half of the height and width of the luma array, while a chroma format index of 2, i.e., 4:2:2 sampling of chroma array type 2, indicates that the height of the two chroma arrays is equal to the height of the luma array and its width is half of the width of the luma array.
[0267] A chroma format index of 3, i.e., 4:4:4 sampling of chroma array type 3, indicates that the height and width of the chroma array are the same as the height and width of the luma array.
[0268] According to the example, when the tree type of the second target block is not a single tree or its chrominance array type is not 3, the specific region may include the sample position of (xCb + cbWidth / 2, yCb + cbHeight / 2). Here, (xCb, yCb) may indicate the upper left position of the chrominance block in the luma sample unit, cbWidth may indicate the width of the corresponding luma block corresponding to the chrominance block, and cbHeight may indicate the height of the corresponding luma block.
[0269] Figures 14a to 14c The variable MIP flag of the corresponding luma block shown as the first target block is used to derive the intra prediction mode for the chrominance block shown as the second target block. Figures 14a to 14c According to the tree type and color format of the luma block and the chrominance block, the sample position of (xCb + cbWidth / 2, yCb + cbHeight / 2) included in the specific region is shown.
[0270] Figure 14a The luma block and the chrominance block with a single tree type and a color format of 4:2:0 are shown. That is, Figure 14a The case where the condition that the chrominance array type of the second target block is not 3 is satisfied among the conditions that the tree type of the second target block is not a single tree or the chrominance array type of the second target block is not 3 is shown.
[0271] The first sample position (I) of the chrominance block indicates the upper left position of the chrominance block, and the second sample position (II(xCb, yCb)) of the luma block indicates the upper left position of the chrominance block in the luma sample unit. cbWidth indicates the width of the corresponding luma block corresponding to the chrominance block, and cbHeight indicates the height of the corresponding luma block. Since the color format is 4:2:0, the width and height (cbWidth and cbHeight) of the corresponding luma block corresponding to the chrominance block are twice the width and height (cbWidth / 2 and cbHeight / 2) of the chrominance block.
[0272] The variable MIP flag value of the third sample position (III) indicated by (xCb + cbWidth / 2, yCb + cbHeight / 2) in the luma block can be used to derive the intra prediction mode for the chrominance block. That is, the candidate intra prediction mode for the chrominance block shown as the second target block can be derived based on the variable MIP flag of the sample position of (xCb + cbWidth / 2, yCb + cbHeight / 2) included in the luma block shown as the first target block, and the intra prediction mode for the chrominance block can be derived based on the candidate intra prediction mode.
[0273] Figure 14bShows a luminance block and a chrominance block having a dual - tree type and a color format of 4:4:4. That is, Figure 14b Shows a case where the condition that the tree type of the second target block is not a single - tree is satisfied among the conditions that the tree type of the second target block is not a single - tree or the chrominance array type of the second target block is not 3.
[0274] As shown, the chrominance block for the luminance block indicated by the dashed line is indicated by the dashed line, and since the color format is 4:4:4, the widths and heights of the luminance array and the chrominance array are the same. Although the luminance block is not divided, the chrominance block is divided and compiled.
[0275] When the lower - right block of the divided chrominance block is the second target block to be predicted, the first sample position (I) indicates the upper - left position of the second target block as shown, and the second sample position (II(xCb,yCb)) indicates the upper - left position of the chrominance block in the luminance sample unit.
[0276] In addition, since the color format is 4:4:4, the third sample position (III) indicated by (xCb + cbWidth / 2,yCb + cbHeight / 2) in the luminance block is located in the corresponding luminance block, which is at the lower - right of the luminance block.
[0277] The decoding device can derive an intra - prediction mode for the chrominance block based on the variable MIP flag value of the third sample position (III). That is, a candidate intra - prediction mode for the chrominance block as the second target block can be derived based on the variable MIP flag of the sample position of (xCb + cbWidth / 2,yCb + cbHeight / 2) included in the luminance block as the first target block, and the intra - prediction mode for the chrominance block can be derived based on the candidate intra - prediction mode.
[0278] Figure 14c Shows a luminance block and a chrominance block having a dual - tree type and a color format of 4:2:0. That is, Figure 14c Shows a case where both the condition that the tree type of the second target block is not a single - tree and the condition that the chrominance array type of the second target block is not 3 are satisfied. In other words, Figure 14c Shows a case where the tree type of the second target block is not a single - tree and its chrominance array type is not 3.
[0279] The first sample position (I) of the chrominance block indicates the upper - left position of the chrominance block, while the second sample position (II(xCb,yCb)) of the luminance block indicates the upper - left position of the chrominance block in the luminance sample unit. cbWidth indicates the width of the corresponding luminance block corresponding to the chrominance block, and cbHeight indicates the height of the corresponding luminance block.
[0280] Therefore, the variable MIP flag value at the third sample position (III) indicated by (xCb + cbWidth / 2, yCb + cbHeight / 2) in the luminance block can be used to derive the intra prediction mode for the chrominance block. That is, the candidate intra prediction mode for the chrominance block as the second target block can be derived based on the variable MIP flag at the sample position of (xCb + cbWidth / 2, yCb + cbHeight / 2) included in the luminance block as the first target block, and the intra prediction mode for the chrominance block can be derived based on the candidate intra prediction mode.
[0281] The decoding device can derive the prediction samples for the second target block based on the intra prediction mode for the second target block (S1150), and can generate the reconstructed block based on the prediction samples (S1160).
[0282] The decoding device can generate the reconstructed block based on the received residual information and the prediction samples. The decoding device can derive the transform coefficients through the transform process based on the residual information, and when the intra prediction mode for the second target block is required in the transform process, the variable MIP flag value of the first target block can be used.
[0283] According to another example, when the intra prediction type information includes the flag syntax element (intra_subpartitions_mode_flag) for the intra sub - partition (ISP) mode of the first target block, the decoding device can use the variable value set based on the value of intra_subpartitions_mode_flag to derive the intra prediction mode for the second target block.
[0284] That is, the decoding device can derive the value of the flag syntax element (intra_subpartitions_mode_flag), can set the variable ISP flag (IntraSubPartitionsModeFlag) for a preset specific region that is the same as the region of the first target block based on the value of the flag syntax element, and can derive the intra prediction mode for the second target block based on the value of the variable ISP flag.
[0285] The variable ISP flag can also be set for the specific region based on the tree type of the first target block not being a dual - tree chrominance. That is, the variable ISP flag can be set to the received value of the flag syntax element for the specific region only when the tree type of the first target block is a single - tree or dual - tree luminance rather than a dual - tree chrominance.
[0286] The following drawings are provided to describe specific examples of the present disclosure. Since specific terms of the devices illustrated in the drawings or specific terms of signals / messages / fields are provided for illustration, the technical features of the present disclosure are not limited to the specific terms used in the following drawings.
[0287] Figure 15 is a flowchart illustrating the operation of a video encoding device according to an embodiment of the present disclosure.
[0288] In Figure 15 each process disclosed is based on some details described with reference to Figures 5 to 10 Therefore, descriptions of specific details overlapping those described with reference to Figure 2 and Figures 5 to 10 will be omitted or will be presented schematically.
[0289] When applying the intra MIP mode to the first target block, the encoding device 200 according to an embodiment can derive prediction samples for the first target block and can derive the value of the intra MIP flag for the first target block (S1510).
[0290] When ISP is applied to the current block, the encoding device can perform prediction through each sub-partition transform block.
[0291] The encoding device can determine whether to compile ISP or apply the ISP mode to the current block, i.e., the compile block, and can determine the direction in which the current block is divided, and can derive the size and number of the divided sub-blocks according to the determination result.
[0292] The same intra prediction mode can be applied to the sub-partition transform blocks into which the current block is divided, and the encoding device can derive prediction samples for each sub-partition transform block. That is, the encoding device sequentially performs intra prediction according to the division form of the sub-partition transform block, for example, horizontally or vertically, or from left to right or from top to bottom. For the leftmost or topmost sub-blocks, as in the traditional intra prediction method, the reconstructed pixels of the already compiled compile block are referred to. Further, for each side of the subsequent internal sub-partition transform block not adjacent to the previous sub-partition transform block, in order to derive the reference pixels adjacent to that side, the reconstructed pixels of the already compiled adjacent compile block are referred to as in the traditional intra prediction method.
[0293] The encoding device can set the variable MIP flag for a preset specific area identical to the area of the first target block based on the value of the intra MIP flag (S1520).
[0294] As described with reference to Table 6, the variable MIP flag (IntraLumaMipFlag[x][y]) can be set in the semantics of the intra MIP syntax element and can be utilized in subsequent processes for decoding an image.
[0295] A specific area can be set to be the same as the area where the sample is in the first target block (x = x0..x0 + cbWidth – 1 and y = y0..y0 + cbHeight – 1). Here, cbWidth and cbHeight indicate the width and height of the first target block.
[0296] The variable MIP flag for the specific area can be set based on that the tree type of the first target block is not a dual-tree chrominance. That is, only when the tree type of the first target block is a single-tree or dual-tree luma rather than a dual-tree chrominance, can the variable MIP flag be set for the specific area to the value of the received intra-MIP syntax element.
[0297] The intra-MIP syntax element intra_mip_flag is signaled only when the tree type of the corresponding coded block is a single-tree or dual-tree luma, and is not signaled when the tree type is a dual-tree chrominance and thus is inferred to be 0. Therefore, when the tree type of the first target block is a dual-tree chrominance, the value of intra_mip_flag is 0. In intra prediction or transformation, when the intra prediction mode for the chrominance block is derived, the intra prediction mode for the luma block is adopted, and in this case, the newly set variable MIP flag can be used. When the condition for setting the variable MIP flag does not have the condition of setting the variable MIP flag only in the case of "single-tree and dual-tree luma", the variable MIP flag can also be set in the dual-tree chrominance. In this case, since the value of intra_mip_flag is 0, the variable MIP flag does not include information about the luma component, and thus the luma information cannot be used when deriving the intra prediction mode for the chrominance block. To prevent this problem from occurring, the variable MIP flag can be set only when the corresponding coded block (i.e., the first target block) is not a dual-tree chrominance.
[0298] The encoding device can derive the intra prediction mode for the second target block based on the variable MIP flag of the first target block (S1530).
[0299] Reference Figures 12a to 14c The described details can be applied to the intra prediction mode derivation process performed by the encoding device.
[0300] The encoding device can derive the prediction samples for the second target block based on the derived intra prediction mode for the second target block (S1540), and can derive the residual samples for the second target block based on these prediction samples (S1550).
[0301] In addition, as described above, when the intra prediction type information for the first target block is in the intra sub-partition (ISP) mode, the variable ISP flag value of the first target block can be set based on a flag value indicating whether to perform the intra sub-partition (ISP) mode. The intra prediction mode for the second target block can be derived based on the variable ISP flag value.
[0302] The encoding device can encode and output the transform coefficient information generated based on the intra MIP flag and the residual samples (S1560).
[0303] The encoding device can derive the quantized transform coefficients by performing quantization on the modified transform coefficients for the current block, and can generate and output image information including the intra MIP flag.
[0304] The encoding device can generate residual information including information about the quantized transform coefficients. The residual information can include the aforementioned transform-related information / syntax elements. The encoding device can encode the image / video information including the residual information and can output the encoded image / video information in the form of a bitstream.
[0305] Specifically, the encoding device 200 can generate information about the quantized transform coefficients and can encode the information about the quantization of the generated transform coefficients.
[0306] In the present disclosure, at least one of quantization / dequantization and / or transformation / inverse transformation can be omitted. When quantization / dequantization is omitted, the quantized transform coefficients can be referred to as transform coefficients. When transformation / inverse transformation is omitted, the transform coefficients can be referred to as coefficients or residual coefficients, or for the sake of consistency in expression, can still be referred to as transform coefficients.
[0307] In addition, in the present disclosure, the quantized transform coefficients and the transform coefficients can be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, the residual information can include information about the transform coefficients, and the information about the transform coefficients can be signaled through the residual compilation syntax. The transform coefficients can be derived based on the residual information (or the information about the transform coefficients), and the scaled transform coefficients can be derived through the inverse transform (scaling) of the transform coefficients. The residual samples can be derived based on the inverse transform (transformation) of the scaled transform coefficients. These details can also be applied / expressed in other parts of the present disclosure.
[0308] In the above embodiments, a method is illustrated based on a flowchart by means of a series of steps or blocks, but the present disclosure is not limited to the order of the steps, and a certain step may be performed in an order or steps different from the above order or steps or simultaneously with another step. In addition, those of ordinary skill in the art can understand that the steps shown in the flowchart are not exclusive, and one or more steps of the flowchart may be incorporated or removed without affecting the scope of the present disclosure.
[0309] The above method according to the present disclosure can be implemented in the form of software, and the encoding device and / or decoding device according to the present disclosure may be included in a device for image processing such as a TV, a computer, a smart phone, a set-top box, a display device, etc.
[0310] When the embodiments in the present disclosure are specifically implemented by software, the above method can be specifically embodied as a module (procedure, function, etc.) for performing the above functions. The module can be stored in a memory and can be executed by a processor. The memory can be inside or outside the processor and can be connected to the processor in various well-known ways. The processor may include an application specific integrated circuit (ASIC), other chip sets, logic circuits, and / or data processing devices. The memory may include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in the present disclosure can be specifically implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing can be specifically implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip.
[0311] In addition, the decoding device and the encoding device applying the present disclosure may be included in a multimedia broadcast transceiver, a mobile communication terminal, a home theater video device, a digital cinema video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camera, a video on demand (VoD) service providing device, an over-the-top (OTT) video device, an Internet streaming service providing device, a three-dimensional (3D) video device, a video phone video device, and a medical video device, and can be used to process video signals or data signals. For example, an over-the-top (OTT) video device may include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smart phone, a tablet PC, a digital video recorder (DVR), etc.
[0312] In addition, the processing method of the present disclosure can be generated in the form of a program executable by a computer and stored in a computer-readable recording medium. Multimedia data having a data structure according to the present disclosure can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all kinds of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium may include, for example, Blu-ray Disc (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage device. In addition, the computer-readable recording medium includes a medium embodied in the form of a carrier wave (e.g., transmission via the Internet). Additionally, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network. Additionally, embodiments of the present disclosure can be embodied as a computer program product by program code, and the program code can be executed on a computer by embodiments of the present disclosure. The program code can be stored on a computer-readable carrier.
[0313] The claims disclosed herein can be combined in various ways. For example, the technical features of the method claims of the present disclosure can be combined to be implemented or executed in a device, and the technical features of the device claims can be combined to be implemented or executed in a method. In addition, the technical features of the method claims and the device claims can be combined to be implemented or executed in a device, and the technical features of the method claims and the device claims can be combined to be implemented or executed in a method.
Claims
1. An image decoding method performed by a decoding device, the method comprising: receiving image information including intra prediction type information from a bitstream, the intra prediction type information including an intra MIP syntax element for a first coding block, the intra MIP syntax element indicating whether matrix-based intra prediction is applied to the first coding block; decoding the intra MIP syntax element; setting a MIP flag variable for the first coding block based on the value of the decoded intra MIP syntax element, the MIP flag variable corresponding to a sample position in the first coding block respectively, wherein each of the MIP flag variables is set to be equal to the value of the intra MIP syntax element indicating whether matrix-based intra prediction is applied; deriving an intra prediction mode for a second coding block based on one of the MIP flag variables; deriving predicted samples for the second coding block based on the intra prediction mode for the second coding block; and generating a reconstructed block based on the predicted samples.
2. The image decoding method according to claim 1, wherein, the first coding block is a neighboring block of the second coding block, and wherein deriving the intra prediction mode for the second coding block includes: deriving a candidate intra prediction mode based on one of the MIP flag variables; and deriving the intra prediction mode for the second coding block based on the candidate intra prediction mode.
3. The image decoding method according to claim 2, wherein, in response to the first coding block being a left neighboring block adjacent to the second coding block, one of the MIP flag variables represents the MIP flag variable corresponding to the sample position of (xCb - 1, yCb + cbHeight - 1), and wherein (xCb, yCb) is the upper left sample position of the second coding block, and cbHeight indicates the height of the second coding block.
4. The image decoding method according to claim 2, wherein, in response to the first coding block being an upper neighboring block adjacent to the second coding block, one of the MIP flag variables represents the MIP flag variable corresponding to the sample position of (xCb + cbWidth - 1, yCb - 1), and wherein (xCb, yCb) is the upper left sample position of the second coding block, and cbWidth indicates the width of the second coding block.
5. The image decoding method according to claim 1, wherein, the first coding block represents a luminance block and the second coding block represents a chrominance block corresponding to the first coding block, and wherein deriving the intra prediction mode for the second coding block includes: deriving a corresponding luminance intra prediction mode for the first coding block based on one of the MIP flag variables; and deriving the intra prediction mode for the second coding block based on the corresponding luminance intra prediction mode.
6. The image decoding method according to claim 5, wherein, In response to the tree type of the second compiled block not being a single tree or its chrominance array type not being 3, deriving the corresponding intra prediction mode of the luminance frame based on the MIP flag variable among the MIP flag variables corresponding to the sample position of (xCb + cbWidth / 2, yCb + cbHeight / 2), and where (xCb, yCb) indicates the top-left sample position of the chrominance block, cbWidth indicates the width of the luminance block, and cbHeight indicates the height of the luminance block.
7. The image decoding method according to claim 1, wherein, Based on the tree type of the first compiled block not being a dual-tree chrominance, setting the MIP flag variable for the first compiled block.
8. An image encoding method executed by an encoding device, the method comprising: Determining the value of the intra MIP syntax element for the first compiled block; Based on the value of the intra MIP flag, setting the MIP flag variables for the first compiled block, the MIP flag variables corresponding to the sample positions in the first compiled block respectively, wherein each of the MIP flag variables is set to be equal to the value of the intra MIP syntax element indicating whether matrix-based intra prediction is applied; Based on one of the MIP flag variables, determining the intra prediction mode for the second compiled block; Based on the predicted samples for the second compiled block, deriving the residual samples for the second compiled block, the predicted samples being obtained based on the intra prediction mode for the second compiled block; and Encoding the transform coefficient information generated based on the residual samples.
9. The image encoding method according to claim 8, wherein, The first compiled block is a neighboring block of the second compiled block, and where determining the intra prediction mode for the second compiled block includes: Deriving candidate intra prediction modes based on one of the MIP flag variables; and Based on the candidate intra prediction modes, determining the intra prediction mode for the second compiled block.
10. The image encoding method according to claim 9, wherein, In response to the first compiled block being the left neighboring block adjacent to the second compiled block, one of the MIP flag variables represents the MIP flag variable corresponding to the sample position of (xCb - 1, yCb + cbHeight - 1), where in response to the first compiled block being the upper neighboring block adjacent to the second compiled block, one of the MIP flag variables represents the MIP flag variable corresponding to the sample position of (xCb + cbWidth - 1, yCb - 1), and where (xCb, yCb) is the top-left sample position of the second compiled block, and cbHeight and cbWidth respectively indicate the height and width of the second compiled block.
11. The image encoding method according to claim 8, wherein, The first compiled block represents a luminance block, and the second compiled block represents a chrominance block corresponding to the first compiled block, and where determining the intra prediction mode for the second compiled block includes: Derive a corresponding intra - luminance prediction mode for the first compiled block based on one of the MIP flag variables; and Determine the intra - prediction mode for the second compiled block based on the corresponding intra - luminance prediction mode.
12. The image coding method according to claim 11, wherein, In response to the tree type of the second compiled block not being a single tree or the chrominance array type of the second compiled block not being 3, derive the corresponding intra - luminance prediction mode based on the MIP flag variable among the MIP flag variables corresponding to the sample position of (xCb + cbWidth / 2, yCb + cbHeight / 2), and where (xCb, yCb) indicates the upper - left sample position of the chrominance block, cbWidth indicates the width of the luminance block, and cbHeight indicates the height of the luminance block.
13. The image coding method according to claim 8, wherein, Set the MIP flag variable for the first compiled block based on the tree type of the first compiled block not being a dual - tree chrominance.
14. A non - transitory computer - readable digital storage medium for storing a computer program, the computer program being executed by a processor to generate a bitstream according to the image coding method of claim 8.
15. A method for transmitting data of image information, the method comprises: Determine the value of the intra - frame MIP syntax element for the first compiled block; Based on the value of the intra - frame MIP flag, set the MIP flag variables for the first compiled block, the MIP flag variables corresponding to the sample positions in the first compiled block respectively, wherein each of the MIP flag variables is set to be equal to the value of the intra - frame MIP syntax element indicating whether matrix - based intra - prediction is applied; Determine the intra - prediction mode for the second compiled block based on one of the MIP flag variables; Derive the residual samples for the second compiled block based on the prediction samples for the second compiled block, the prediction samples being obtained based on the intra - prediction mode for the second compiled block; Encode the transform coefficient information generated based on the residual samples to generate a bitstream; and Transmit the data including the bitstream.
Citation Information
Patent Citations
CONSTRAINED BLOCK-LEVEL OPTIMIZATION AND SIGNALING FOR VIDEO decoding TOOLS
CN108781289A