Image encoding / decoding method and apparatus, and recording medium for storing bitstream
By determining the initial intra-frame prediction mode and optimizing adjacent candidate modes in image coding, the problem of low intra-frame prediction accuracy is solved, coding efficiency is improved, and transmission and storage costs are reduced.
Patent Information
- Application Number
- CN202480025049.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-08
- Filing Date
- 2024-06-05
- Publication Date
- 2025-12-12
AI Technical Summary
Existing image coding techniques suffer from low intra-frame prediction accuracy, which limits coding efficiency. Furthermore, the increased data volume of high-resolution, high-quality images leads to higher transmission and storage costs.
By determining the initial intra-prediction mode of the current block and identifying adjacent candidate modes based on this mode, the optimal intra-prediction mode is finally determined. Combined with the intra-prediction mode refinement method, the prediction accuracy is improved.
This improves the prediction accuracy of intra-frame prediction mode, thereby enhancing overall coding efficiency and reducing transmission and storage costs.
Smart Images

Figure CN121128165A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to an image encoding / decoding method, an apparatus, and a recording medium for storing a bitstream. More particularly, the present disclosure relates to an image encoding / decoding method, an apparatus, and a recording medium for storing a bitstream using an intra prediction method. BACKGROUND
[0002] In recent years, there has been an increasing demand for high-resolution, high-quality images, such as ultra-high definition (UHD) images, in various application fields. As the resolution and quality of image data increase, the amount of data also relatively increases compared to existing image data. Therefore, when transmitting image data using existing wired and wireless broadband lines or the like, or storing image data using existing storage media, the transmission and storage costs increase. In order to solve these problems caused by the increase in the resolution and quality of image data, efficient image encoding / decoding technology needs to be developed to implement higher resolution and higher quality images.
[0003] In video encoding and decoding, intra prediction is a technique of predicting a current block using reconstructed reference pixels of a current image. At this time, according to the intra prediction, a prediction block is generated from neighboring reference pixels around the current block based on a predetermined non-directional mode or a directional mode. Compared to inter prediction, the prediction accuracy of the intra prediction can be low, thereby causing a limitation in coding efficiency. Therefore, various methods for improving the prediction accuracy of the intra prediction are being discussed. SUMMARY
[0004] TECHNICAL PROBLEM
[0005] The present disclosure aims to provide an image encoding / decoding method and apparatus having improved encoding / decoding efficiency.
[0006] In addition, the present disclosure aims to provide a recording medium for storing a bitstream generated by a method or apparatus for decoding an image provided by the present disclosure.
[0007] TECHNICAL SOLUTION
[0008] The method of decoding an image according to an embodiment of the present disclosure can include determining an initial intra prediction mode of a current block, determining one or more neighboring candidate modes adjacent to the initial intra prediction mode based on the initial intra prediction mode of the current block, determining a final intra prediction mode of the current block among candidate intra prediction modes including the one or more neighboring candidate modes, and predicting the current block based on the final intra prediction mode of the current block.
[0009] According to an embodiment, the method further includes: determining whether to perform intra prediction refinement on the current block; when the intra prediction refinement is performed on the current block, determining one or more neighboring candidate modes adjacent to the initial intra prediction mode based on the initial intra prediction mode of the current block; determining a final intra prediction mode of the current block among the candidate intra prediction modes including the one or more neighboring candidate modes; and predicting the current block based on the final intra prediction mode of the current block; when the intra prediction refinement is not performed on the current block, predicting the current block based on the initial intra prediction mode of the current block.
[0010] According to an embodiment, when the initial intra prediction mode is a non-directional intra prediction mode, it can be determined not to perform the intra prediction refinement on the current block.
[0011] According to an embodiment, when an index of the initial intra prediction mode is not included in a predetermined range, it can be determined not to perform the intra prediction refinement on the current block.
[0012] According to an embodiment, whether to perform the intra prediction refinement on the current block can be determined according to at least one of a width, a height, an area, and an aspect ratio of the current block.
[0013] According to an embodiment, when the intra prediction refinement is applied to a reference block of the current block, the initial intra prediction mode of the current block can be determined based on an initial intra prediction mode of the reference block.
[0014] According to an embodiment, the one or more neighboring candidate modes can include: a first neighboring candidate mode having a prediction direction between a prediction direction of a first neighboring intra prediction mode of the current block and a prediction direction of the initial intra prediction mode; and a second neighboring candidate mode having a prediction direction between a prediction direction of a second neighboring intra prediction mode of the current block and the prediction direction of the initial intra prediction mode; the first neighboring intra prediction mode can have an index that is one predetermined value smaller than an index of the initial intra prediction mode; and the second neighboring intra prediction mode can have an index that is one predetermined value larger than the index of the initial intra prediction mode.
[0015] According to an embodiment, the first neighboring candidate mode can be determined by equally dividing a direction difference between the prediction direction of the first neighboring intra prediction mode of the current block and the prediction direction of the initial intra prediction mode according to a number of the first neighboring candidate modes; and the second neighboring candidate mode can be determined by equally dividing a direction difference between the prediction direction of the second neighboring intra prediction mode of the current block and the prediction direction of the initial intra prediction mode according to a number of the second neighboring candidate modes.
[0016] According to one implementation, when the initial intra-prediction mode is the intra-prediction mode with the smallest index among the directional modes within a predetermined range, the one or more adjacent candidate modes may include only the second adjacent intra-prediction mode; when the initial intra-prediction mode is the intra-prediction mode with the largest index among the directional modes within a predetermined range, the one or more adjacent candidate modes may include only the first adjacent intra-prediction mode.
[0017] According to one implementation, among the candidate intra-prediction modes of the current block, the mode with the least distortion is determined as the final intra-prediction mode of the current block.
[0018] According to one implementation, the distortion of the candidate intra-frame prediction mode can be determined based on the difference between the reconstructed sample contained in the adjacent template of the current block and the prediction sample corresponding to the reconstructed sample.
[0019] According to one implementation, the prediction sample can be determined based on the template reference sample adjacent to the template and the prediction direction of the candidate intra-frame prediction mode.
[0020] According to one embodiment, the method for decoding an image further includes obtaining intra-prediction mode refinement index information from the bitstream, the intra-prediction mode refinement index information indicating the final intra-prediction mode of the current block among the candidate intra-prediction modes of the current block, and the final intra-prediction mode of the current block can be determined based on the intra-prediction mode refinement index information.
[0021] According to one implementation, the intra-prediction mode refinement index information may include: intra-prediction mode refinement sign information, which indicates whether the final intra-prediction mode is positive (+) or negative (-) relative to the initial intra-prediction mode; and intra-prediction mode refinement index difference information, which indicates the index difference between the final intra-prediction mode and the initial intra-prediction mode.
[0022] According to one implementation, in the intra-prediction mode refinement index difference information, the codeword length allocated to the candidate intra-prediction mode can be determined based on the index difference between the initial intra-prediction mode and the candidate intra-prediction mode.
[0023] A method for encoding an image according to an embodiment of the present disclosure may include: determining an initial intra-prediction mode for a current block; determining one or more neighboring candidate modes adjacent to the initial intra-prediction mode based on the initial intra-prediction mode of the current block; determining a final intra-prediction mode for the current block among the candidate intra-prediction modes including one or more neighboring candidate modes; and predicting the current block based on the final intra-prediction mode of the current block.
[0024] According to an embodiment of the present disclosure, a non-transitory computer-readable recording medium can store a bitstream generated by the method of encoding the image.
[0025] According to an embodiment of the present disclosure, a transmission method can transmit a bit stream generated by the method of encoding the image.
[0026] The features of the present invention briefly summarized above are provided as examples only to explain the detailed description and do not constitute a limitation on the scope of the present invention.
[0027] Beneficial effects
[0028] This invention proposes an embodiment of an intra-prediction method based on decoder-side intra-prediction mode refinement.
[0029] Furthermore, this invention also proposes an embodiment of a method for encoder-side intra-frame prediction mode refinement.
[0030] According to the above embodiments, the prediction accuracy of intra-frame prediction mode can be improved, thereby improving the overall coding efficiency. Attached Figure Description
[0031] Figure 1 This is a block diagram illustrating the configuration of an encoding apparatus according to an embodiment of the present disclosure.
[0032] Figure 2 This is a block diagram illustrating the configuration of a decoding apparatus according to an embodiment of the present disclosure.
[0033] Figure 3 This is a schematic diagram illustrating a video coding system to which this disclosure applies.
[0034] Figure 4 This demonstrates a template-based intra-mode derivation (TIMD) method.
[0035] Figure 5 and Figure 6 An implementation of a method for determining neighboring candidate modes of an initial intra-frame prediction mode is shown.
[0036] Figure 7 and Figure 8 An implementation of a method for determining adjacent candidate modes of an initial intra-prediction mode is shown, wherein the initial intra-prediction mode is the intra-prediction mode with the smallest index among directional modes within a predetermined range.
[0037] Figure 9 and Figure 10 This paper illustrates an implementation of a method for determining adjacent candidate modes of the initial intra-prediction mode when the initial intra-prediction mode is the intra-prediction mode with the largest index among directional modes within a predetermined range.
[0038] Figure 11This illustrates an implementation of an intra-prediction method refined by applying intra-prediction modes.
[0039] Figure 12 An exemplary content streaming system to which embodiments of this disclosure may be applied is shown. Detailed Implementation
[0040] Best mode
[0041] A method for decoding an image according to embodiments of the present disclosure may include: determining an initial intra-prediction mode for a current block; determining one or more neighboring candidate modes adjacent to the initial intra-prediction mode based on the initial intra-prediction mode of the current block; determining a final intra-prediction mode for the current block among the candidate intra-prediction modes including one or more neighboring candidate modes; and predicting the current block based on the final intra-prediction mode of the current block.
[0042] Invention Model
[0043] This disclosure can have various modifications and implementations, which are illustrated in the accompanying drawings and described in detail in the specification. However, this is not to imply that this disclosure is limited to the specific implementations, but should be understood to include all modifications, equivalents, or substitutions within the spirit and technical scope of this disclosure. Similar reference numerals in the drawings indicate the same or similar functions in various aspects. For clarity of description, the shapes and dimensions of elements in the drawings may be provided by way of example. The detailed description of exemplary embodiments described below refers to the accompanying drawings, which illustrate specific embodiments by way of example. These embodiments are described in sufficient detail to enable those skilled in the art to practice them. It should be understood that the various embodiments differ from one another, but are not necessarily mutually exclusive. For example, specific shapes, structures, and features described herein may be implemented in other embodiments without departing from the spirit and scope of this disclosure with respect to one embodiment. It should also be understood that the position or arrangement of various components in each disclosed embodiment may be changed without departing from the spirit and scope of the embodiments. Therefore, the detailed description described below is not intended to be limiting, and the scope of the exemplary embodiments is defined only by the full scope of the appended claims and their equivalents (if properly described).
[0044] In this disclosure, the terms "first," "second," etc., may be used to describe various components, but components should not be limited by these terms. These terms are used only to distinguish one component from another. For example, without departing from the scope of this disclosure, a first component may be referred to as a second component, and similarly, a second component may be referred to as a first component. The term is and / or includes a combination of multiple related descriptive items or any item among multiple related descriptive items.
[0045] The components shown in the embodiments of this disclosure are depicted independently to indicate different functional characteristics, and this does not mean that each component constitutes a separate hardware or software configuration unit. That is, for ease of explanation, each component is listed and treated as a separate component, and at least two components can be combined to form a single component, or a component can be broken down into multiple components to perform functions. Furthermore, implementations of component integration and implementations of each component being broken down are also included within the scope of this disclosure, as long as they do not depart from the essence of this disclosure.
[0046] The terminology used in this disclosure is for descriptive purposes only and is not intended to limit the scope of the disclosure. Singular expressions include plural expressions unless the context clearly indicates otherwise. Furthermore, some components of this disclosure are not essential for performing the basic functions of this disclosure, but may be optional components used only to improve performance. This disclosure can be implemented by excluding components used only to improve performance and including only the essential components necessary to implement the essence of this disclosure, and structures that include only essential components and exclude optional components used only to improve performance are also included within the scope of this disclosure.
[0047] In one implementation, the term "at least one" may refer to a numerical value greater than or equal to 1, such as 1, 2, 3, and 4. In one implementation, the term "a plurality of" may refer to a numerical value greater than or equal to 2, such as 2, 3, and 4.
[0048] The embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. In describing the embodiments of this specification, if it is determined that a detailed description of a related known configuration or function might obscure the subject matter of this specification, such a detailed description will be omitted, and the same reference numerals will be used for the same components in the drawings, and repeated descriptions of the same components will be omitted.
[0049] Terminology Explanation
[0050] In the following text, "image" can refer to a single frame that makes up a video, or it can refer to the video itself. For example, "encoding and / or decoding of an image" can refer to "encoding and / or decoding of a video," or it can refer to "encoding and / or decoding of one of the images that make up a video."
[0051] In the following text, "moving image" and "video" have the same meaning and can be used interchangeably. Furthermore, the target image can be an encoded target image that serves as the encoding target and / or a decoded target image that serves as the decoding target. Additionally, the target image can be an input image input to the encoding device or an input image input to the decoding device. Here, the meaning of target image can be the same as the current image.
[0052] In the following text, "image," "picture," "frame," and "screen" have the same meaning and can be used interchangeably.
[0053] In the following text, "target block" can refer to an encoding target block that is the target of encoding and / or a decoding target block that is the target of decoding. Furthermore, a target block can be the current block, i.e., the target block for the current encoding and / or decoding. For example, "target block" and "current block" have the same meaning and can be used interchangeably.
[0054] In the following text, "block" and "unit" have the same meaning and can be used interchangeably. Furthermore, "unit" can refer to a block containing a luma component block and a corresponding chroma component block, thus distinguishing it from a block. For example, a coding tree unit (CTU) can consist of a luma component (Y) coding tree block (CTB) and two associated chroma component (Cb, Cr) coding tree blocks.
[0055] In the following text, "sample," "image element," and "pixel" have the same meaning and can be used interchangeably. Here, a sample can represent the basic unit that makes up a block.
[0056] In the following text, "inter-frame" and "inter-screen" have the same meaning and can be used interchangeably.
[0057] In the following text, "intra-frame" and "intra-screen" have the same meaning and can be used interchangeably.
[0058] Figure 1 This is a block diagram illustrating the configuration of an encoding apparatus according to an embodiment of the present disclosure.
[0059] The encoding device 100 can be an encoder, a video encoding device, or an image encoding device. The video can include one or more images. The encoding device 100 can encode one or more images sequentially.
[0060] refer to Figure 1 The encoding device 100 may include an image segmentation unit 110, an intra-frame prediction unit 120, a motion prediction unit 121, a motion compensation unit 122, a switch 115, a subtractor 113, a transform unit 130, a quantization unit 140, an entropy coding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 117, a filtering unit 180, and a reference image buffer 190.
[0061] Furthermore, the encoding device 100 can generate a bitstream containing information encoded by encoding the input image and output the generated bitstream. The generated bitstream can be stored in a computer-readable recording medium or streamed via a wired / wireless transmission medium.
[0062] Image segmentation unit 110 can segment the input image into various forms to improve the efficiency of video encoding / decoding. That is, the input video consists of multiple images, and individual images can be segmented and processed hierarchically to improve compression efficiency, achieve parallel processing, etc. For example, a single image can be segmented into one or more tiles or slices, and then further segmented into multiple CTUs (Coding Tree Units). Alternatively, a single image can first be segmented into multiple sub-images defined as groups of rectangular slices, and then each sub-image can be segmented into tiles / slices. In this case, the sub-images can be used to support partially independent encoding / decoding and transmission of images. Since multiple sub-images can be reconstructed individually, it has the advantage of ease of editing in applications that configure multi-channel input as a single image. Furthermore, tiles can be horizontally segmented into bricks. In this case, bricks can be used as the basic unit for parallel processing within the image. Additionally, a CTU can be recursively segmented into a quadtree (QT), and the segmented terminal node can be defined as a coding unit (CU). A CU can be segmented into prediction units (PUs) and transform units (TUs) to perform prediction and segmentation. Simultaneously, the CU can be used as a prediction unit and / or the transform unit itself. Here, for flexible segmentation, each CTU can be recursively segmented into a multi-type tree (MTT) and a quadtree (QT). Segmenting the CTU into a multi-type tree can start from the terminal node of the QT. The MTT can consist of a binary tree (BT) and a ternary tree (TT). For example, the MTT structure can be divided into a vertical binary segmentation mode (SPLIT_BT_VER), a horizontal binary segmentation mode (SPLIT_BT_HOR), a vertical ternary segmentation mode (SPLIT_TT_VER), and a horizontal ternary segmentation mode (SPLIT_TT_HOR). Furthermore, during segmentation, the minimum block size (MinQTSize) of the quadtree for the luma block can be set to 16x16, the maximum block size (MaxBtSize) of the binary tree can be set to 128x128, and the maximum block size (MaxTtSize) of the ternary tree can be set to 64x64. Furthermore, the minimum block size (MinBtSize) of the binary tree and the minimum block size (MinTtSize) of the ternary tree can be specified as 4x4, and the maximum depth (MaxMttDepth) of the multi-type tree can be specified as 4. Additionally, to improve the coding efficiency of I-slices, a dual-tree approach can be adopted, using a CTU partitioning structure for both luma and chroma components. On the other hand, in P-slices and B-slices, the luma and chroma CTBs (Coding Tree Blocks) within the CTU can be divided into single trees sharing a coding tree structure.
[0063] The encoding device 100 can encode the input image in intra-frame mode and / or inter-frame mode. Alternatively, the encoding device 100 can encode the input image in a third mode other than intra-frame mode and inter-frame mode (e.g., IBC mode, palette mode, etc.). However, if the third mode has similar functional characteristics to the intra-frame mode or inter-frame mode, it can be classified as an intra-frame mode or inter-frame mode for ease of explanation. In this disclosure, the third mode will only be classified and described separately when it is necessary to specifically describe it.
[0064] When using intra-frame mode as the prediction mode, switch 115 can switch to intra-frame mode; when using inter-frame mode as the prediction mode, switch 115 can switch to inter-frame mode. Here, intra-frame mode can refer to intra-frame prediction mode, and inter-frame mode can refer to inter-frame prediction mode. The encoding device 100 can generate prediction blocks for input blocks of the input image. Furthermore, after generating prediction blocks, the encoding device 100 can encode residual blocks using the residuals between the input blocks and the prediction blocks. The input image can be called the current image, i.e., the current encoding target. The input block can be called the current block, i.e., the current encoding target or encoding target block.
[0065] When the prediction mode is intra-frame mode, the intra-frame prediction unit 120 can use samples from encoded / decoded blocks surrounding the current block as reference samples. The intra-frame prediction unit 120 can use the reference samples to perform spatial prediction on the current block, or generate prediction samples for the input block through spatial prediction. Here, intra-frame prediction can refer to in-screen prediction.
[0066] Intra-frame prediction methods can apply non-directional prediction modes (such as DC mode and planar mode) as well as directional prediction modes (such as 65 directions). Here, intra-frame prediction methods can be represented as intra-frame prediction modes or in-screen prediction modes.
[0067] When the prediction mode is inter-frame mode, the motion prediction unit 121 can retrieve the region in the reference image that best matches the input block during motion prediction and derive the motion vector using the retrieved region. In this case, the search region can be used as the region. The reference image can be stored in the reference image buffer 190. Here, it can be stored in the reference image buffer 190 when encoding / decoding the reference image.
[0068] The motion compensation unit 122 can generate a prediction block for the current block by using motion vectors for motion compensation. Here, inter-frame prediction can refer to inter-screen prediction or motion compensation.
[0069] When the value of the motion vector is not an integer, the motion prediction unit 121 and the motion compensation unit 122 can generate prediction blocks by applying an interpolation filter to a portion of the reference image. To perform inter-frame prediction or motion compensation, the motion prediction and motion compensation modes of the prediction units included in the coding unit can be determined based on the coding unit, such as skip mode, merge mode, advanced motion vector prediction (AMVP) mode, or intra-block copy (IBC) mode, and inter-frame prediction or motion compensation can be performed according to each mode.
[0070] Furthermore, based on the aforementioned inter-frame prediction methods, other modes can be applied, including AFFINE mode based on sub-PU prediction, SbTMVP (sub-block temporal motion vector prediction), MMVD (merged with MVD) mode based on PU prediction, and Geometric Partitioning (GPM). Additionally, to improve the performance of each mode, other modes can be applied, such as HMVP (history-based MVP), PAMVP (pairwise average MVP), CIIP (intra-frame / inter-frame combined prediction), AMVR (adaptive motion vector resolution), BDOF (bidirectional optical flow), BCW (bidirectional prediction with CU weights), LIC (local illumination compensation), TM (template matching), and OBMC (overlapping block motion compensation).
[0071] Subtractor 113 can generate a residual block using the difference between the input block and the prediction block. The residual block can be called a residual signal. The residual signal can represent the difference between the original signal and the prediction signal. Alternatively, the residual signal can be a signal generated by transforming or quantizing the difference between the original signal and the prediction signal, or a signal generated by transforming and quantizing the difference between the original signal and the prediction signal. The residual block can be a residual signal on a block-by-block basis.
[0072] Transform unit 130 can generate transform coefficients by transforming the residual block and output the generated transform coefficients. Here, the transform coefficients can be coefficient values generated by transforming the residual block. When transform skip mode is applied, transform unit 130 can skip the transformation of the residual block.
[0073] Quantization levels can be generated by applying quantization to the transform coefficients or residual signals. In the following implementation, quantization levels may also be referred to as transform coefficients.
[0074] For example, a 4x4 lumen residual block generated by intra-frame prediction is transformed using DST (Discrete Sine Transform)-based basis vectors, and then the remaining residual blocks are transformed using DCT (Discrete Cosine Transform)-based basis vectors. Furthermore, RQT (Residual Quadtree) technology is used to partition the transformed blocks into quadtree shapes. After transforming and quantizing each transformed block partitioned by RQT, a coded block flag (CBF) can be transmitted when all coefficients become 0 to improve coding efficiency.
[0075] Another alternative is to apply the Multiple Transform Selection (MTS) technique, which selectively uses multiple transform bases for transformation. That is, instead of segmenting the CU into TUs using RQT, a function similar to TU segmentation can be performed using Sub-Block Transform (SBT). Specifically, SBT applies only to inter-frame prediction blocks. Unlike RQT, the current block can be segmented into 1 / 2 or 1 / 4 blocks vertically or horizontally, and then the transformation is performed on only one of these blocks. For example, in a vertical segmentation, the transformation can be performed on the leftmost or rightmost block; in a horizontal segmentation, the transformation can be performed on the topmost or bottommost block.
[0076] In addition, LFNST (Low Frequency Non-Separable Transform) can be applied. This is a secondary transform technique that performs an additional transform on the residual signal transformed to the frequency domain by DCT or DST. LFNST also performs an additional transform on the 4x4 or 8x8 low-frequency region in the upper left corner, thereby concentrating the residual coefficients in the upper left corner.
[0077] The quantization unit 140 can quantize the transform coefficients or residual signal according to the quantization parameters (QP) to generate a quantization level and output the generated quantization level. Here, the quantization unit 140 can use a quantization matrix to quantize the transform coefficients.
[0078] For example, quantizers with QP values from 0 to 51 can be used. Alternatively, if the image size is large and high coding efficiency is required, QP values from 0 to 63 can be used. Furthermore, the DQ (correlated quantization) method, which uses two quantizers (instead of one), can also be employed. DQ uses two quantizers (e.g., Q0 and Q1) for quantization, but even without conveying information about using a specific quantizer, the quantizer used for the next transform coefficient can be selected based on the current state through a state transition model.
[0079] The entropy coding unit 150 can generate a bitstream by entropy coding based on the value calculated by the quantization unit 140 or the probability distribution of the coding parameter values calculated during encoding, and then output the bitstream. The entropy coding unit 150 can entropy code image sample information and information used for decoding the image. For example, the information used for decoding the image may include syntax elements.
[0080] When applying entropy coding, symbols are represented by allocating fewer bits to symbols with high occurrence probabilities and more bits to symbols with low occurrence probabilities, thereby reducing the size of the bitstream to be encoded. The entropy coding unit 150 can perform entropy coding using methods such as exponential Golomb coding, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). For example, the entropy coding unit 150 can perform entropy coding using a variable-length coding (VLC) table. Furthermore, the entropy coding unit 150 can derive a binarization method for the target symbol and a probability model for the target symbol / bit, and perform arithmetic coding using the derived binarization method and context model.
[0081] Relatedly, when applying CABAC, to reduce the size of the probability table stored in the decoding device, the table probability update method can be changed to a table update method using simple equations. Furthermore, two different probability models can be used to obtain more accurate symbol probability values.
[0082] In order to encode the transform coefficient levels (quantization levels), the entropy coding unit 150 can convert the two-dimensional block form coefficients into one-dimensional vector form by the transform coefficient scanning method.
[0083] The encoding parameters may include information (flags, indexes, etc.) encoded in the encoding device 100 and transmitted to the decoding device 200 via signals, such as syntax elements, as well as information derived during the encoding or decoding process, and may represent the information required when encoding or decoding an image.
[0084] In this paper, sending a signal flag or index can indicate that the corresponding flag or index is entropy encoded in the encoder and included in the bitstream, or it can indicate that the corresponding flag or index is entropy decoded from the bitstream in the decoder.
[0085] The encoded current image can be used as a reference image for another image to be processed later. Therefore, the encoding device 100 can reconstruct or decode the encoded current image again and store the reconstructed or decoded image as a reference image in the reference image buffer 190.
[0086] The quantization level can be dequantized in dequantization unit 160 or inverse transformed in inverse transform unit 170. The coefficients after dequantization and / or inverse transform can be added to the prediction block via adder 117. Here, the coefficients after dequantization and / or inverse transform can refer to the coefficients of at least one of the dequantization and / or inverse transform operations, or they can refer to the reconstructed residual block. Dequantization unit 160 and inverse transform unit 170 can be performed as the inverse process of quantization unit 140 and transform unit 130.
[0087] The reconstructed blocks can be processed by filtering unit 180. Filtering unit 180 can use all or some filtering techniques, applying deblocking filters, sample adaptive offset (SAO), adaptive loop filter (ALF), bilateral filter (BIF), luma mapping and chroma scaling (LMCS), etc., to the reconstructed samples, reconstructed blocks, or reconstructed images. Filtering unit 180 can be referred to as a loop filter. In this case, loop filter is also used as a name, but LMCS is not included.
[0088] Deblocking filters can eliminate block distortion generated at the boundaries between blocks. To determine whether to apply a deblocking filter, samples from several rows or columns contained in the current block can be used. When applying a deblocking filter to a block, different filters can be applied depending on the desired deblocking intensity.
[0089] To compensate for coding errors using sample-adaptive offsets, an appropriate offset value can be added to the sample values. Sample-adaptive offsets can correct the offset between the deblocked image and the original image on a sample-by-sample basis. This can be achieved by dividing the samples in the image into a predetermined number of regions, determining the regions to which the offset will be applied, and then applying the offset to the determined regions; alternatively, the edge information of each sample can be considered when applying the offset.
[0090] For images that have undergone deblocking, the bilateral filter (BIF) can also correct the offset from the original image sample by sample.
[0091] Adaptive loop filters (ALFs) can perform filtering based on a comparison between the reconstructed image and the original image. The samples contained in the image can be divided into predetermined groups, the filter to be applied to each group can be determined, and differential filtering can be performed on each group. Information regarding whether to apply the ALF can be signaled by the coding unit (CU), and the form and coefficients of the adaptive loop filter applied to each block can vary.
[0092] In LMCS (Luminosity Mapping and Chromaticity Scaling), Luminosity Mapping (LM) refers to remapping luminance values using a piecewise linear model, while Chromaticity Scaling (CS) refers to scaling the residual chrominance components based on the average luminance value of the predicted signal. Specifically, LMCS can be used as an HDR correction technique to reflect the characteristics of HDR (High Dynamic Range) images.
[0093] The reconstructed blocks or reconstructed image processed by filtering unit 180 can be stored in reference image buffer 190. The reconstructed blocks processed by filtering unit 180 can be a portion of the reference image. That is, the reference image is a reconstructed image composed of the reconstructed blocks processed by filtering unit 180. The stored reference image can later be used for inter-frame prediction or motion compensation.
[0094] Figure 2 This is a block diagram illustrating the configuration of a decoding apparatus according to an embodiment of the present disclosure.
[0095] The decoding device 200 can be a decoder, a video decoding device, or an image decoding device.
[0096] refer to Figure 2 The decoding device 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra-frame prediction unit 240, a motion compensation unit 250, an adder 201, a switch 203, a filtering unit 260, and a reference image buffer 270.
[0097] The decoding device 200 can receive a bitstream output from the encoding device 100. The decoding device 200 can receive a bitstream stored in a computer-readable recording medium, or it can receive a bitstream streamed via a wired / wireless transmission medium. The decoding device 200 can decode the bitstream in intra-frame mode or inter-frame mode. Furthermore, the decoding device 200 can generate a reconstructed image or a decoded image through decoding, and output the reconstructed image or the decoded image.
[0098] When the prediction mode used for decoding is intra-frame mode, switch 203 can switch to intra-frame mode. Alternatively, when the prediction mode used for decoding is inter-frame mode, switch 203 can switch to inter-frame mode.
[0099] Decoding device 200 can obtain a reconstructed residual block and generate a prediction block by decoding the input bitstream. When the reconstructed residual block and prediction block are obtained, decoding device 200 can generate a reconstructed block by adding the reconstructed residual block and prediction block; this reconstructed block serves as the decoding target. The decoding target block can be referred to as the current block.
[0100] Entropy decoding unit 210 can generate symbols by performing entropy decoding on the bitstream according to the probability distribution. The generated symbols may include symbols in the form of quantization levels. Here, the entropy decoding method can be the inverse process of the entropy encoding method described above.
[0101] The entropy decoding unit 210 can convert one-dimensional vector coefficients into two-dimensional block coefficients by the transform coefficient scanning method to decode the transform coefficient level (quantization level).
[0102] The quantization level can be dequantized in the dequantization unit 220 or inverse transformed in the inverse transform unit 230. The quantization level can be the result of dequantization and / or inverse transform, and can be generated as a reconstruction residual block. Here, the dequantization unit 220 can apply the quantization matrix to the quantization level. The dequantization unit 220 and inverse transform unit 230 applied to the decoding device can employ the same techniques as the dequantization unit 160 and inverse transform unit 170 applied to the encoding device described above.
[0103] When using intra-frame mode, intra-frame prediction unit 240 can generate a prediction block by performing spatial prediction on the current block, which uses sample values from decoded blocks surrounding the target block. Intra-frame prediction unit 240 applied to the decoding apparatus can employ the same techniques as intra-frame prediction unit 120 applied to the encoding apparatus described above.
[0104] When using inter-frame mode, motion compensation unit 250 can generate a prediction block by performing motion compensation on the current block, using motion vectors and a reference image stored in reference image buffer 270. When the value of the motion vector is not an integer, motion compensation unit 250 can generate a prediction block by applying an interpolation filter to a portion of the reference image. To perform motion compensation, the motion compensation method of the prediction unit included in the corresponding coding unit can be determined based on the coding unit: skip mode, merge mode, AMVP mode, or current image reference mode, and motion compensation can be performed according to each mode. The motion compensation unit 250 applied to the decoding device can employ the same technique as the motion compensation unit 122 applied to the coding device described above.
[0105] Adder 201 generates a reconstructed block by adding the reconstructed residual block and the predicted block. Filtering unit 260 can apply at least one of inverse LMCS, deblocking filter, sample adaptive offset, and adaptive loop filter to the reconstructed block or reconstructed image. Filtering unit 260 applied to the decoding device can employ the same filtering technique as that used by filtering unit 180 applied to the encoding device described above.
[0106] The filtering unit 260 can output a reconstructed image. The reconstructed blocks or reconstructed image can be stored in the reference image buffer 270 and used for inter-frame prediction. The reconstructed blocks processed by the filtering unit 260 can be used as part of the reference image. That is, the reference image can be a reconstructed image composed of the reconstructed blocks processed by the filtering unit 260. The stored reference image can later be used for inter-frame prediction or motion compensation.
[0107] Figure 3 This is a schematic diagram illustrating a video coding system to which this disclosure applies.
[0108] According to one embodiment, a video encoding system may include an encoding device 10 and a decoding device 20. The encoding device 10 may transmit encoded video and / or image information or data to the decoding device 20 in the form of a file or stream via a digital storage medium or network.
[0109] The encoding apparatus 10 according to the embodiments may include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. The decoding apparatus 20 according to the embodiments may include a receiving unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 may be referred to as a video / image encoding unit, and the decoding unit 22 may be referred to as a video / image decoding unit. The transmission unit 13 may be included in the encoding unit 12. The receiving unit 21 may be included in the decoding unit 22. The rendering unit 23 may include a display unit, and the display unit may be configured as a separate device or an external component.
[0110] The video source generation unit 11 can acquire video / images through a process of capturing, synthesizing, or generating video / images. The video source generation unit 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device may include, for example, a computer, tablet computer, and smartphone, etc., and can generate video / images (electronically). For example, virtual video / images can be generated by a computer, etc., in which case the video / image capture process can be replaced by a process of generating related data.
[0111] Encoding unit 12 can encode the input video / image. Encoding unit 12 can perform a series of processes, such as prediction, transformation, and quantization, to improve compression and encoding efficiency. Encoding unit 12 can output encoded data (encoded video / image information) in the form of a bitstream. Detailed configuration of encoding unit 12 can also be as described above. Figure 1 The encoding device 100 is configured in the same way.
[0112] The transmitting unit 13 can output encoded video / image information or data as a bitstream, and transmit it to the receiving unit 21 of the decoding device 20 in the form of a file or stream via a digital storage medium or network. The digital storage medium can include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit 13 can include elements for generating media files according to a predetermined file format, and can also include elements for transmission via a broadcast / communication network. The receiving unit 21 can extract / receive the bitstream from the storage medium or network and send it to the decoding unit 22.
[0113] Decoding unit 22 can decode video / images by performing a series of steps (e.g., dequantization, inverse transform, and prediction) corresponding to the operations of encoding unit 12. The detailed configuration of decoding unit 22 can also be as described above. Figure 2 The decoding device 200 is configured in the same way.
[0114] Rendering unit 23 can render the decoded video / image. The rendered video / image can be displayed through the display unit.
[0115] This disclosure describes a method for refining the determined intra-prediction mode in intra-prediction by correcting the intra-prediction mode, thereby improving the prediction accuracy of intra-prediction. To correct the intra-prediction mode, the decoder can select the candidate intra-prediction mode with the highest prediction accuracy from multiple candidate intra-prediction modes based on a template. Alternatively, the decoder can select one from multiple candidate intra-prediction modes based on information sent from the encoder. As described above, the prediction accuracy of intra-prediction can be improved by correcting the intra-prediction mode.
[0116] The following describes various implementations of the decoder-side intra-prediction mode refinement (DIMR) method.
[0117] Figure 4 This demonstrates a template-based intra-frame mode derivation (TIMD) method.
[0118] Figure 4 This diagram shows an M x N size template 402 for the current block 400 and reference pixels 404 for generating the pixel prediction values of template 402. Template 402 for the current block 400 includes at least one of the left and upper regions of the current block 400 and has been reconstructed prior to the current block 400. The left region is L1 x N in size, and the upper region is composed of an upper template of size M x L2. Here, M, N, L1, and L2 are arbitrary positive integers.
[0119] Based on a template-based intra-mode derivation method, the applicability of each candidate mode in the Most Probable Mode (MPM) list for the current block 400 is calculated. To calculate applicability, a predicted value for the template is generated from the reference pixel of the template based on the directionality of the corresponding candidate mode. Furthermore, the sum of absolute transform differences (SATD) between the predicted and reconstructed values of the template pixels is calculated. Finally, among all candidate modes in the MPM list, the mode with the smallest SATD is selected as the intra-prediction mode.
[0120] In template-based intra-mode derivation methods, various distortion measurement methods (such as Sum of Absolute Differences (SAD) or Sum of Squared Differences (SSD)) can be used instead of SATD. Furthermore, in template-based intra-mode derivation methods, distortion is measured not only for candidate modes in the MPM list but also for all intra-prediction modes, thus allowing selection of an intra-prediction mode from among all intra-prediction modes. Moreover, since distortion is measured on a list consisting of arbitrary intra-prediction candidate modes, an intra-prediction mode can be selected from any of the intra-prediction modes.
[0121] Based on the decoder-side intra-prediction mode refinement (DIMR) method, an initial intra-prediction mode is first determined. Furthermore, the initial intra-prediction mode is corrected using a template-based intra-prediction method. Specifically, among candidate intra-prediction modes that include the initial intra-prediction mode, the optimal candidate mode is selected using a template-based intra-prediction method. Finally, the optimal candidate mode is determined as the final intra-prediction mode.
[0122] When the decoder-side intra-prediction mode refinement mode is allowed to be determined by the syntax elements of the higher-level data units of the block (e.g., sequences, images, stripes, tiles, and coding tree units), the candidate intra-prediction modes may include the initial intra-prediction mode and its neighboring candidate modes. This is because the decoder needs to determine whether the initial intra-prediction mode has lower distortion than all neighboring candidate modes without block unit data.
[0123] Alternatively, when the refinement mode of the decoder-side intra-prediction mode is ultimately determined by the syntax elements of the block unit, the candidate intra-prediction modes may only include the candidate modes adjacent to the initial intra-prediction mode. This is because when the encoder determines that the initial intra-prediction mode of the current block has the lowest distortion, the syntax elements in the block unit encoded by the encoder will be encoded to indicate that decoder-side intra-prediction mode refinement is not required.
[0124] The prediction direction of the adjacent candidate modes of the initial intra-prediction mode can be between the prediction direction of the first adjacent intra-prediction mode (whose index is 1 less than the index of the initial intra-prediction mode) and the prediction direction of the initial intra-prediction mode. Here, the adjacent candidate modes of the initial intra-prediction mode can include S intra-prediction modes, whose prediction directions are between the prediction direction of the first adjacent intra-prediction mode and the prediction direction of the initial intra-prediction mode. For example, when the index of the initial intra-prediction mode is 34, S intra-prediction modes can be included as adjacent candidate modes, whose prediction directions are between the prediction direction of intra-prediction mode 33 (which is the first adjacent intra-prediction mode) and the prediction direction of intra-prediction mode 34 (which is the initial intra-prediction mode). S is any positive integer greater than or equal to 1.
[0125] The prediction direction of the adjacent candidate modes of the initial intra-prediction mode can be between the prediction direction of the second adjacent intra-prediction mode (whose index is one greater than the initial intra-prediction mode's index) and the prediction direction of the initial intra-prediction mode. Here, the adjacent candidate modes of the initial intra-prediction mode can include S intra-prediction modes, whose prediction directions are between the prediction directions of the second adjacent intra-prediction mode and the prediction direction of the initial intra-prediction mode. For example, when the index of the initial intra-prediction mode is 34, S intra-prediction modes can be included as adjacent candidate modes, whose prediction directions are between the prediction directions of intra-prediction mode 34 (which is the initial intra-prediction mode) and the prediction directions of intra-prediction mode 35 (which is the second adjacent intra-prediction mode). S is any positive integer greater than or equal to 1.
[0126] Figure 5 An implementation of a method for determining neighboring candidate modes of an initial intra-frame prediction mode is shown. According to... Figure 5 The initial intra-frame prediction mode is mode 63, and the first adjacent intra-frame prediction mode and the second adjacent intra-frame prediction mode are mode 62 and mode 64, respectively.
[0127] Adjacent candidate modes can include mode 63-1 / 2 (prediction direction between the prediction direction of mode 62 and the prediction direction of mode 63) and mode 63+1 / 2 (prediction direction between the prediction direction of mode 63 and the prediction direction of mode 64). Furthermore, according to the template-based intra-mode derivation method, among modes 63, 63-1 / 2, and 63+1 / 2, the intra-prediction mode with the smallest distortion value can be determined as the intra-prediction mode for the current block.
[0128] Figure 6 Another implementation of a method for determining neighboring candidate modes of an initial intra-frame prediction mode is shown. According to... Figure 6 ,and Figure 5 Similarly, the initial intra-frame prediction mode is mode 63, and the first adjacent intra-frame prediction mode and the second adjacent intra-frame prediction mode are mode 62 and mode 64, respectively.
[0129] Adjacent candidate modes can include modes 63-1 / 3 and 63-2 / 3 (which exist between the prediction directions of mode 62 and mode 63), and modes 63+1 / 3 and 63+2 / 3 (which exist between the prediction directions of mode 63 and mode 64). Furthermore, according to the template-based intra-mode derivation method, the intra-prediction mode with the smallest distortion value among modes 63, 63-1 / 3, 63-2 / 3, 63+1 / 3, and 63+2 / 3 can be determined as the intra-prediction mode for the current block.
[0130] Figure 5 The total number of adjacent candidate patterns is 2, while Figure 6 The total number of adjacent candidate patterns is 4. For example... Figure 6 As shown, increasing the number of neighboring candidate modes may improve the intra-prediction accuracy of the current block. However, the computational burden on the decoder may increase during the process of selecting the best neighboring candidate mode based on distortion. Therefore, considering the computational burden on the decoder, the number of neighboring candidate modes can be increased. However, unlike selecting neighboring candidate modes based on distortion in the decoder, when selecting neighboring candidate modes based on neighboring candidate mode information encoded in the encoder, even increasing the number of neighboring candidate modes does not impose any computational burden on the decoder that calculates the distortion of neighboring candidate modes. Therefore, when selecting neighboring candidate modes based on neighboring candidate mode information, more neighboring candidate modes can be used for intra-prediction mode refinement.
[0131] Figure 7 An implementation of a method for determining neighboring candidate modes of an initial intra-prediction mode is shown, wherein the initial intra-prediction mode is the intra-prediction mode with the smallest index among directional modes within a predetermined range. According to... Figure 7 The intra-prediction mode with the smallest index among the directional modes within the predetermined range is mode 2.
[0132] Adjacent candidate modes can include mode 2+1 / 2, whose prediction direction lies between the prediction direction of mode 2 and the prediction direction of mode 3. Furthermore, based on a template-based intra-mode derivation method, the intra-prediction mode with the smallest distortion value between mode 2 and mode 2+1 / 2 can be determined as the intra-prediction mode for the current block.
[0133] Figure 8 Another embodiment of a method for determining neighboring candidate modes of an initial intra-prediction mode is shown, wherein the initial intra-prediction mode is the intra-prediction mode with the smallest index among directional modes within a predetermined range. According to Figure 8 Among the directional modes within the predetermined range, the intra-prediction mode with the smallest index is mode 2.
[0134] Adjacent candidate modes can include mode 2+1 / 3 and mode 2+2 / 3, whose prediction directions lie between the prediction directions of mode 2 and mode 3. Furthermore, based on a template-based intra-mode derivation method, the intra-prediction mode with the smallest distortion value among mode 2, mode 2+1 / 3, and mode 2+2 / 3 can be determined as the intra-prediction mode for the current block.
[0135] exist Figure 7 and Figure 8 In the example above, pattern 2 is described as having the smallest index in the direction patterns, but this is just an example; the smallest index in the direction patterns can be different values. Figure 5 andFigure 6 The implementation method is similar, in Figure 7 and Figure 8 In this implementation, while increasing the number of neighboring candidate modes improves prediction accuracy, it also increases the computational burden on the decoder. Therefore, considering the computational burden on the decoder, the number of neighboring candidate modes can be increased. Furthermore, when neighboring candidate modes are selected based on the neighboring candidate mode information encoded in the encoder, the computational burden on the decoder decreases, thus allowing for the use of more neighboring candidate modes for intra-frame prediction mode refinement.
[0136] Figure 9 This illustrates an implementation of a method for determining neighboring candidate modes of the initial intra-prediction mode when the initial intra-prediction mode is the intra-prediction mode with the largest index among directional modes within a predetermined range. According to... Figure 9 Among the directional modes within the predetermined range, the intra-prediction mode with the largest index is mode 66.
[0137] Adjacent candidate modes may include mode 66-1 / 2, whose prediction direction lies between the prediction directions of mode 65 and mode 66. Furthermore, based on a template-based intra-mode derivation method, the intra-prediction mode with the smallest distortion value between mode 66-1 / 2 and mode 66 can be determined as the intra-prediction mode for the current block.
[0138] Figure 10 This illustrates another implementation of a method for determining neighboring candidate modes of the initial intra-prediction mode when the initial intra-prediction mode is the intra-prediction mode with the largest index among directional modes within a predetermined range. According to... Figure 10 Among the directional patterns within the predetermined range, the intra-prediction pattern with the largest index is pattern 66.
[0139] Neighboring candidate modes can include modes 66-1 / 3 and 66-2 / 3, whose prediction directions lie between the prediction directions of mode 65 and mode 66. Furthermore, based on a template-based intra-mode derivation method, the intra-prediction mode with the smallest distortion value among modes 66-1 / 3, 66-2 / 3, and 66 can be determined as the intra-prediction mode for the current block.
[0140] exist Figure 9 and Figure 10 In this example, mode 66 has a large index among directional modes, but this is just an example, and the maximum index of the directional mode can be determined to be a different value depending on the total number of intra-frame prediction modes. (As mentioned above...) Figure 5 and Figure 6 The implementation method is similar, in Figure 9 and Figure 10In this implementation, while increasing the number of neighboring candidate modes improves prediction accuracy, it also increases the computational burden on the decoder. Therefore, considering the computational burden on the decoder, the number of neighboring candidate modes can be increased. Furthermore, when neighboring candidate modes are selected based on the neighboring candidate mode information encoded in the encoder, the computational burden on the decoder decreases, thus allowing for the use of more neighboring candidate modes for intra-frame prediction mode refinement.
[0141] exist Figures 5 to 10 In this implementation, one or two intra-prediction modes are included as adjacent candidate modes in each index direction of the initial intra-prediction mode. However, this is merely an example; S intra-prediction modes can be included as adjacent candidate modes in each index direction of the initial intra-prediction mode. S is any positive integer greater than or equal to 3.
[0142] exist Figures 5 to 10 In this method, neighboring candidate modes are determined by equally dividing the prediction direction difference between the initial intra-frame prediction mode and adjacent intra-frame prediction modes. However, according to the implementation, neighboring candidate modes can be determined by unevenly dividing the prediction direction difference. Furthermore, neighboring candidate modes can be determined based on any direction adjacent to the direction indicated by the initial intra-frame prediction mode.
[0143] Equal segmentation can represent the equal distribution of the angular difference in the prediction direction between the initial intra-prediction mode and adjacent intra-prediction modes. Alternatively, equal segmentation can represent the equal distribution of the positional difference between the reference positions indicated by the prediction directions of the initial intra-prediction mode and adjacent intra-prediction modes.
[0144] exist Figures 5 to 10 In this approach, adjacent candidate modes are determined based on neighboring intra-prediction modes immediately adjacent to the initial intra-prediction mode. However, according to an implementation, adjacent candidate modes can be determined based on adjacent intra-prediction modes whose indices differ from the index of the initial intra-prediction mode by a factor of A. For example, when the initial intra-prediction mode is 32 and A is 2, the first adjacent intra-prediction mode can be determined as 30, and the second adjacent intra-prediction mode can be determined as 34. Furthermore, the adjacent candidate modes can include multiple intra-prediction modes with index values ranging from 30 to 34. Here, A is any positive integer greater than or equal to 1.
[0145] Whether to apply the decoder-side intra-predictive mode refinement (DIMR) method can be determined based on the following conditions.
[0146] According to one implementation, when the initial intra-prediction mode of the current block is a non-directional mode (e.g., planar mode or DC mode), the decoder-side intra-prediction mode refinement method may not be applied to the current block.
[0147] According to one implementation, the application of a decoder-side intra-prediction mode refinement method to the current block can be determined based on the size of the current block. For example, when the area of the current block is equal to or less than M luma samples (M is any positive integer greater than or equal to 1), the decoder-side intra-prediction mode refinement method may not be applied to the current block. Here, the area of the current block can be measured by the number of luma samples contained in the current block. As another example, when the smaller of the width and height of the current block is equal to or less than K (K is any positive integer greater than or equal to 1), the decoder-side intra-prediction mode refinement method may not be applicable to the current block. Alternatively, when the aspect ratio (width / height or height / width) of the current block is within a predetermined range, the decoder-side intra-prediction mode refinement method may not be applicable to the current block.
[0148] According to one implementation, when applying IntraTMP (Intra-Temporal Matching Prediction) mode or IntraBlock Copy (IBC) mode to the current block, the decoder-side intra-prediction mode refinement method may not be applicable to the current block.
[0149] According to one implementation, when the index value of the initial intra-prediction mode is equal to or greater than T or equal to or less than U (T and U are mutually exclusive arbitrary integers), the decoder-side intra-prediction mode refinement method may not be applicable to the current block. The values of T and U can be determined based on the total number of intra-prediction modes.
[0150] When generating the MPM list for the current block and determining the direct mode (DM) for the chroma components, the intra-prediction mode of the reference block is referenced. At this point, when the decoder-side intra-prediction mode refinement method is applied to the reference block, the initial intra-prediction mode of the reference block (rather than its final intra-prediction mode) can be used to generate the MPM list and determine the direct mode (DM) for the chroma components.
[0151] According to the decoder-side intra-prediction mode refinement method, the computational burden on the decoder side may increase when the decoder calculates the distortion of adjacent candidate modes. Therefore, the decoder can choose not to calculate the distortion of adjacent candidate modes, and instead, the encoder can encode the intra-prediction mode refinement information for the adjacent candidate modes applied to the current block. Furthermore, the decoder can determine the optimal adjacent candidate modes for encoder-side intra-prediction mode refinement (EIMR) based on the intra-prediction mode refinement information.
[0152] According to one implementation, intra-prediction mode refinement information may include intra-prediction mode refinement application information indicating whether intra-prediction mode refinement is applied, and intra-prediction mode refinement index information indicating the best neighboring candidate mode among neighboring candidate modes. When the intra-prediction mode refinement application information indicates that intra-prediction mode refinement is applied to the current block or its upper-level units (e.g., coding tree blocks, stripes, tiles, or pictures), the best candidate intra-prediction mode can be selected as the final intra-prediction mode for the current block based on the intra-prediction mode refinement index information. On the other hand, when the intra-prediction mode refinement application information indicates that intra-prediction mode refinement is not applied to the current block or its upper-level units, an initial intra-prediction mode can be selected as the final intra-prediction mode for the current block.
[0153] According to the decoder-side intra-prediction mode refinement method, the initial intra-prediction mode is also included as a candidate intra-prediction mode. However, according to the decoder-side intra-prediction mode refinement method, when the intra-prediction mode refinement application information indicates that intra-prediction mode refinement is to be applied, the initial intra-prediction mode will be excluded from the candidate intra-prediction modes.
[0154] Intra-prediction mode refinement application information can be encoded as a 1-bit codeword. The intra-prediction mode refinement index information indicates the best candidate intra-prediction mode among the candidate intra-prediction modes. Taking into account the number of neighboring candidate modes, the best candidate intra-prediction mode is mapped to a predetermined index. For example, when there are four neighboring candidate modes, the best neighboring candidate mode is mapped to a value between 0 and 3. Furthermore, the intra-prediction mode refinement index information is configured to indicate the mapping value. The intra-prediction mode refinement index information can be encoded based on fixed-length code (FLC), truncated unary code (TU), or signed exponential-Golomb code.
[0155] Intra-prediction mode refinement index information can be encoded by taking into account the frequency of occurrence of neighboring candidate modes. Specifically, shorter codewords can be assigned to neighboring candidate modes with higher occurrence frequencies, while longer codewords can be assigned to neighboring candidate modes with lower occurrence frequencies. Furthermore, the frequency of occurrence of neighboring candidate modes may differ due to their index difference from the initial intra-prediction mode. For example, the neighboring candidate mode with the smallest absolute index difference from the initial intra-prediction mode may have the highest occurrence frequency. Therefore, a shorter codeword can be assigned to the neighboring candidate mode with the smallest absolute index difference from the initial intra-prediction mode. Conversely, a longer codeword can be assigned to the neighboring candidate mode with the largest absolute index difference from the initial intra-prediction mode.
[0156] According to one implementation, the intra-prediction mode refinement index information may include intra-prediction mode refinement sign information and intra-prediction mode refinement index difference information. The intra-prediction mode refinement sign information indicates whether the best candidate intra-prediction mode is in a positive (+) or negative (-) direction relative to the initial intra-prediction mode. The intra-prediction mode refinement index difference information indicates the index difference between the best candidate intra-prediction mode and the initial intra-prediction mode.
[0157] The intra-prediction mode refinement symbol information can be encoded as a 1-bit codeword. When the initial intra-prediction mode is the intra-prediction mode with the smallest index or the intra-prediction mode with the largest index among the directional modes within a predetermined range, the intra-prediction mode refinement symbol information does not need to be encoded, and only the intra-prediction mode refinement index difference information needs to be encoded.
[0158] When there is only one adjacent candidate mode in each direction, the intra-prediction mode refinement index difference information does not need to be encoded. When there are two or more adjacent candidate modes in each direction, the intra-prediction mode refinement index difference information can be encoded based on fixed-length code (FLC), truncated unary code (TU), or signed exponential-Golomb code.
[0159] Refining the index difference information in intra-frame prediction modes can take into account the frequency of occurrence of neighboring candidate modes when encoding. Furthermore, the frequency of occurrence of neighboring candidate modes may differ due to their index difference from the initial intra-frame prediction mode. For example, the neighboring candidate mode with the smallest absolute index difference from the initial intra-frame prediction mode may have the highest frequency of occurrence. Therefore, a short codeword can be assigned to the index of the neighboring candidate mode with the smallest absolute index difference from the initial intra-frame prediction mode. Conversely, a long codeword can be assigned to the index of the neighboring candidate mode with the largest absolute index difference from the initial intra-frame prediction mode.
[0160] Figure 11 This illustrates an implementation of an intra-prediction method refined by applying intra-prediction modes.
[0161] In step S1102, the initial intra-prediction mode of the current block is determined. The initial intra-prediction mode of the current block can be determined by referring to the intra-prediction mode of a reference block. When intra-prediction refinement is applied to the reference block of the current block, the initial intra-prediction mode of the current block can be determined based on the initial intra-prediction mode of the reference block. The reference block can be a reference block used to determine the most probable mode or direct mode of the current block. When determining the MPM of the current block, the reference block can be a spatially adjacent block of the current block. When determining the DM of the chrominance component of the current block, the reference block can be the corresponding block of the luma component at the corresponding position of the current block.
[0162] In step S1104, based on the initial intra-prediction mode of the current block, one or more adjacent candidate modes adjacent to the initial intra-prediction mode are determined.
[0163] According to one implementation, one or more adjacent candidate modes may include: a first adjacent candidate mode whose prediction direction is between the prediction direction of a first adjacent intra-prediction mode in the current block and the prediction direction of an initial intra-prediction mode; and a second adjacent candidate mode whose prediction direction is between the prediction direction of a second adjacent intra-prediction mode in the current block and the prediction direction of the initial intra-prediction mode. The index of the first adjacent intra-prediction mode is smaller than the index of the initial intra-prediction mode by a predetermined value, and the index of the second adjacent intra-prediction mode is larger than the index of the initial intra-prediction mode by a predetermined value.
[0164] According to one implementation, the first neighboring candidate mode is determined by equally dividing the prediction direction of the first neighboring intra-prediction mode in the current block and the prediction direction of the initial intra-prediction mode based on the number of first neighboring candidate modes. That is, the direction difference between the first neighboring candidate modes and the direction difference between the first neighboring candidate mode and the initial intra-prediction mode are determined equally.
[0165] Similarly, the second neighboring candidate mode is determined by equally dividing the prediction direction of the second neighboring intra-prediction mode in the current block with the prediction direction of the initial intra-prediction mode, based on the number of second neighboring candidate modes. In other words, the direction difference between the second neighboring candidate modes and the direction difference between the second neighboring candidate mode and the initial intra-prediction mode are determined equally.
[0166] According to one implementation, when the initial intra-prediction mode is the intra-prediction mode with the smallest index among the directional modes within a predetermined range, only the second adjacent intra-prediction mode can be included in one or more adjacent candidate modes. On the other hand, when the initial intra-prediction mode is the intra-prediction mode with the largest index among the directional modes within a predetermined range, only the first adjacent intra-prediction mode can be included in one or more adjacent candidate modes.
[0167] In step S1106, the final intra-prediction mode of the current block is determined from among the candidate intra-prediction modes that include one or more adjacent candidate modes.
[0168] According to one implementation, among the candidate intra-prediction modes of the current block, the mode with the least distortion is determined as the final intra-prediction mode for the current block. Here, the distortion of the candidate intra-prediction modes can be determined based on the difference between the reconstructed samples contained in the adjacent templates of the current block and the predicted samples corresponding to the reconstructed samples. Furthermore, the predicted samples can be determined based on the template reference samples adjacent to the templates and the prediction direction of the candidate intra-prediction modes.
[0169] According to one implementation, in the candidate intra-prediction modes of the current block, intra-prediction mode refinement index information indicating the final intra-prediction mode of the current block can be obtained from the bitstream. Furthermore, the final intra-prediction mode of the current block can be determined based on the intra-prediction mode refinement index information.
[0170] According to one implementation, the intra-prediction mode refinement index information may include: intra-prediction mode refinement sign information, indicating whether the final intra-prediction mode is positive (+) or negative (-) relative to the initial intra-prediction mode; and intra-prediction mode refinement index difference information, indicating the index difference between the final intra-prediction mode and the initial intra-prediction mode. Therefore, the final intra-prediction mode of the current block can be derived based on the sign information based on the intra-prediction mode refinement sign information and the index difference information based on the intra-prediction mode refinement index difference information.
[0171] According to one implementation, codewords of predetermined length are assigned to each symbol of an index difference indicated by the refined index difference information of the intra-prediction mode. Here, the codeword length assigned to the candidate intra-prediction mode can be determined based on the index difference between the initial intra-prediction mode and the candidate intra-prediction mode. For example, shorter codewords can be assigned to symbols with smaller index differences.
[0172] In step S1108, the current block is predicted based on the final intra-frame prediction mode of the current block.
[0173] According to one embodiment, the image decoding method may further include determining whether to perform intra-prediction refinement on the current block. When intra-prediction refinement is performed on the current block, steps S1104 to S1108 may be executed. When intra-prediction refinement is not performed on the current block, steps S1104 to S1108 are omitted, and the current block is predicted based on the initial intra-prediction mode of the current block.
[0174] According to one implementation, when the initial intra-prediction mode is a non-directional intra-prediction mode, it can be determined that intra-prediction refinement will not be performed on the current block.
[0175] According to one implementation, when the index of the initial intra-prediction mode is not included in a predetermined range, it can be determined that intra-prediction refinement should not be performed on the current block.
[0176] According to one implementation, it can be determined whether to perform intra-frame prediction refinement on the current block based on at least one of the current block's width, height, area, and aspect ratio.
[0177] Based on the prediction method executed in steps S1102 to S1108, the current block can be encoded or decoded. In some embodiments, the intra-prediction mode refinement index information is encoded in the encoder according to the prediction result. Furthermore, the intra-prediction mode refinement index information can be decoded in the decoder and used to derive the final intra-prediction mode.
[0178] Furthermore, the bit stream generated by the encoder according to the prediction method performed in steps S1102 to S1108 can be stored in a recording medium or transmitted to the outside of the encoder.
[0179] Figure 12 An exemplary content streaming system to which embodiments of the present disclosure may be applied is shown.
[0180] like Figure 12 As shown, the content streaming media system using the embodiments of this disclosure mainly includes an encoding server, a streaming media server, a web server, a media storage device, a user device, and a multimedia input device.
[0181] The encoding server compresses content received from multimedia input devices such as smartphones, cameras, and CCTV into digital data to generate a bitstream, which is then transmitted to the streaming media server. Alternatively, if the multimedia input devices such as smartphones, cameras, and CCTV directly generate the bitstream, the encoding server can be omitted.
[0182] The bitstream can be generated by the image encoding method and / or image encoding apparatus applied to the embodiments of this disclosure, and the streaming media server can temporarily store the bitstream during transmission or reception.
[0183] The streaming media server transmits multimedia data to the user's device via a web server based on the user's request. The web server can also act as an intermediary, informing the user of any available services. When a user requests a service from the web server, the web server transmits it to the streaming media server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, which can control the commands / responses between devices within the content streaming system.
[0184] A streaming media server can receive content from media storage and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming media server can store the bitstream for a period of time.
[0185] Examples of user devices may include mobile phones, smartphones, laptops, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, board PCs, tablets, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc.
[0186] Each server in the aforementioned content streaming system can operate as a distributed server, in which case data received from each server can be distributed and processed.
[0187] The above embodiments can be performed in the same or corresponding manner in the encoding and decoding devices. Furthermore, at least one or a combination of at least one of the above embodiments can be used to encode / decode images.
[0188] The order in which the above embodiments are applied in the encoding and decoding devices may be different. Alternatively, the order in which the above embodiments are applied in the encoding and decoding devices may be the same.
[0189] The above implementation methods can be performed separately for luminance signals and chrominance signals. Alternatively, the above implementation methods for luminance signals and chrominance signals can also be performed in the same way.
[0190] In the above embodiments, the method is described based on a flowchart containing a series of steps or units. However, this disclosure is not limited to the order of the steps; on the contrary, some steps may be performed simultaneously or in a different order than other steps. Furthermore, those skilled in the art should understand that the steps in the flowchart are not mutually exclusive, and other steps may be added to or some steps may be deleted from the flowchart without affecting the scope of this disclosure.
[0191] These implementations can be implemented in the form of program instructions, which can be executed by various computer components and recorded in a computer-readable recording medium. The computer-readable recording medium may include individual program instructions, data files, data structures, etc., or combinations of such program instructions, data files, data structures, etc. The program instructions recorded in the computer-readable recording medium may be specifically designed and constructed for this disclosure, or may be well known to those skilled in the art of computer software.
[0192] The bitstream generated by the encoding method according to the above embodiments can be stored in a non-transitory computer-readable recording medium. Furthermore, the bitstream stored in the non-transitory computer-readable recording medium can be decoded by the decoding method according to the above embodiments.
[0193] Examples of computer-readable recording media include: magnetic recording media, such as hard disks, floppy disks, and magnetic tapes; optical data storage media, such as CD-ROMs or DVD-ROMs; magnetically optimized media, such as optical-to-floppy disks; and hardware devices specifically designed for storing and executing program instructions, such as read-only memory (ROM), random access memory (RAM), flash memory, etc. Examples of program instructions include not only machine language code formatted by a compiler but also high-level language code executable by a computer using an interpreter. Hardware devices can be configured to be operated by one or more software modules, and vice versa, to perform processes according to this disclosure.
[0194] Although this disclosure has been described with reference to specific items (e.g., detailed elements) and limited embodiments and drawings, these descriptions are provided only to help to more fully understand the invention, and this disclosure is not limited to the described embodiments. Those skilled in the art to which this disclosure pertains will understand that various modifications and changes can be made based on the above description.
[0195] Therefore, the spirit of this disclosure should not be limited to the above-described embodiments, and the entire scope of the appended claims and their equivalents should fall within the scope and spirit of this invention.
[0196] Industrial applicability
[0197] This disclosure can be used for apparatus for encoding / decoding images and for recording media for storing bit streams.
Claims
1. A method for decoding an image, the method comprising the following steps: Determine the initial intra-prediction mode for the current block; Based on the initial intra-prediction mode of the current block, determine one or more adjacent candidate modes that are adjacent to the initial intra-prediction mode. Among the candidate intra-prediction modes that include the one or more adjacent candidate modes, determine the final intra-prediction mode of the current block; as well as Predict the current block based on the final intra-frame prediction mode of the current block.
2. The method according to claim 1, further comprising the following step: Determine whether to perform intra-frame prediction refinement on the current block. Specifically, when performing the intra-prediction refinement on the current block, based on the initial intra-prediction mode of the current block, one or more adjacent candidate modes adjacent to the initial intra-prediction mode are determined; among the candidate intra-prediction modes including the one or more adjacent candidate modes, the final intra-prediction mode of the current block is determined; and the current block is predicted based on the final intra-prediction mode of the current block. Specifically, when the intra-prediction refinement is not performed on the current block, the current block is predicted based on the initial intra-prediction mode of the current block.
3. The method according to claim 2, wherein, When the initial intra-prediction mode is a non-directional intra-prediction mode, it is determined that intra-prediction refinement will not be performed on the current block.
4. The method according to claim 2, wherein, If the index of the initial intra-prediction mode is not included in the predetermined range, it is determined that intra-prediction refinement will not be performed on the current block.
5. The method according to claim 2, wherein, Whether to perform intra-frame prediction refinement on the current block is determined based on at least one of the current block's width, height, area, and aspect ratio.
6. The method according to claim 1, wherein, When intra-prediction refinement is applied to a reference block of the current block, the initial intra-prediction mode of the current block is determined based on the initial intra-prediction mode of the reference block.
7. The method according to claim 1, wherein, The one or more adjacent candidate modes include: a first adjacent candidate mode having a prediction direction between the prediction direction of the first adjacent intra-prediction mode of the current block and the prediction direction of the initial intra-prediction mode; and a second adjacent candidate mode having a prediction direction between the prediction direction of the second adjacent intra-prediction mode of the current block and the prediction direction of the initial intra-prediction mode. Wherein, the first adjacent intra-prediction mode has an index that is one predetermined value smaller than the index of the initial intra-prediction mode; The second adjacent intra-prediction mode has an index that is one predetermined value larger than the index of the initial intra-prediction mode.
8. The method according to claim 7, wherein, The first neighboring candidate mode is determined by equally dividing the prediction direction of the first neighboring intra-prediction mode of the current block and the prediction direction of the initial intra-prediction mode according to the number of the first neighboring candidate modes; and The second adjacent candidate mode is determined by equally dividing the prediction direction of the second adjacent intra-prediction mode of the current block and the prediction direction of the initial intra-prediction mode according to the number of the second adjacent candidate modes.
9. The method according to claim 7, wherein, When the initial intra-prediction mode is the intra-prediction mode with the smallest index among the directional modes within a predetermined range, the one or more adjacent candidate modes include only the second adjacent intra-prediction mode. and Wherein, when the initial intra-frame prediction mode is the intra-frame prediction mode with the largest index among the directional modes within a predetermined range, the one or more adjacent candidate modes only include the first adjacent intra-frame prediction mode.
10. The method according to claim 1, wherein, Among the candidate intra-prediction modes of the current block, the mode with the least distortion is determined as the final intra-prediction mode of the current block.
11. The method according to claim 10, wherein, The distortion of the candidate intra-prediction mode is determined based on the difference between the reconstructed sample contained in the adjacent template of the current block and the prediction sample corresponding to the reconstructed sample.
12. The method according to claim 11, wherein, The predicted sample is determined based on the template reference sample adjacent to the template and the prediction direction of the candidate intra-frame prediction mode.
13. The method according to claim 1, further comprising: Intra-prediction mode refinement index information is obtained from the bitstream. This intra-prediction mode refinement index information indicates the final intra-prediction mode of the current block among the candidate intra-prediction modes of the current block. The final intra-prediction mode of the current block is determined based on the refined index information of the intra-prediction mode.
14. The method according to claim 13, wherein, The intra-prediction mode refinement index information includes: intra-prediction mode refinement sign information, which indicates whether the final intra-prediction mode is positive (+) or negative (-) relative to the initial intra-prediction mode; and intra-prediction mode refinement index difference information, which indicates the index difference between the final intra-prediction mode and the initial intra-prediction mode.
15. The method according to claim 14, wherein, In the refined index difference information of the intra-prediction mode, the codeword length allocated to the candidate intra-prediction mode is determined based on the index difference between the initial intra-prediction mode and the candidate intra-prediction mode.
16. A method for encoding an image, the method comprising the following steps: Determine the initial intra-prediction mode for the current block; Based on the initial intra-prediction mode of the current block, determine one or more adjacent candidate modes that are adjacent to the initial intra-prediction mode. Among the candidate intra-prediction modes that include the one or more adjacent candidate modes, determine the final intra-prediction mode of the current block; as well as Predict the current block based on the final intra-frame prediction mode of the current block.
17. A computer-readable recording medium for storing a bitstream generated by a method for encoding images, in, The method for encoding the image includes: Determine the initial intra-prediction mode for the current block; Based on the initial intra-prediction mode of the current block, determine one or more adjacent candidate modes that are adjacent to the initial intra-prediction mode. Among candidate intra-prediction modes including the one or more adjacent candidate modes, determine the final intra-prediction mode of the current block; and Predict the current block based on the final intra-frame prediction mode of the current block.
18. A method for transmitting a bitstream generated by a method of encoding an image, the method comprising: The image is encoded based on the aforementioned image encoding method; as well as Transmit a bitstream containing the encoded image. The method for encoding the image includes: Determine the initial intra-prediction mode for the current block; Based on the initial intra-prediction mode of the current block, determine one or more adjacent candidate modes that are adjacent to the initial intra-prediction mode. Among candidate intra-prediction modes including the one or more adjacent candidate modes, determine the final intra-prediction mode of the current block; and Predict the current block based on the final intra-frame prediction mode of the current block.