Prediction image generation device, video decoding device, video encoding device, and prediction image generation method
Through the adaptive reference area DIMD method and gradient derivation unit, different filters are selected, which solves the problem of insufficient intra-mode export accuracy in video encoding, and achieves higher encoding quality and lower computing complexity.
Patent Information
- Application Number
- CN202380085539.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-06
- Filing Date
- 2023-11-17
- Publication Date
- 2025-07-18
AI Technical Summary
In the existing video encoding technology, the decoder-side intra-mode export method has problems such as inconsistent prediction angles of adjacent blocks and fixed filter sizes, resulting in insufficient encoding accuracy.
Adaptive reference area DIMD method is used to switch the filter according to the size of the target block, and different filters are selected through the gradient derivation unit for brightness and chrominance prediction, thereby improving the accuracy of the intra prediction mode.
Without increasing the calculation amount, the encoding loss is reduced and the codec quality is improved, thereby improving the accuracy of the intra prediction mode.
Smart Images

Figure CN120345241A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to a predicted image generation device, a video decoding device, a video encoding device, and a predicted image generation method. Background Art
[0002] A video encoding device that generates encoded data by encoding a video and a video decoding device that generates a decoded image by decoding the encoded data are used for efficient transmission or recording of videos.
[0003] For example, specific video encoding schemes include H.264 / AVC, High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC) schemes, etc.
[0004] In such a video encoding scheme, the images (pictures) constituting the video are managed in a hierarchical structure, which includes slices obtained by dividing the image, coding tree units (CTUs) obtained by dividing the slices, coding units (coding units; will be referred to as CUs) obtained by dividing the coding tree units, and transform units (TUs) obtained by dividing the coding units, and encoding / decoding is performed for each CU.
[0005] In such a video encoding scheme, a predicted image is usually generated based on a locally decoded image obtained after encoding / decoding an input image (source image), and an encoding is performed on a prediction error component (which may also be referred to as a "differential image" or "residual image") obtained by subtracting the predicted image from the input image. Methods for generating a predicted image include inter-picture prediction (inter-frame prediction) and intra-picture prediction (intra-frame prediction).
[0006] In recent video encoding and decoding technologies, NPL1 proposed a decoder-side intra-mode derivation (DIMD) prediction method, in which the decoder derives an intra-angle prediction mode by performing luminance prediction using pixels in an adjacent region, thereby deriving a predicted image. In addition, NPL2 proposed an improvement to the DIMD method, in which different reference regions are used for prediction. In addition, NPL3 proposed a method for chrominance prediction using the DIMD method. Citation List Non-Patent Literature
[0007] NPL 1: M. Abdoli, T. Guionnet, E. Mora et al., "Non-CE3: Decoder-Side Intra-Mode Derivation and Prediction Fusion Using Planar", JVET-O0449, Gothenburg, July 2019. NPL 2: Z. Fan, Y. YASUGI, T. IKAI, "Non-EE2: Adaptive Reference Region DIMD", JVET-AB0065, Mainz, Germany, October 20 - 28, 2022 NPL 3: X. Li, R. Liao, J. Chen, Y. Ye, "Non-EE2: Regarding Chrominance Intra Prediction Mode", JVET-Y0092, Conference Call, January 12 - 21, 2022. Summary of the Invention Technical Problem
[0008] The pixel value gradients in different adjacent regions are used on the decoder side to derive the intra mode. However, in NPL2, there is a problem that the prediction angles of adjacent blocks may not be consistent with the prediction angle of the target block. And the sizes and values of the derivation filters used in NPL1 and NPL3 are fixed, and this is not the best choice for some cases.
[0009] The present invention aims to improve the accuracy by deriving different filters in the decoder-side intra mode derivation and switching the filter used for calculating the gradient according to the size of the target block. Problem Solution
[0010] The Adaptive Reference Region DIMD (NPL2) method allows the selection of a reference region with three modes: 1: Use the top-left, top, and left adjacent regions. 2: Use the top-left, top, and top-right adjacent regions. 3: Use the top-left, left, and bottom-left adjacent regions. They are named DIMD_TL, DIMD_T, and DIMD_L respectively. The DIMD (NPL1) and DIMD Chrominance (NPL3) methods derive the intra prediction direction by using the gradients of pixel values in different adjacent regions for luminance and chrominance prediction respectively.
[0011] The filter selection unit 3104611 is included in the gradient derivation unit 310461. The filter selection unit 3104611 selects a filter based on the size of the target block.
[0012] The filter selection unit 3104611 is characterized in that different filters are used for target blocks of different sizes. Advantageous Effects of the Invention
[0013] According to one aspect of the present invention, the computational amount for intra prediction mode derivation can be reduced while limiting the coding loss, and the quality of the codec can be improved without adding additional computations. Brief Description of the Drawings
[0014] Figure 1 Figure 1 It is a schematic diagram illustrating the configuration of the image transmission system according to this embodiment. Figure 2 Figure 2 It is a diagram showing the hierarchical structure of the encoded stream data. Figure 3 Figure 3 It is a schematic diagram showing the types (mode numbers) of the intra prediction modes. Figure 4 Figure 4 It is a schematic diagram of a video decoder. Figure 5 Figure 5 It is an example of the DIMD syntax. Figure 6 Figure 6 It is a binarized diagram of the dimd_mode syntax used in the DIMD prediction unit 31046. Figure 7 Figure 7 It is another example of the DIMD syntax. Figure 8 Figure 8 It shows the structure of the intra prediction image generation unit 310. Figure 9 Figure 9 It is a diagram showing the details of the DIMD prediction unit 31046. Figure 10 Figure 10 It shows an example of the reference region referred to by the DIMD prediction unit 31046. Figure 11 Figure 11 It shows an example of a set of spatial filters. Figure 12 Figure 12 It shows an example of the target pixels for gradient derivation when using a 3x3 filter. Figure 13 Figure 13 It shows an example of the target pixels for gradient derivation when using a 2x2 filter. Figure 14 Figure 14 It shows an example of the target pixels for gradient derivation when using a 2x2 filter. Figure 15 Figure 15 It is a block diagram illustrating the relationship between the gradient and the region. Figure 16 Figure 16 It is a block diagram showing the structure of the angle derivation unit. Figure 17 Figure 17 Shows an example of a reference region used in the gradient derivation of the DIMD prediction unit 31046. Figure 18 Figure 18 Is a block diagram showing an example of the configuration of the inverse quantization and inverse transformation processing unit 311. Figure 19 Figure 19 Is a block diagram showing the structure of the video coding system. Figure 20 Figure 20 Is a diagram showing the details of the DIMD chrominance prediction unit 31047. Figure 21 Figure 21 Shows an example of the reference region referred to by the DIMD prediction unit 31046. Figure 22 Figure 22 Shows some examples of filters for gradient derivation. Figure 23 Figure 23 Shows some examples of filters for gradient derivation. Figure 24 Figure 24 Shows an example of gradient-derived pixels for gradient derivation when using a filter. Figure 25 Figure 25 Shows an example of gradient-derived pixels for gradient derivation when using a filter. Figure 26 Figure 26 Shows an example of the reference region for gradient derivation of the DIMD prediction unit 31046. Figure 27 Figure 27 Is a block diagram showing the structure of the angle mode derivation device 310475. Figure 28 Figure 28 Shows an example of the reference region for gradient derivation of the DIMD chrominance prediction unit 31047. DETAILED DESCRIPTION First Embodiment
[0015] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings.
[0016] Figure 1 Is a schematic diagram illustrating the configuration of the image transmission system 1 according to this embodiment.
[0017] The image transmission system 1 is a system that transmits an encoded stream obtained by encoding a target image to be encoded, decodes the transmitted encoded stream, and displays the image. The image transmission system 1 includes a video encoding device (image encoding device) 11, a network 21, a video decoding device (image decoding device) 31, and a video display device (image display device) 41.
[0018] The image T is input to the video encoding device 11.
[0019] The network 21 sends the encoded stream Te generated by the video encoding device 11 to the video decoding device 31. The network 21 is the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. The network 21 is not necessarily limited to a two-way communication network and may be a one-way communication network configured to transmit broadcast waves such as digital terrestrial television broadcasts and satellite broadcasts. In addition, the network 21 may be replaced by a storage medium that records the encoded stream Te, such as a digital versatile disc (DVD: trademark) or a Blu-ray disc (BD: trademark).
[0020] The video decoding device 31 decodes each encoded stream Te transmitted from the network 21 and generates one or more decoded images Td that have been decoded.
[0021] The video display device 41 displays all or part of the one or more decoded images Td generated by the video decoding device 31. For example, the video display device 41 includes a display device such as a liquid crystal display and an organic electroluminescence (EL) display. The form of the display includes a fixed type, a mobile type, an HMD type, etc. In addition, when the video decoding device 31 has a high processing capacity, a high-definition image is displayed; and when the processing capacity of the video decoding device 31 is low, an image that does not require high processing capacity and display capacity is displayed. Operator
[0022] The operators and symbols used in this specification are as described below. >> is an arithmetic right shift, << is an arithmetic left shift, & is a bitwise AND, | is a bitwise OR, ^ is a bitwise exclusive OR, |= is an or assignment operator, and || indicates a logical OR.
[0023] x? y:z is a ternary operator that takes y when x is true (non-0) and takes z when x is false (0).
[0024] The Clip3(a, b, c) function is used to intercept the value c that is greater than or equal to a and less than or equal to b; returns a when c is less than a (c < a); returns b when c is greater than b (c > b); returns c in other cases (a is less than or equal to b (a <= b)).
[0025] The abs(a) function is used to return the absolute value of a.
[0026] The Int(a) function is used to return the integer value of a.
[0027] The floor(a) function is used to return the largest integer that is equal to or less than a.
[0028] The ceil(a) function is used to return the smallest integer that is equal to or greater than a.
[0029] a / d represents a divided by d (rounded down after the decimal point).
[0030] x = y..z means that x takes integer values from y to z, including y and z, where x, y, and z are integers and z is greater than or equal to y. Structure of the encoded stream Te
[0031] Before explaining the video encoding device 11 and the video decoding device 31 of the present embodiment in detail, the data structure of the encoded stream Te generated by the video encoding device 11 and decoded by the video decoding device 31 will be described.
[0032] Figure 4 is a diagram illustrating the data hierarchical structure of the encoded stream Te. The encoded stream Te illustratively includes a sequence and a plurality of pictures constituting the sequence. Figure 4 The (a) to (f) of are respectively diagrams illustrating an encoded video sequence defining a sequence SEQ, an encoded picture specifying a picture PICT, an encoded slice specifying a slice S, encoded slice data specifying the slice data, an encoded tree unit included in the encoded slice data, and an encoded unit (CU) included in each encoded tree unit. Encoded video sequence
[0033] In the encoded video sequence (CVS, encoded stream), a set of data is defined that is referred to by the video decoding device 31 for decoding the CVS to be processed. As Figure 2 shown, the CVS includes a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a picture (PICT), and supplementary enhancement information (SEI).
[0034] In the video parameter set VPS, in a video including multiple layers, a set of encoding parameters common to multiple videos and a set of encoding parameters associated with the multiple layers and individual layers included in the video are defined.
[0035] In the sequence parameter set SPS, a set of encoding parameters is defined that is referred to by the video decoding device 31 for decoding the target sequence. For example, the width and height of the picture are defined. Note that there may be multiple SPSs. In this case, any one of the multiple SPSs is selected from the PPS.
[0036] In the Picture Parameter Set (PPS), a set of coding parameters is defined that is referenced by the video decoding device 31 to decode each picture in the target sequence. For example, it includes a reference value for the quantization step size for picture decoding (pic_init_qp_minus26) and a flag indicating the application of weighted prediction (weighted_pred_flag). Note that there may be multiple PPSs. In this case, any one of the multiple PPSs is selected for each picture in the target sequence. Coded picture
[0037] In the coded picture, a set of data is defined that is referenced by the video decoding device 31 to decode the picture PICT to be processed. As Figure 2 shown, the picture PICT includes slices 0 to slice NS-1 (NS is the total number of slices included in the picture PICT).
[0038] Note that when it is not necessary to distinguish each of the following slices 0 to slice NS-1, the subscript of the reference symbol can be omitted. Additionally, this also applies to other data with subscripts included in the coded stream Te described below. Coded slice
[0039] In the coded slice, a set of data is defined that is referenced by the video decoding device 31 to decode the slice S to be processed. As Figure 2 shown, the slice includes a slice header and slice data.
[0040] The slice header includes a set of coding parameters that the video decoding device 31 references to determine the decoding method of the target slice. The slice type designation information (slice_type) indicating the slice type is an example of the coding parameters included in the slice header.
[0041] Examples of slice types that can be designated by the slice type designation information include (1) I slice that only uses intra prediction in coding, (2) P slice that uses unidirectional prediction or intra prediction in coding, and (3) B slice that uses unidirectional prediction, bidirectional prediction, or intra prediction in coding, etc. Note that inter prediction is not limited to unidirectional prediction and bidirectional prediction, and more reference pictures can also be used to generate the predicted image. Hereinafter, when a certain slice is referred to as a P slice or a B slice, the slice indicates a slice including blocks that can use inter prediction.
[0042] Note that the slice header may include a reference to the Picture Parameter Set (PPS) (pic_parameter_set_id). Coded slice data
[0043] In the coded slice data, a set of data is defined that is referenced by the video decoding device 31 to decode the slice data to be processed. The slice data includes asFigure 2 The CTU shown. A CTU is a block of a fixed size (e.g., 64x64) that constitutes a slice and may be referred to as the largest coding unit (LCU). Coding tree unit
[0044] In Figure 2 it, a set of data is defined that is referenced by the video decoding device 31 to decode the CTU to be processed. The CTU is divided into coding units CU, which are the basic units of coding processing, by recursive quadtree splitting (QT splitting), binary tree splitting (BT splitting), or ternary tree splitting (TT splitting). BT splitting and TT splitting are collectively referred to as multi-tree splitting (MT splitting). The nodes of the tree structure obtained by recursive quadtree splitting are called coding nodes. The intermediate nodes of the quadtree, binary tree, and ternary tree are coding nodes, and the CTU itself is also defined as the highest coding node. Coding unit
[0045] As Figure 2 shown, a set of data is defined that is referenced by the video decoding device 31 to decode the coding unit to be processed. Specifically, the CU includes a CU header CUH, prediction parameters, transform parameters, quantized transform coefficients, etc. In the CU header, prediction modes, etc. are defined.
[0046] Prediction processing is sometimes performed in units of CU and sometimes in units of sub-CUs obtained by further splitting the CU. When the size of the CU is equal to the size of the sub-CU, the number of sub-CUs in the CU is 1. When the size of the CU is larger than the size of the sub-CU, the CU is split into sub-CUs. For example, when the size of the CU is 8x8 and the size of the sub-CU is 4x4, the CU is split into four sub-CUs, including two horizontal splits and two vertical splits.
[0047] There are two types of prediction types (prediction modes): intra-frame prediction and inter-frame prediction. Intra-frame prediction refers to prediction within the same picture, and inter-frame prediction refers to prediction processing between different pictures (e.g., between pictures at different display times).
[0048] Transform and quantization processing are performed in units of CU, but the quantized transform coefficients can be entropy-coded in units of sub-blocks (such as 4x4). Prediction parameters
[0049] The predicted image is derived by the prediction parameters attached to the block. The prediction parameters include prediction parameters for intra-frame prediction and inter-frame prediction.
[0050] The prediction parameters for intra-frame prediction will be introduced below. The intra-frame prediction parameters include the luminance intra-frame prediction mode IntraPredModeY and the chrominance prediction mode IntraPredModeC. Figure 3It is a schematic diagram indicating the type (mode number) of intra prediction mode. As shown in the figure, for example, there are 67 types (0 to 66) of intra prediction modes. For example, there are planar prediction (0), DC prediction (1), and angular prediction (2 to 66). In addition, for chrominance, CCLM (Cross Component Linear Model) prediction modes (81 to 83), MMLM (Multi-Mode Linear Model) prediction modes, and LM (Linear Model) prediction modes can be added. Configuration of Video Decoding Device
[0051] The configuration of the video decoding device 31 ( Figure 4 ) according to the present embodiment will be described below.
[0052] The video decoding device 31 includes an entropy decoding unit 301, a parameter decoding unit (predicted image decoding device) 302, a loop filter 305, a reference picture memory 306, a prediction parameter memory 307, a predicted image generation unit 308, an inverse quantization and inverse transform processing unit 311, an addition unit 312, and a prediction parameter derivation unit 320. Note that, according to the video encoding device 11 described later, a configuration in which the loop filter 305 is not included in the video decoding device 31 is also used.
[0053] The parameter decoding unit 302 further includes a header decoding unit 3020, a CT information decoding unit 3021, and a CU decoding unit 3022 (prediction mode decoding unit). The CU decoding unit 3022 further includes a TU decoding unit 3024. These can be collectively referred to as decoding modules. The header decoding unit 3020 decodes parameter set information such as VPS, SPS, and PPS, and slice headers (slice information) from the encoded data. The CT information decoding unit 3021 decodes CT from the encoded data. The CU decoding unit 3022 decodes CU from the encoded data. When the TU includes prediction errors, the TU decoding unit 3024 decodes QP update information (quantization correction value) and quantized prediction errors (residual_coding) from the encoded data.
[0054] In addition, examples of using CTU and CU as processing units are described below, but the processing is not limited to this example, and processing can be performed in units of sub-CUs. Alternatively, by replacing CTU and CU with blocks, and sub-CU with sub-blocks, processing can be performed in units of blocks or sub-blocks.
[0055] The entropy decoding unit 301 performs entropy decoding on the encoded stream Te input from the outside, and separates and decodes each code (syntax element). The separated codes include prediction information for generating a predicted image, prediction errors for generating a differential image, and the like. Entropy coding has a variable-length coding method for syntax elements according to context (probability model) (the context is adaptively selected according to the type of syntax element and surrounding conditions), and a variable-length coding method for syntax elements using a predetermined table or formula.
[0056] In the entropy decoding unit 301, for the adaptive reference region DIMD(NPL2), there is a syntax element named dimd_mode. dimd_mode is a parameter for selecting the DIMD method reference region. dimd_mode includes the DIMD_MODE_TOP_LEFT mode, the DIMD_MODE_TOP mode, and the DIMD_MODE_LEFT mode. These three modes are represented by 0, 1, and 2, respectively.
[0057] Figure 6 An example of the binarization of dimd_mode is shown. In Figure 6 , binId is a variable indicating the bit position, and bin0 (binidx == 0) and Bin1 (binidx == 1) of the syntax element refer to the first bit and the next bit.
[0058] Bin0 is a flag for indicating whether to select the DIMD_MODE_TOP_LEFT. When Bin0 is 0, the DIMD_MODE_TOP_LEFT mode will be selected, and when Bin0 is 1, the DIMD_MODE_TOP_LEFT mode will not be selected.
[0059] Bin1 is a flag for indicating which one of the DIMD_MODE_TOP mode and the DIMD_MODE_LEFT mode to select. When Bin1 is 0, the DIMD_MODE_TOP mode will be selected, and when Bin1 is 1, the DIMD_MODE_LEFT mode will be selected.
[0060] In addition, it is worth noting that Bin0 and Bin1 do not form a single syntax element, but the syntax elements are assigned to Bin0 and Bin1 respectively. Therefore, dimd_mode can be seen through two syntax elements. Here, the syntax element assigned to Bin0 is named dimd_mode_flag, and the syntax element assigned to Bin1 is named dimd_mode_dir (as Figure 7As shown). In this case, the entropy decoding unit 301 can obtain dimd_mode according to dimd_mode_flag and dimd_mode_dir through the following formula. And when dimd_mode_flag is 0, dimd_mode_dir will also be set to 0.
[0061] dimd_mode = ((dimd_mode_flag == 0)? 0 : 1)+dimd_mode_dir In this example, DIMD_MODE_TOP_LEFT is represented by 1 bit (such as "0"), and 1 bit is allocated after "0" to represent DIMD_MODE_TOP and DIMD_MODE_LEFT. In the binarization of dimd_mode, for the DIMD_MODE_TOP_LEFT mode with a high selection rate, a shorter bit is used than the selection of the DIMD_MODE_TOP and DIMD_MODE_LEFT modes. Thus, the average coding amount can be shortened to improve the coding efficiency.
[0062] The parameter decoding unit 302 notifies the entropy decoding unit 301 of the syntax elements to be decoded. The entropy decoding unit 301 outputs the syntax elements to the prediction parameter derivation unit 320. Configuration of the prediction parameter derivation unit 320
[0063] The prediction parameter derivation unit 320 can derive prediction parameters based on the output of the parameter decoding unit 302 and the prediction parameters stored in the prediction parameter memory 307. The derived prediction parameters will be output to the prediction image generation unit 308 and stored in the prediction parameter memory 307. The prediction parameter derivation unit 320 can derive different prediction modes for luminance prediction and chrominance prediction.
[0064] The loop filter 305 is a filter set in the coding loop and is a filter that removes block distortion and ringing distortion and improves the image quality. The loop filter 305 applies filters such as a deblocking filter, sample adaptive offset (SAO), and adaptive loop filter (ALF) to the decoded image of the CU generated by the addition unit 312.
[0065] The reference picture memory 306 stores the decoded image of the CU generated by the addition unit 312 at a predetermined position for each target picture and target CU.
[0066] The prediction parameter memory 307 stores the prediction parameters at a predetermined position for each CTU or CU to be decoded. Specifically, the prediction parameter memory 307 stores the parameters derived by the prediction parameter derivation unit 320, the prediction mode predMode separated by the entropy decoding unit 301, etc.
[0067] The prediction image generation unit 308 receives the input of the prediction parameters derived by the prediction parameter derivation unit 320, etc. In addition, the prediction image generation unit 308 reads a reference picture from the reference picture memory 306. The prediction image generation unit 308 generates a prediction image of a block or a sub-block in the prediction mode indicated by the prediction mode predMode by using the prediction parameters and the read reference picture (reference picture block). Here, the reference picture block refers to a set of pixels on the reference picture (since they are usually rectangular, they are called blocks), and is the area used to generate the prediction image. Prediction image generation unit 318
[0068] When the prediction mode predMode indicates the intra prediction mode, the intra prediction image generation unit 310 performs intra prediction by using the intra prediction parameters (luminance intra prediction mode IntraPredModeY and / or chrominance intra prediction mode IntraPredModeC) input from the prediction parameter derivation unit 320 and the reference pixels read from the reference picture memory 306. When the prediction mode predMode indicates the inter prediction mode, the inter prediction image generation unit performs inter prediction by using the inter prediction parameters input from the prediction parameter derivation unit 320 and the reference pixels read from the reference picture memory 306.
[0069] Specifically, the prediction image generation unit 308 reads adjacent blocks within a predetermined range from the target block on the target picture from the reference picture memory 306. The predetermined range is the adjacent blocks on the left, upper left, upper, and upper right of the target block, and the reference area varies according to the intra prediction mode.
[0070] The prediction image generation unit 308 generates a prediction image of the target block by referring to the read decoded pixel values and the prediction mode indicated by predMode, IntraPredModeY, and / or IntraPredModeC. The prediction image generation unit 308 outputs the generated prediction image of the block to the addition unit 312.
[0071] The generation of the prediction image based on the intra prediction mode will be described below. In planar prediction, DC prediction, and angular prediction, the decoded peripheral area adjacent (close) to the prediction target block is configured as the reference area R. Then, the pixels on the reference area R are extrapolated in a specific direction to generate the prediction image. For example, the reference area R can be configured to include an L-shaped area on the left and above (or further, upper left, upper right, lower left) of the prediction target block. Intra prediction image generation unit 310
[0072] will useFigure 8 Describe the configuration of the intra-prediction image generation unit 310. The intra-prediction image generation unit 310 includes a reference sample filter unit 3103 (second reference image configuration unit), an intra-prediction unit 3104, and a prediction image corrector 3105 (prediction image corrector, filter switching unit, weight coefficient changing unit).
[0073] The intra-prediction unit 3104 generates a prediction image of the target block based on each reference pixel (unfiltered reference image) in the reference region R, the filtered reference image generated by applying a reference pixel filter (first filter), and the intra-prediction mode, and outputs it to the prediction image corrector 3105. The prediction image corrector 3105 corrects the prediction image according to the intra-prediction mode and outputs the corrected prediction image.
[0074] Each unit included in the intra-prediction image generation unit 310 will be described below. Reference sample filter unit 3103
[0075] The reference sample filter unit 3103 applies a reference pixel filter (first filter) to the unfiltered reference image according to the intra-prediction mode to derive the filtered reference image s[x][y] at each position (x, y) in the reference region R. Specifically, a low-pass filter is applied to the unfiltered reference image at and around each position (x, y), and the filtered reference image is derived. Note that the low-pass filter does not necessarily need to be applied to all intra-prediction modes, and the low-pass filter can be applied to some intra-prediction modes. Note that the filter applied to the unfiltered reference image in the reference region R in the reference sample filter unit 3103 is called the "reference pixel filter (first filter)", and the filter that corrects the prediction image in the prediction image corrector 3105 described below is called the "boundary filter (second filter)". Configuration of the intra-prediction unit 3104
[0076] The intra prediction unit 3104 generates a predicted image (predicted pixel values, uncorrected predicted image) of a target block to be predicted based on an intra prediction mode, an unfiltered reference image, and filtered reference pixel values, and outputs it to the predicted image corrector 3105. The intra prediction unit 3104 includes, inside thereof, a planar prediction unit 31041, a DC prediction unit 31042, an angular prediction unit 31043, an LM prediction unit 31044, a MIP prediction unit (matrix-based intra prediction) 31045, a DIMD (decoder-side intra mode derivation, DIMD) prediction unit 31046, and a DIMD chrominance prediction unit 31047. The intra prediction unit 3104 selects a specific predictor according to the intra prediction mode, and inputs the unfiltered reference image and the filtered reference image. The relationship between the intra prediction mode and the corresponding predictor is as follows. - Planar prediction... Planar prediction unit 31041 - DC prediction. DC prediction unit 31042 - Angular prediction... Angular prediction unit 31043 - LM prediction... LM prediction unit 31044 - MIP prediction... MIP prediction unit 31045 - DIMD prediction... DIMD prediction unit 31046 - DIMD chrominance prediction... DIMD chrominance prediction unit 31047 Planar prediction
[0077] The planar prediction unit 31041 generates a predicted image q[x][y] by linearly adding a plurality of filtered reference images s[x][y] according to the distance between the predicted pixel position and the reference pixel position, and outputs the generated image to the predicted image corrector 3105. DC prediction
[0078] The DC prediction unit 31042 derives a DC prediction value corresponding to the average value of the filtered reference image s[x][y], and outputs a predicted image q[x][y] having the DC prediction value as the pixel value. Angular prediction
[0079] The angular prediction unit 31043 generates a predicted image q[x][y] in a prediction direction (reference direction) indicated by the intra prediction mode using the filtered reference image s[x][y], and outputs the generated image to the predicted image corrector 3105. LM prediction
[0080] The LM prediction unit 31044 predicts the pixel values of chrominance based on the pixel values of luminance. More specifically, a linear model is used to generate a predicted chrominance image (Cb, Cr) based on the decoded luminance image. As an example of LM prediction, there is CCLM (Cross-Component Linear Model prediction) prediction. CCLM prediction is a prediction method that uses a linear model to predict the chrominance of the same block based on luminance. MIP Prediction
[0081] The MIP prediction unit 31045 generates a predicted image q[x][y] by performing a sum-of-products operation on the reference sample s[x][y] and the weight matrix derived from the adjacent region, and outputs the predicted image q[x][y] to the predicted image corrector 3105. DIMD Prediction
[0082] The DIMD prediction unit 31046 generates a luminance predicted image using the intra prediction mode of the target block. Information from the adjacent region is used to derive an intra prediction mode suitable for the target block, and the DIMD prediction unit 31046 uses the derived intra prediction mode to generate a predicted luminance image. Details will be described later. DIMD Chrominance Prediction
[0083] The DIMD chrominance prediction unit 31047 generates a chrominance predicted image using the intra prediction mode of the target block. Information from the adjacent region is used to derive an intra prediction mode suitable for the target block, and the DIMD chrominance prediction unit 31047 uses the derived intra prediction mode to generate a chrominance predicted image. Details will be described later. Configuration of the Predicted Image Corrector 3105
[0084] The predicted image corrector 3105 corrects the predicted image output from the intra prediction unit 3104 according to the intra prediction mode. Specifically, the predicted image corrector 3105 derives a predicted image (corrected predicted image) Pred by performing weighted addition (weighted average) on the unfiltered reference image and the predicted image for each pixel of the predicted image according to the distance between the reference region R and the target predicted pixel, where the predicted image is modified. Note that in some intra prediction modes (e.g., planar prediction, DC prediction, etc.), the predicted image corrector 3105 may not correct the predicted image, and the output of the intra prediction unit 3104 can be used as the predicted image. Application Example
[0085] Figure 9Shows the configuration of the DIMD prediction unit 31046 of the present embodiment. The DIMD prediction unit 31046 includes a reference sample derivation unit 310460, an angle mode derivation device 310465, a prediction mode selection unit 310463, and a predicted image generation unit 310464. The angle mode derivation device 310465 includes a gradient derivation unit 310461 and an angle mode derivation unit 310462. The gradient derivation unit 310461 includes a filter selection unit 3104611. The angle mode derivation device 310465 may include the prediction mode selection unit 310463.
[0086] Figure 5 Shows an example of encoded data regarding the DIMD method. For each target block, the prediction parameter derivation unit 320 decodes a flag named dimd_flag, which is used to indicate whether the DIMD method is used for the target block. When dimd_flag is 1, some syntax elements related to intra prediction (intra_mip_flag, intra_luma_mpm_flag, intra_luma_mpm_idx, intra_luma_mpm_reminder) are not decoded by the parameter decoding unit 302. intra_mip_flag indicates whether MIP prediction is used. intra_luma_mpm_flag is a flag used to indicate whether the most probable mode (MPM) is used. intra_luma_mpm_idx is an index used to indicate the MPM candidate when MPM is used. intra_luma_mpm_reminder is an index used to select an intra prediction mode from the remaining modes when MPM is not used. If dimd_flag is 0, intra_luma_mpm_flag is decoded, and if intra_luma_mpm_flag is 0, intra_luma_mpm_reminder is decoded. When the dimd_flag of the target block is 1, the dimd_mode_flag of the target block is decoded. The dimd_mode flag is a flag used to indicate the reference region for deriving the intra prediction mode. The dimd_mode flag can be set to the following values.
[0087] dimd_mode = 0 DIMD_MODE_TOP_LEFT (using the upper adjacent reference region and the left adjacent reference region) dimd_mode = 1 DIMD_MODE_LEFT (using the left adjacent reference region) dimd_mode = 2 DIMD_MODE_TOP (using the upper adjacent reference region) When the dimd_flag is 1, the DIMD prediction unit 31046 derives an angular mode indicating the texture direction in the adjacent region based on the pixel values. This angular mode will be used to generate an intra-predicted image. More specifically, (Step 1) pixel gradients are derived using the pixel values at a given position, (Step 2) the derived pixel gradients are converted into an angular prediction mode, (Step 3) a histogram is constructed using all the angular prediction modes obtained from Steps 1 and 2. (Step 4) One or more angular prediction modes are selected from the histogram and used to generate the predicted image. Figure 9 The DIMD prediction unit 31046 is shown, and each part of the DIMD prediction unit 31046 and the processing of each part are described in more detail below.
[0088] The processing of the DIMD prediction unit 31046 is as follows: Step 1: Pixel gradients are derived using the pixel values at a given position. Step 2: The derived pixel gradients are converted into an angular prediction mode. Step 3: A histogram is constructed using all the angular prediction modes obtained from Steps 1 and 2. Step 4: One or more angular prediction modes are selected from the histogram and used to generate the predicted image.
[0089] Figure 20 The configuration of the DIMD chrominance prediction unit 31047 of this embodiment is shown. The structure of the DIMD chrominance prediction unit 31047 is the same as that of the DIMD prediction unit 31046. The differences between DIMD chrominance and the DIMD method will be described later. The DIMD chrominance prediction unit 31047 applies the same process to the chrominance component of the target block. Reference sample derivation unit
[0090] The reference sample derivation unit 310460 derives a reference sample refUnit from the previously decoded pixels recSamples adjacent to the target block. Note that the reference sample derivation unit 310460 may be included in the reference sample filter unit 3103. Figure 10 An example of the reference region of the DIMD prediction unit 31046 is shown. The reference sample derivation unit 310460 stores the recSamples adjacent to the target block in the reference sample refUnit, and this refUnit is used by the gradient derivation unit 310461 and the prediction image generation unit 310. Example of reference region selection based on dimd_mode
[0091] When dimd_mode == DIMD_MODE_TOP_LEFT, the reference sample export unit 310460 exports the reference sample refUnit from the left and upper regions of the target block as described below.
[0092] First, the following process is performed at the position (x, y) within the left range (hereinafter referred to as RL) of the target block. refUnit[x][y] = recSamples[xC + x][yC + y] Here, RL has a range of (x = -1 - refIdxW... -1, y = 0... refH - 1). (xC, yC) are the upper left coordinates of the target block, and refIdxW is a constant indicating the width of the reference region on the left side of the target block. refIdxW can be 2, and refH is equal to the height of the target block (i.e., bH).
[0093] Second, the following process is performed in RT, where RT represents the upper range of the target block. refUnit[x][y] = recSamples[xC + x][yC + y] Here, RT has a range of (x = 0... refW - 1, y = -1 - refIdxH... -1). refIdxH is a constant indicating the height of the reference region above the target block. refIdxH can be 2, and refW is equal to the width of the target block (i.e., bW).
[0094] Third, the following process is performed to export the reference sample in the upper left range of the target block. refUnit[x][y] = recSamples[xC + x][yC + y] Here, x = -1 - refIdxW... -1, y = -1 - refIdxH... -1. (x and y cannot both be equal to -1) RTL is a range that includes all the above three ranges.
[0095] When dimd_mode == DIMD_MODE_LEFT, the reference sample export unit 310460 exports the left and lower left adjacent regions of the target block. refUnit[x][y] = recSamples[xC + x][yC + y] Here, x = -1 - refIdxW... -1, y = -1 - refIdxH... refH * 2 - 1. In this case, refIdxW and refIdxH are constants indicating the width and height of the reference area to the left of the target block, respectively. refIdxW and refIdxH can be equal to 3, and refH is equal to the height of the target block (i.e., bH).
[0096] When dimd_mode == DIMD_MODE_TOP, the reference sample derivation unit 310460 derives the adjacent areas above and to the upper right of the target block. refUnit[x][y] = recSamples[xC + x][yC + y] Here, x = -1 - refIdxW...refW * 2 - 1, y = -1 - refIdxH... -1. In this case, refIdxW and refIdxH are constants indicating the width and height of the reference area above the target block, respectively. refIdxW and refIdxH can be equal to 3, and refW is equal to the width of the target block (i.e., bW). It can also be written as follows.
[0097] In Figure 21 there are multiple reference ranges for the target block (e.g., 5): left, above, upper left, upper right, and lower left. If dimd_flag == 1, the reference sample derivation unit 310460 derives these multiple adjacent areas of the target block as follows.
[0098] First, the reference sample derivation unit 310460 derives the left adjacent area and the lower left adjacent area of the target block. The following processing is performed at the position (x, y) within the left range of the target block (hereinafter, the left range is referred to as RL (left area), and the lower left range is referred to as RLB (lower left area)). refUnit[x][y] = recSamples[xC + x][yC + y] Here, x = -refIdxW.. -1, y = 0..refH * 2 - 1. (xC, yC) is the upper left coordinate of the target block.
[0099] In this case, refIdxW and refH are constants indicating the width and height of the reference area to the left of the target block, respectively. refIdxW can be equal to 3, and refH is equal to the height of the target block (i.e., bH).
[0100] Second, the reference sample derivation unit 310460 derives the adjacent areas above and to the upper right of the target block. The following processing is performed in RT and RTR. RT (upper area) represents the upper range of the target block, and RTR (upper right area) represents the upper right range of the target block. refUnit[x][y] = recSamples[xC + x][yC + y] Here, x = 0..refW * 2 - 1, y = -refIdxH..-1. (xC, yC) is the upper-left coordinate of the target block.
[0101] In this case, refW and refIdxH are constants indicating the width and height of the reference area above the target block respectively. refIdxH can be equal to 3, and refW is equal to the width of the target block (i.e., bW).
[0102] Third, the following process is performed to derive the reference samples in the upper-left range of the target block. refUnit[x][y] = recSamples[xC + x][yC + y] Here, x = -refIdxW..-1, y = -refIdxH..-1. (x and y cannot be equal to -1 at the same time) (xC, yC) is the upper-left coordinate of the target block.
[0103] RTL (upper-left area) is an area including all the above three ranges. refIdxH and refIdxW can be equal to 3.
[0104] The above reference areas include RT, RL, RTL, RTR, and RLB, where refIdxH = 3, refIdxW = 3, which is called the normal reference range. In addition to the above reference area selection method, the reference area selection method can be set in the following two ways. The first is to reduce the span of the reference range and only select RT, RL, and RTL as the reference ranges (small reference range). The second is to increase the span of the reference range, select RT, RL, RTL, RTR, and RLB as the reference ranges, and increase the values of refIdxH and refIdxW, such as 4 (large reference range).
[0105] In this embodiment, the reference sample derivation unit 310460 can use the above different reference areas, that is, the small reference area, the normal reference area, and the large reference area. The reference sample derivation unit 310460 can use the regionType value to distinguish these areas, where regionType = 0 represents the small reference area, regionType = 1 represents the normal reference area, and regionType = 2 represents the large reference area. Selection of the reference range according to the size of the target block
[0106] Specifically, the reference sample derivation unit 310460 derives the blockType value by comparing one or more thresholds with the size of the target block. If the width and height of the target block are less than or equal to the corresponding thresholds (bW <= THW0 && bH <= THH0), the target block is a small block (blockType = 0); otherwise, if the width and height of the target block are less than or equal to another corresponding threshold (bW <= THW1 && bH <= THH1), the target block is a normal block (blockType = 1); otherwise, that is, the width and height of the target block are greater than another corresponding threshold (bW > THW1 || bH > THH1), the target block is a large block (blockType = 2). THW0 and THH0 can be 16, and THW1 and THH1 can be 64. It should be noted that the thresholds are not limited to the above thresholds and can be set in other ways. THW0 and THH0 may not be the same value, and THW1 and THH1 may not be the same value. For example, THW0 can be 8 and THH0 can be 16.
[0107] The reference sample derivation unit 310460 can derive the reference region range according to the size and / or shape of the target block. The reference region range can be a small reference range, a normal reference range, or a large reference range, which can be distinguished by the region type value (regionType). In one embodiment, the reference sample derivation unit 310460 derives a large reference range (regionType = 2) for a large target block (blockType is 2), a normal reference range (regionType = 1) for a normal target block (blockType is 1), and a small reference range (regionType = 0) for a small target block (blockType is 0). In another embodiment, the reference sample derivation unit 310460 uses a small reference range for a large target block, a normal reference range for a normal target block, and a large reference range for a small target block. Alternatively, the reference sample derivation unit 310460 uses the same reference range regardless of the size of the target range. The selection of the reference range is not limited to the above examples and can be in other ways. When a region is selected (deriving regionType), the gradient derivation unit 310461 uses the selected reference range to derive the direction. It should be noted that without introducing intermediate values (i.e., regionType and blockType), the reference range can be directly selected according to the width and height of the target block. Gradient derivation unit
[0108] The gradient derivation unit 310461 derives gradient values in two or more specific directions (e.g., Dx and Dy) based on the pixel values of adjacent blocks of the target block and corresponding spatial filters (e.g., the filter for Dx and the filter for Dy), and derives angle information representing the direction of the texture pattern according to the gradient values (Dx and Dy). The precision of the angle information can be set in other ways. For example, it can be set to about 1 / 36 degree units.
[0109] The spatial filter is used to derive the gradient. The spatial filter can be a 3x3 pixel filter or a 2x2 pixel filter corresponding to the horizontal direction and the vertical direction, as Figure 11 (a, b, e, f) shows. The gradient derivation unit 310461 derives the gradient value of the point P[x][y] (hereinafter simply referred to as P) in the reference sample refUnit[x][y] derived by the reference sample derivation unit 310460. The point recSamples[xC+x][yC+y] can be used instead of the point refUnit[x][y] as the point P.
[0110] Figure 12 and Figure 13 respectively show examples of the gradient target image in an 8x8 pixel target block of a 2x2 filter and a 3x3 filter. The pixels in the gradient target image are used to calculate the gradient. When the angle mode derivation unit 310462 is used for intra prediction, the meshed image in the adjacent region of the target block is the gradient target image of the target block (as Figure 12 and Figure 13 show). The gradient target image is the luminance image of the target block. The number of pixels used for gradient derivation to select the angle mode (modeVal) can be changed according to the size of the target block (bW, bH). The number of pixels used for gradient derivation is determined by the spatial filter. Therefore, the tap number (spatial filter size) of the spatial filter can be changed according to the size of the target block. The minimum number of reference rows in the required reference range varies according to the size of the filter. For example, a 2x2 filter requires at least 2 rows, while a 3x3 filter requires at least 3 rows. The number of reference rows obtained from the neighborhood of the current block must be greater than or equal to the minimum number of reference rows required by the selected filter. For rectangular filters such as 3x2 and 2x3, the minimum required reference rows in the horizontal and vertical directions are different. The specific number of reference rows can be determined based on the filter. In this case, the reference row count is defined as the minimum number of reference rows required by the filter, as Figure 12 (a), (b) show. Alternatively, the number of reference rows can be defined based on the DIMD mode. For example, when using the DIMD_MODE_TOP_LEFT mode, the row count can be set to "n" (where n == 3). Figure 13Illustrates the influence area of a 2x2 filter when the reference line count is set to 3. However, when DIMD is in DIMD_MODE_TOP or DIMD_MODE_LEFT mode, the reference line count is set to "n + 1" (4). Figure 14 (a), (b) illustrate the influence areas of 2x2 and 3x3 filters when the reference line count is set to 4. The value of "n" can be set in various ways (not only equal to 3), and there are multiple ways to configure the reference line counts for DIMD_MODE_TOP_LEFT, DIMD_MODE_TOP, and DIMD_MODE_LEFT modes.
[0111] The filter selection unit 3104611 determines the size of the filter (derives the parameter filterIdx) based on the size (bW, bH) of the target block. A variable filterIdx is defined to indicate the currently selected filter. The value of filterIdx is obtained according to the following equation. filterIdx = (bW <= TW1 && bH <= TH1)? 0 : 1 where TW1 = 16 and TH1 = 16. In other words, the filter selection unit 3104611 determines the size of the filter according to the size of the target block (e.g., whether to use a 2x2 filter or a 3x3 filter). For small target blocks with width and height equal to or less than 16, a smaller 2x2 filter is used, and for non-small (larger) target blocks with other sizes, a larger 3x3 filter is used. The gradient derivation unit 310461 derives the gradients Dx and Dy in the horizontal and vertical directions for each point P respectively according to the following formula based on filterIdx. Therefore, the gradient derivation unit 310461 derives Dx and Dy according to the size of the target block, where a larger filter is used for a larger block.
[0112] If filterIdx == 0 (for small target blocks), then the 2x2 filter is applied as follows Dx = P[x][y] + P[x][y + 1] - P[x + 1][y] - P[x + 1][y + 1] Dy = P[x][y] + P[x + 1][y] - P[x][y + 1] - P[x + 1][y + 1] If filterIdx == 1 (for other blocks), then the 3x3 filter is applied as follows Dx = P[x - 1][y + 1] + 2 * P[x][y + 1] + P[x + 1][y + 1] - P[x - 1][y - 1] - 2 * P[x][y - 1] - P[x + 1][y - 1] Dy = P[x - 1][y - 1] + 2 * P[x - 1][y] + P[x - 1][y + 1] - P[x + 1][y - 1] - 2 * P[x + 1][y] - P[x + 1][y + 1] Note that a 2x2 filter can use four positions: (x, y), (x + 1, y), (x, y + 1), (x + 1, y + 1), while a 3x3 filter can use nine positions: (x - 1, y - 1), (x, y - 1), (x + 1, y - 1), (x - 1, y), (x, y), (x + 1, y), (x - 1, y + 1), (x, y + 1), (x + 1, y + 1). A 4x4 filter can use 16 positions: (x - 1, y - 1), (x, y - 1), (x + 1, y - 1), (x + 2, y - 1), (x - 1, y), (x, y), (x + 1, y), (x + 2, y), (x - 1, y + 1), (x, y + 1), (x + 1, y + 1), (x + 2, y + 1), (x - 1, y + 2), (x, y + 2), (x + 1, y + 2), (x + 2, y + 2). Similarly, a 2x3 filter and a 3x2 filter can use 6 positions respectively: (x, y - 1), (x + 1, y - 1), (x, y), (x + 1, y), (x, y + 1), (x + 1, y + 1) and (x - 1, y), (x, y), (x + 1, y), (x - 1, y + 1), (x, y + 1), (x + 1, y + 1). Generally, an MxN filter can use (x + xx, y + yy), where xx = -(M - 1) / 2..M / 2, yy = -(N - 1) / 2..N / 2, and / represents integer division with the result truncated towards zero. The signs of each component in the above equation can be reversed, that is, Figure 11 The filter of (a, b, e, f) can be rotated 180 degrees to become Figure 11 the filter of (c, d, g, h), and these filters can be used to derive gradients. In this case, the following equations are used to derive Dx and Dy. If filterIdx == 0 (for small target blocks), then the 2x2 filter is applied as follows Dx = P[x + 1][y] + P[x + 1][y + 1] - P[x][y] - P[x][y + 1] Dy = P[x][y + 1] + P[x + 1][y + 1] - P[x][y] - P[x + 1][y] If filterIdx == 1 (for other blocks), then the 3x3 filter is applied as follows Dx = P[x - 1][y - 1] + 2 * P[x][y - 1] + P[x + 1][y - 1] - P[x - 1][y + 1] - 2 * P[x][y + 1] - P[x + 1][y + 1] Dy = P[x + 1][y - 1] + 2 * P[x + 1][y] + P[x + 1][y + 1] - P[x - 1][y - 1] - 2 * P[x - 1][y] - P[x - 1][y + 1]
[0113] The gradient derivation method is not limited to the above, and can also be configured in other ways. Examples of other thresholds The thresholds (TW1, TH1) can be set to be equal to (32, 32), (16, 32), (32, 16), etc. "<=" can be replaced by "<", as shown below filterIdx = (bW < TW1 && bH < TH1)? 0 : 1 The determination of the target block size can be based on bW + bH. filterIdx = (bW + bH <= TWH1)? 0 : 1 The threshold TWH1 can be set to be equal to 16, 32, 64, etc. The determination of the target block size can be based on bW * bH. filterIdx = (bW * bH <= TWH2)? 0 : 1 The threshold TWH2 can be set to be equal to 256, 512, 1024, etc. The determination of the target block size can be based on log2(bW) + log2(bH). filterIdx = (log2(bW) + log2(bH) <= TWH3)? 0 : 1 The threshold TWH3 can be equal to 8, 9, 10, etc. Note that the thresholds are not limited to the above values. 3 filters
[0114] The filter selection unit 3104611 can use two thresholds to select one filter from three types according to the size of the target block: large, medium, and small. The filter selection unit 3104611 selects 4x4, 3x3, and 2x2 filters for large, medium, and small target blocks respectively. The filter sizes can be set as follows. filterIdx = (bW < TW1 && bH < TH1)? 0 : (bW < TW2 && bH < TH2)? 1 : 2 where TW1, TH1, TW2, TH2 can be 16, 16, 32, 32 Rectangular filter
[0115] In addition to the square filter, a rectangular (non-square) filter can also be used. Figure 11 (i, j, m, l) show examples of two sets of rectangular filters. Taking the Figure 11 filter shown in (i, j) as an example. It is a rectangular filter with a width of 3 and a height of 2 (3x2 filter). The corresponding Dx and Dy can be calculated according to the following formulas. Dx = P[x - 1][y + 1] + 2 * P[x][y + 1] + P[x + 1][y + 1] - P[x - 1][y] - 2 * P[x][y] - P[x + 1][y] Dy = 2 * P[x - 1][y] + 2 * P[x - 1][y + 1] - 2 * P[x + 1][y - 1] - 2 * P[x + 1][y + 1] Similarly, Figure 11 the gradient derivation formula for the filter shown in (m, l) (2x3 filter) is as follows. Dx = 2 * P[x][y - 1] + 2 * P[x + 1][y - 1] - 2 * P[x][y + 1] - 2 * P[x + 1][y + 1] Dy = P[x][y - 1] + 2 * P[x][y] + P[x][y + 1] - P[x + 1][y - 1] - 2 * P[x + 1][y] - P[x + 1][y + 1] Rectangular filters are not limited to the above two types and can have other varieties. According to the block size, the filter selection unit 3104611 determines the size of the filter according to the size (bW, bH) of the target block. filterIdx = (bW <= 8 && bH <= 8)? 0 : ((bW <= 16 && bH <= 16)? (bW > bH? 3 : 4) : 1) The gradient derivation unit 310461 derives Dx and Dy according to the derived filterIdx. If filterIdx == 0 (for small target blocks, e.g., bW <= 8 && bH <= 8), a 2x2 filter is applied. Dx = P[x][y] + P[x][y + 1] - P[x + 1][y] - P[x + 1][y + 1] Dy = P[x][y] + P[x + 1][y] - P[x][y + 1] - P[x + 1][y + 1] Otherwise, if filterIdx == 3 (for rectangular blocks that are longer in the horizontal direction, bW > bH && bW <= 16 && bH <= 16), a 3x2 filter is applied. Dx = P[x - 1][y + 1] + 2 * P[x][y + 1] + P[x + 1][y + 1] - P[x - 1][y] - 2 * P[x][y] - P[x + 1][y] Dy = 2 * P[x - 1][y] + 2 * P[x - 1][y + 1] - 2 * P[x + 1][y - 1] - 2 * P[x + 1][y + 1] Otherwise, if filterIdx == 4 (for a vertically long rectangular block, bW < bH && bW <= 16 && bH <= 16), then apply a 2x3 filter. Dx = 2 * P[x][y - 1] + 2 * P[x + 1][y - 1] - 2 * P[x][y + 1] - 2 * P[x + 1][y + 1] Dy = P[x][y - 1] + 2 * P[x][y] + P[x][y + 1] - P[x + 1][y - 1] - 2 * P[x + 1][y] - P[x + 1][y + 1] Otherwise, if filterIdx == 1 (for other blocks), then apply a 3x3 filter. Dx = P[x - 1][y + 1] + 2 * P[x][y + 1] + P[x + 1][y + 1] - P[x - 1][y - 1] - 2 * P[x][y - 1] - P[x + 1][y - 1] Dy = P[x - 1][y - 1] + 2 * P[x - 1][y] + P[x - 1][y + 1] - P[x + 1][y - 1] - 2 * P[x + 1][y] - P[x + 1][y + 1]
[0116] The filter selection unit 3104611 can determine the size of the filter based on dimd_mode. If both the upper and left adjacent blocks are referenced, use a larger filter (e.g., 3x3), otherwise use a smaller filter (e.g., 2x2). filterIdx = (dimd_mode != DIMD_MODE_TOP_LEFT) ? 0 : 1 The gradient derivation unit 310461 derives Dx and Dy based on the derived filterIdx.
[0117] To reduce complexity, the filter selection unit 3104611 can determine the size of the filter based on dimd_mode, in other words, use smaller blocks (e.g., DIMD_MODE_TOP_LEFT) in the case of having more reference images. If both the upper and left adjacent blocks are referenced, use a larger filter (e.g., 2x2), otherwise use a smaller filter (e.g., 3x3). filterIdx = (dimd_mode == DIMD_MODE_TOP_LEFT)? 0 : 1 The gradient derivation unit 310461 derives Dx and Dy based on the derived filterIdx.
[0118] The filter selection unit 3104611 can determine the size of a filter including a rectangular kernel (e.g., 3x2, 2x3) based on dimd_mode. If both the upper and left adjacent blocks are referred to, a larger filter (e.g., 2x2) is used, otherwise a smaller filter (e.g., 3x2) is used for DIMD_MODE_LEFT, and another smaller filter (e.g., 2x3) is used for DIMD_MODE_TOP. filterIdx = (dimd_mode == DIMD_MODE_TOP_LEFT)? 0 : (dimd_mode == DIMD_MODE_LEFT)? 3 : 4. Additional spatial filter
[0119] Figure 22 and Figure 23 shows an example of a spatial filter. Figure 22 Shows 14 pairs of 3x3 filters. Each pair contains two filters, i.e., a Dx filter and a Dy filter corresponding to the horizontal and vertical gradients Dx and Dy. Each pair is assigned a number (filterType = #1..#14) or filterType. Figure 23 Shows 14 pairs of filters (filterType = #15..#28), including 2x2, 3x2, 4x4, and 2x3 filters. Note that the weights of the filters are not limited to these examples, and other filterType values can be used (assigned). A variable named "filterIdx" can be used to indicate the selected filter. The range of the filterIdx value is determined by the number of candidate filters. If the gradient derivation unit 310461 uses N sets of filters, filterIdx can be 0..N - 1. For example, in the case of selecting a filter based on the size of the target block, there may be three candidate filters with filterType #3, #2, and #1. In this case, filterIdx can be defined such that filterIdx == 0 represents filterType #3, filterIdx == 1 represents filterType #2, and filterIdx == 2 represents filterType #1.
[0120] The gradient derivation unit 310461 derives the gradient value of the point P[x][y] (hereinafter simply referred to as P) within the reference sample refUnit derived by the reference sample derivation unit 310460. The point recSamples[xC+x][yC+y] can be used instead of the point refUnit[x][y] as the point P. (xC, yC) are the upper left coordinates of the target block. The gradient derivation unit 310461 uses the pixels of adjacent images and uses different filters according to the target block to derive the gradient of the target block. The gradient derivation unit 310461 may include a filter selection unit 3104611, and the filter selection unit 3104611 may change the weights of the filters of the target block according to the target block. The filter selection unit 3104611 may derive the filter, and the gradient derivation unit 310461 may use the derived filter to derive the gradient. The gradient derivation unit 310461 may directly apply the selected filter without deriving the filterType value.
[0121] Figure 24 and Figure 25 respectively show examples of gradient target images of the filter in an 8x8 pixel target block. The pixels in the gradient target image are used to calculate the gradient of each gradient-derived pixel. The grid image of the adjacent region of the target block is the gradient target image of the target block. The gradient target image is a luminance image. The number of pixels used for gradient derivation to select the angle mode (modeVal) may vary according to the size of the target block (bW, bH). It is possible to change the number of taps of the spatial filter (spatial filter size) according to the size of the target block.
[0122] The filter selection unit 3104611 determines filters with different weights (a set of filters for Dx and Dy) based on the size (bW, bH) of the target block. It is possible to select Figure 22 and Figure 23 the candidate filters shown in. The selection can be made according to the size of the target block (for example, the sum of the width and height of the target block). The filter selection unit 3104611 can determine filters with different weights as follows. if(bW + bH <= a threshold){filterIdx = 0, use a weight values at certain positions, e.g. 3} else{filterIdx = 1, use a different weight values at certain positions, e.g. 2} The weight values at specific positions can be central elements or edge elements, where the central element is defined as Figure 22The positions (x, y-1), (x-1, y), (x+1, y), and (x, y+1) in the 3x3 filter shown. #1, #2, #3, #4, #5, and Figure 22 。#1, #10, #11, and Figure 22 。#1, #12, #13. The edge elements are defined as Figure 22 The positions (x-1, y-1), (x+1, y-1), (x-1, y+1), and (x+1, y+1) in the 3x3 filter shown. #1, #6, #7, #8, #9. The diagonal arrows indicate the center element, while the vertical arrows indicate the edge elements.
[0123] For example, the filter selection unit 3104611 can select a spatial filter such that the absolute weight value of the center element (also known as the center weight) decreases as the block size increases. if(bW + bH <= 40){filterIdx = 0, select filterType#3, center weight = 3} else if(bW + bH <= 64){filterIdx = 1, select filterType#2, center weight = 2} else{filterIdx = 2, select filterType#1, center weight = 1}
[0124] Again, the filter selection unit 3104611 can select a spatial filter such that the absolute weight value of the edge elements in the 3x3 filter increases as the block size increases. if(bW + bH <= 128){filterIdx = 2, select filterType#1, center weight = 1} else if(bW + bH <= 256){filterIdx = 1, select filterType#2, center weight = 2} else{filterIdx = 0, select filterType#3, center weight = 3}
[0125] Again, the filter selection unit 3104611 can select a spatial filter such that the weight of the center element in the 3x3 filter decreases to a certain size as the block size increases, and then increases as the block size increases.
[0126] The filter selection unit 3104611 can determine the filter or filterType as follows. if (bW + bH <= 40) {filterIdx = 0, select filterType#3, center weight = 3} else if (bW + bH <= 64) {filterIdx = 1, select filterType#2, center weight = 2} else if (bW + bH <= 128) {filterIdx = 2, select filterType#1, center weight = 1} else if (bW + bH <= 256) {filterIdx = 1, select filterType#2, center weight = 2} else {filterIdx = 0, select filterType#3, center weight = 3}
[0127] For another example, the filter selection unit 3104611 may select a spatial filter such that the weights of the edge elements in the 3x3 filter decrease with the increase of the block size to a certain size and then increase with the increase of the block size.
[0128] The filter selection unit 3104611 may determine the filter or filterType such that the weights of the edge elements in the 3x3 filter decrease with the increase of the block size. if (bW + bH <= 40) {filterIdx = 0, select filterType#7, side weight = 3} else if (bW + bH <= 64) {filterIdx = 1, select filterType#6, side weight = 2} else {filterIdx = 2, select filterType#1, side weight = 1}
[0129] The filter selection unit 3104611 may determine the filter or filterType such that the weights of the edge elements in the 3x3 filter decrease with the increase of the block size. if (bW + bH <= 40) {filterIdx = 0, select filterType#1, side weight = 1} else if(bW + bH <= 64){filterIdx = 1, select filterType#6, side weight = 2} else{filterIdx = 2, select filterType#7, side weight = 3}
[0130] These thresholds and filterType can be represented using two arrays, THArray[40, 64, 128, 256] and filterCandList[#3, #2, #1, #2, #3], and the process can be expressed in the following form. if(bW + bH <= THArray[0]){filterType = filterCandList[0]} else if(bW + bH <= THArray[1]){filterType = filterCandList[1]} else if(bW + bH <= THArray[2]){filterType = filterCandList[2]} else if(bW + bH <= THArray[3]){filterType = filterCandList[3]} else{filterType = filterCandList[4]} Or for i = 0..3 If bW + bH <= THArray[i], filterType is set equal to filterCandList[i] Otherwise (i == 4), filterType is set equal to filterCandList[i].
[0131] Note that the filter selection unit 3104611 can directly use the corresponding spatial filter and conditions without deriving intermediate values, namely THArray and filterCandList.
[0132] Different embodiments of the filter selection unit 3104611 are shown below. 1. As the size of the target block increases, the weight of the central element in the 3x3 filter first increases and then decreases. filterCandList can be [#1, #2, #3, #2, #1] or [#1, #3, #5, #3, #1]. 2. As the size of the target block increases, the weight of the central element in the 3x3 filter increases. filterCandList can be [#1, #2, #3, #4, #5]. 3. As the size of the target block increases, the weight of the central element in the 3x3 filter decreases. filterCandList can be [#5, #4, #3, #2, #1]. 4. As the size of the target block increases, the weight of the elements at the edges in the 3x3 filter first decreases and then increases. filterCandList can be [#7, #6, #1, #6, #7]. 5. As the size of the target block increases, the weight of the edge elements in the 3x3 filter increases. filterCandList can be [#1, #6, #7, #8, #9].
[0133] The thresholds in the above examples can be changed from [40, 64, 128, 256] to [32, 64, 128, 256] or other values. The number and values of the thresholds can be used. Filter Selection Method 2
[0134] In one embodiment, the filter selection unit 3104611 selects filters with different sizes (the size of the filter) (a set of filters for Dx and Dy), depending on the size of the target block. For example, for small target blocks, 2x2 filters are selected, for medium target blocks, 3x3 filters are selected, and for large target blocks, 4x4 filters are selected. THArray can be [64, 128], and filterCandList can be [#15, #1, #22].
[0135] In other embodiments, the filter selection unit 3104611 selects small filters for large blocks and large filters for small blocks. THArray can be set to [#64, #128], and filterCandList can be set to [#22, #1, #15]. The filter selection unit 3104611 can select filters based on the product of the block width and the block height. For example, bW x bH, where THArray can be [1024, 4096].
[0136] The filter selection unit 3104611 can select filters based on the sum of the logarithm of the block width and the block height, for example, log(bW) + log(bH), where THArray can be [#10, #12]. The filter selection unit 3104611 can select a filter according to the product of the logarithmic block width and the block height. For example, log(bW) x log(bH), where THArray can be [25, 36]. The filter selection unit 3104611 can select a filter according to the block width and the block height. For example, bW <= THArray[i] && bH <= THArray[i], where i = 0, 1 and THArray can be [32, 64]. The filter selection unit 3104611 can select a filter according to the logarithmic block width and the logarithmic block height. For example, log(bW) <= THArray[i] && log(bH) <= THArray[i], where i = 0, 1 and THArray can be [5, 6].
[0137] Note that MxM and NxN filters (M > N) can be represented by an MxM filter with zero values at specific positions (e.g., a 2x2 filter can be a 3x3 filter with five zeros in an L shape).
[0138] In one embodiment, the filter selection unit 3104611 selects filters with different weights (a set of filters for Dx and Dy), depending on the shape of the target block. if(bW == bH){filterIdx = 0, use a weight values at certain positions, e.g. 1} else{filterIdx = 1, use a weight values at certain positions, e.g. 2} The filter selection unit 3104611 can select a filter such that the weight of the central element can vary according to the block shape. if(bW == bH){filterIdx = 1, select filterType#1, center weight = 1} else{filterIdx = 2, select filterType#2, side weight = 2} The filter selection unit 3104611 can select a filter such that the weight of the edge elements can vary according to the block shape. if(bW == bH){filterIdx = 0, select filterType#1, side weight = 1} else{filterIdx = 1, select filterType#6, side weight = 2}
[0139] In another embodiment, the filter selection unit 3104611 selects filters (a set of filters for Dx and Dy) having different shapes (the shape of the filter) depending on the shape of the target block. if(bW == bH){filterIdx = 0, select square filter, e.g. filterType16} else{filterIdx = 1, select non - square filter, e.g. filterType#23}
[0140] For example, for a target block with width greater than height, a 3x2 filter is used; for a target block with height greater than width, a 2x3 filter is used; and for a target block with height equal to width, a 3x3 filter is used, as follows. if(bW > bH){filterIdx = 0, select 3x2 horizontal long filter, e.g. filterType16} else if(bW < bH){filterIdx = 1, select 2x3 vertical long filter, e.g. filterType#23} else{filterIdx = 2, select 3x3 square filter, e.g. filterType#1} In summary, the filter selection unit 3104611 can select filters as follows. if(bW > bH * K){filterType = filterCandList[0] or MxN horizontal long filter} else if(bW * K < bH){filterType = filterCandList[2] or NxM vertical long filter} else{filterType = filterCandList[1] or LxL square filter, K is a constant, 1, 2, 3…} where M, N (M > N) and L are constants (L is either M or N), for example, filterCandList[] is {#16, #1, #23}.
[0141] Note that MxN and NxN filters can be represented by LxL filters with zero values at specific positions (e.g., a 3x2 filter can be a 3x3 filter with 3x1 zeros, and a 2x3 filter can be a 3x3 filter with 1x3 zeros). Derivation of Dx and Dy
[0142] The gradient derivation unit 310461 derives the gradients Dx and Dy in the horizontal and vertical directions for each point P based on the filter according to the following formulas. Thus, the gradient derivation unit 310461 derives Dx and Dy according to the size or shape of the target block. The size or shape of the target block can be used to derive the weights of the filter. As described above, a variable named "filterIdx" can be used to indicate the selected filter being used. For example, in the case of selecting a filter based on the shape of the target block, there are three candidate filters with filterType #16, #1, and #23, where filterCandList[] is [#16, #1, #23]. In this case, filterIdx can be defined such that filterIdx == 0, 1, and 2 represent filterType #16, filterType #1, and filterType #23 respectively, where the range of filterIdx is 0..2.
[0143] Examples of the Dx and Dy formulas for filterType #15, #2, #17, #24 are as follows: For filterType #15, a 2x2 filter as shown below is applied Dx = P[x][y] + P[x][y + 1] - P[x + 1][y] - P[x + 1][y + 1] Dy = -P[x][y] - P[x + 1][y] + P[x][y + 1] + P[x + 1][y + 1] For filterType #2, a 3x3 filter as shown below is applied Dy = P[x - 1][y + 1] + 2 * P[x][y + 1] + P[x + 1][y + 1] - P[x - 1][y - 1] - 2 * P[x][y - 1] - P[x + 1][y - 1] Dx = P[x - 1][y - 1] + 2 * P[x - 1][y] + P[x - 1][y + 1] - P[x + 1][y - 1] - 2 * P[x + 1][y] - P[x + 1][y + 1] For filterType#17, apply the 3x2 filter as shown below Dy = P[x-1][y+1] + 2*P[x][y+1] + P[x+1][y+1] - P[x-1][y] - 2*P[x][y] - P[x+1][y] Dx = 2*P[x-1][y] + 2*P[x-1][y+1] - 2*P[x+1][y-1] - 2*P[x+1][y+1] For filterType#24, apply the 2x3 filter as shown below Dy = -2*P[x][y-1] - 2*P[x+1][y-1] + 2*P[x][y+1] + 2*P[x+1][y+1] Dx = P[x][y-1] + 2*P[x][y] + P[x][y+1] - P[x+1][y-1] - 2*P[x+1][y] - P[x+1][y+1] For filterType#22, apply the 4x4 filter as shown below Dy = P[x-2][y+1] + P[x-1][y+1] + P[x][y+1] + P[x+1][y+1] + P[x-2][y] + P[x-1][y] + P[x][y] + P[x+1][y] - P[x-2][y-2] - P[x-1][y-2] - P[x][y-2] - P[x+1][y-2] - P[x-2][y-1] - P[x-1][y-1] - P[x][y-1] - P[x+1][y-1] Dx = P[x-2][y-2] + P[x-2][y-1] + P[x-2][y] + P[x-2][y+1] + P[x-1][y-2] + P[x-1][y-1] + P[x-1][y] + P[x-1][y+1] - P[x][y-2] - P[x][y-1] - P[x][y] - P[x][y+1] - P[x+1][y-2] - P[x+1][y-1] - P[x+1][y] - P[x+1][y+1]
[0144] Note that a 2x2 filter can use four positions: (x,y), (x+1,y), (x,y+1), (x+1,y+1), while a 3x3 filter can use nine positions: (x-1,y-1), (x,y-1), (x+1,y-1), (x-1,y), (x,y), (x+1,y), (x-1,y+1), (x,y+1), (x+1,y+1). A 4x4 filter can use 16 positions: (x-2,y-2), (x-1,y-2), (x,y-2), (x+1,y-2), (x-2,y-1), (x-1,y-1), (x,y-1), (x+1,y-1), (x-2,y), (x-1,y), (x,y), (x+1,y), (x-2,y+1), (x-1,y+1), (x,y+1), (x+1,y+1). Similarly, a 2x3 filter and a 3x2 filter can use 6 positions respectively: (x,y-1), (x+1,y-1), (x,y), (x+1,y), (x,y+1), (x+1,y+1) and (x-1,y), (x,y), (x+1,y), (x-1,y+1), (x,y+1), (x+1,y+1). Generally, an MxN filter can use (x+xx,y+yy), where xx = -(M-1) / 2..M / 2 and yy = -(N-1) / 2..N / 2. The signs of each component in the above equations can be reversed, i.e., the filter can be rotated 180 degrees, which can also be used to derive gradients. In this case, the following equations are used to derive Dx and Dy. Taking the rotation of filterType#15 as an example, the following rotated filter is applied: Dx = -P[x+1][y] - P[x+1][y+1] + P[x][y] + P[x][y+1] Dy = -P[x][y+1] - P[x+1][y+1] + P[x][y] + P[x+1][y] Taking the rotation of filterType#2 as an example, the following rotated filter is applied: Dy = P[x-1][y-1] + 2*P[x][y-1] + P[x+1][y-1] - P[x-1][y+1] - 2*P[x][y+1] - P[x+1][y+1] Dx = P[x+1][y-1] + 2*P[x+1][y] + P[x+1][y+1] - P[x-1][y-1] - 2*P[x-1][y] - P[x-1][y+1] Range of gradient derivation
[0145] The (x, y) position range for histogram counting (for gradient derivation) is located in the reference regions RL, RT, RTL, RTR, and RLB. The gradient derivation ranges (x, y) corresponding to RL, RT, RTL, RTR, and RLB are referred to as RDL, RDT, RDTL, RDTR, and RDLB, as Figure 26 shown, Figure 26 which shows an example of a 3x3 filter.
[0146] Note that if the filter size applied is considered, the range of the reference region should be larger than the range of gradient derivation. For example, for a 3-tap filter accessing positions x - 1, x, x + 1 of x. Thus, the gradient derivation unit 310461 can change / extend the reference region according to the derived filter size. Alternatively, the gradient derivation unit 310461 can change / reduce the gradient derivation region according to the derived filter size.
[0147] For example, if the reference range is x = x0..x1 and y = y0..y1, the following is applied. (1) For a 3x3 filter, the range of gradient derivation can be x = x0 + 1..x1 - 1 and y = y0 + 1..y1 - 1 (by increasing the starting position by 1 and reducing the ending position by 1 with respect to the reference region). (2) For a 2x2 filter, the range of gradient derivation can be x = x0..x1 - 1 and y = y0..y1 - 1 (by reducing the ending position by 1 with respect to the reference region). (3) For a 4x4 filter, the range of gradient derivation can be x = x0 + 2..x1 - 1 and y = y0 + 2..y1 - 1 (by increasing the starting position by 2 and reducing the ending position by 1 with respect to the reference region). (4) For a 3x2 filter, the range of gradient derivation can be x = x0 + 1..x1 - 1 and y = y0..y1 - 1. (5) For a 2x3 filter, the range of gradient derivation can be x = x0..x1 - 1 and y = y0 + 1..y1 - 1.
[0148] Again, if the reference range for gradient derivation is x = x0..x1 and y = y0..y1, the following is applied. (1) For a 3x3 filter, the range of the reference region can be x = x0 - 1..x1 + 1 and y = y0 - 1..y1 + 1. (2) For a 2x2 filter, the range of the reference region can be x = x0..x1 + 1 and y = y0..y1 + 1. (3) For a 4x4 filter, the range of the reference region can be x = x0 - 2..x1 + 1 and y = y0 - 2..y1 + 1. (4) For a 3x2 filter, the range of the reference region can be x = x0 - 1..x1 + 1 and y = y0..y1 + 1. (5) For a 2x3 filter, the range of the reference region can be x = x0..x1 + 1 and y = y0 - 1..y1 + 1.
[0149] The gradient derivation unit 310461 derives angle information composed of the angular quadrant of the target block texture and the angle within the quadrant based on the relationship between Dx and Dy. By using the quadrant, some directions with rotational symmetry or line symmetry relationships can be uniformly processed. However, the angle information is not limited to the quadrant and the angle within the quadrant. For example, the angle information can be set only as an angle, and the quadrant can be derived from this angle as needed. Additionally, in this embodiment, the derived intra prediction direction mode is limited to the direction from bottom left to top right ( Figure 3 from 2 to 66 in Gradient derivation based on quadrant
[0150] Figure 15 (a) is a table showing the relationship between the signs (signx, signy) of Dx and Dy, the magnitude relationship xgty, and the quadrants (Ra to Rd). Figure 15 (b) shows the quadrants of Ra to Rd. The gradient derivation unit 310461 derives signx, signy, and xgty as follows. absx = abs(Dx) absy = abs(Dy) signx = Dx < 0? 1 : 0 signy = Dy < 0? 1 : 0 xgty = absx > absy? 1 : 0 Here, the inequality signs (> <) can be replaced with (>= <=). The angle information can be derived from signx, signy, and xgty.
[0151] The gradient derivation unit 310461 derives the quadrant from signx, signy, and xgty using the following operations or by looking up a table. The gradient derivation unit 310461 can refer to Figure 15 the table in (a) of quadrant = xgty? ((signx ^ signy)? 1 : 0) : ((signx ^ signy)? 2 : 3) The quadrants are represented by values from 0 to 3, and {Ra, Rb, Rc, Rd} = {0, 1, 2, 3}. The values of the quadrants are not limited to the above. Angle mode derivation device 310465
[0152] As described above, in this embodiment, an angle mode derivation device 310465 is used to derive an angle mode. It includes two parts: a gradient derivation unit 310461 and an angle mode derivation unit 310462. The gradient derivation unit 310461 is used to derive a gradient based on pixels, and the angle mode derivation unit 310462 is used to derive an angle based on the gradient. In addition, in this embodiment, the filters and reference ranges used in the gradient derivation unit 310461 can also be changed according to the size and shape of the target block. Angle mode derivation unit
[0153] The angle mode derivation unit 310462 derives an angle mode (a prediction mode corresponding to the target block, such as an intra prediction mode) based on the gradient information of the point P[x][y]. Figure 16 is a block diagram showing a configuration of the angle mode derivation unit 310462. As Figure 16 shown, a first gradient, a second gradient, and two tables are used to derive the angle mode (mode_delta or modeVal). Derivation of mode_delta
[0154] The angle mode derivation unit 310462 includes an angle coefficient derivation unit 310466 and a mode transformation unit 310467. The angle coefficient derivation unit 310466 derives an angle coefficient iRatio (or v) based on two gradients. Here, the gradient iRatio (= absy ÷ absx) is derived based on the absolute value of the first gradient (absx) and the absolute value of the second gradient (absy). iRatio can be approximately represented by a ratio and R_UNIT as follows iRatio = int(R_UNIT * absy / absx) ≒ ratio * R_UNIT R_UNIT is a power of 2 (1 << shiftR). For example, when shiftR = 16, R_UNIT = 65536.
[0155] The method of deriving iRatio is described below, but the derivation method is not limited to this example. Here, gradDivTable = {0, 7, 6, 5, 5, 4, 4, 3, 3, 3, 2, 1, 1, 1, 1, 0} and angTable = {0, 2048, 4096, 6144, 8192, 12288, 16384, 20480, 24576, 28672, 32768, 36864, 40960, 47104, 53248, 59392, 65536}. Additionally, in the expression, "|8" (the OR operation with 8) can be calculated as "+8". Similarly, "|16", "|32", and "|64" in the following descriptions can be calculated as "+16", "+32", and "+64" respectively.
[0156] x is the integer part of the logarithm of the third gradient value s1 (absx or absy) of the pixel. norm_s1 is derived by performing a shift operation on x using the third gradient s1. The angle coefficient v is determined by using norm_s1 and the reference table gradDivTable. Additionally, iRatio is derived by multiplying v and the fourth gradient s0 and performing a shift operation based on x. The angle mode mode_delta is derived using iRatio and angTable. The following clipping operation is performed to ensure that iRatio does not exceed the number range of angTable. iRatio = min((s0 * v) << 3 >> x, N_LUT - 1) Additionally, it is also suitable to clip and ensure that the value of s0 * v does not exceed a predetermined value KK. For example, not exceeding 32 bits. At this time, s0 * v = (min(s0 * v, KK) << 3) >> x, and KK = (1 << (31 - 3)) - 1 = 268435455. Note that iRatio can be derived by reversing the definitions of s0 and s1. Derivation of modeVal
[0157] The mode transformation unit 310467 uses mode_delta to derive and output the angle mode modeVal. modeVal = base_mode[quadrant] + direction[quadrant] * mode_delta where base_mode[4] = {HOR_IDX, HOR_IDX, VER_IDX, VER_IDX}, direction[4] = {-1, 1, -1, 1}. The number of occurrences of modeVal is recorded in the histogram (HistMode[]). The histogram can be calculated by adding 1 to the occurrence value corresponding to modeVal (this operation is referred to as "histogram counting" in the subsequent part of this document). HistMode[modeVal]+=1 Prediction mode selection unit
[0158] The prediction mode selection unit 310463 uses the histogram HistMode[] to derive one or more intra prediction modes dimdModeVal (dimdModeVal0, dimdModeVa1, dimdModeVal2, dimdModeVal3, dimdModeVal4). The histogram is derived from the modeVal (modeVal) calculated by a plurality of points P belonging to the gradient derivation target image. In the present embodiment, dimdModeVal is an estimated value of the dominant texture direction of the target block. dimdModeVal is derived by finding the value with the highest frequency (mode value) in the histogram. In the histogram, the first mode dimdModeVal0 is derived by selecting the mode with the highest frequency (the largest number of occurrences), and the second mode dimdModeVal1 is derived by selecting the mode with the second highest frequency (the second largest number of occurrences). More specifically, HistMode[x] is scanned to give the x with the maximum value in HistMode to determine dimdModeVal0 (= argmax(HistMode)), and the x with the second maximum value in HistMode is given to determine dimdModeVal1.
[0159] Here, cntMode can be 67. The method for deriving dimdModeVal0, dimdModeVal1, etc. is not limited to the histogram. For example, the prediction mode selection unit 310463 can set the average value of modeVal as dimdModeVar0 or dimdModeVal1.
[0160] The prediction mode selection unit 310463 sets a predetermined mode as the third mode dimdModeVal2. In the present embodiment, the third mode is set to the planar mode (0), but it is not limited to the planar mode. Other modes can be adaptively set as the third mode, or the third mode can be not used.
[0161] The angle mode selection unit 310463 can derive the weights of the above three modes used in the predicted image generation unit 310464. The total weight is set to 64, the weight of the third mode is assigned (W2 = 21), and the remaining weights are respectively assigned to weights W0 and W1 according to the frequency ratio of the first mode and the second mode in the histogram. The weighting of the first mode, the second mode, and the third mode is not limited to this, and the weights W0, W1, and W2 of the first mode, the second mode, and the third mode can be changed. For example, W1 can be increased or decreased. Note that setting the weight of the angle mode selection section to 0 means that the mode corresponding to the weight is not used.
[0162] Configurations of the adaptive gradient derivation unit 310461 and the angle mode derivation unit 310465 As described above, in this embodiment, the reference region of the reference image for deriving the intra prediction mode changes according to dimd_mode. Specifically, the position of the point P used in the gradient derivation unit 310461, the angle mode derivation unit 310462, and the angle mode selection unit 310463 changes according to dimd_mode.
[0163] The position range (x, y) for deriving the histogram count of the gradient and the angle mode is located in the reference regions RL, RT, and RTL. For the 3x3 filter, within the range of gradient derivation, the starting point is determined by adding 1, and the ending point is determined by subtracting 1. In the case of the 2x2 filter (filterIdx == 0), within the range of gradient derivation, the starting point remains unchanged, and the ending point is determined by subtracting 1. That is, in the case of the 3x3 filter (filterIdx == 1), if the reference range for dimd prediction is x = x0..x1, y = y0..y1, then the range of gradient derivation can be x = x0 + 1..x1 - 1, y = y0 + 1..y1 - 1. In the case of the 2x2 filter (filterIdx == 0), if the reference range for dimd prediction is x = x0..x1, y = y0..y1, then the range of gradient derivation can be x = x0..x1 - 1, y = y0..y1 - 1. The ranges (x, y) of gradient derivation corresponding to RL, RT, and RTL are referred to as RDL, RDT, and RDTL. Example of selecting the reference range based on dimd_mide
[0164] Figure 17 (b) shows an example of the reference range used in the gradient derivation process of dimd prediction. In this case, the 3x3 filter (filterIdx == 1) is used.
[0165] When dimd_mode == DIMD_MODE_TOP_LEFT, the angle mode export unit 310462 derives Dx and Dy from each point P in the left region RDL of the target block, derives modeVal, and performs a histogram counting operation. Subsequently, Dx and Dy are derived from each point P in the above-mentioned region RDT of the target block, and modeVal is derived and histogram counting is performed. The range of RDL is x = -refIdxW..-2, y = -refIedxH..refH-2. The range of RDT is x = -refIdxW..refW-2, y = -refIdxH..-2 RDTL is the range where RDL and RDT are combined.
[0166] When dimd_mode == DIMD_MODE_LEFT, the angle mode export unit 310462 uses the extended left region RDL_EXT of the target block. For example, Dx and Dy are derived from RDL_EXT for deriving and counting modeVal. The range of RDL_EXT is x = -refIdxW..-2, y = -refIedxH..refH*2-2. When dimd_mode == DIMD_MODE_TOP, the angle mode export unit 310462 derives modeVal from the extended upper region RDT_EXT of the target block and performs histogram counting. The range of RDT_EXT is x = -refIdxW..refW2-2, y = -refIdxH..-2
[0167] Here, refIdxW = 2, refIedxH = 2, refH = bH (height of the target block), refW = bW (width of the target block).
[0168] Similarly, for the case of using a 2x2 filter, the reference regions for gradient derivation are as Figure 16 (a) shown. In the case of using a 2x2 filter: The range of RDL is x = -refIdxW..-2, y = -refIedxH..refH-2 The range of RDT is x = -refIdxW..refW-2, y = -refIdxH..-2 RDTL is the range where RDL and RDT are combined. The range of RDL_EXT is x = -refIdxW..-2, y = -refIedxH..refH*2-2 The range of RDT_EXT is x = -refIdxW..refW2 - 2, y = -refIdxH..-2 Here, refIdxW = 2, refIdxH = 2, refH = bH (height of the target block), refW = bW (width of the target block).
[0169] The prediction mode selection unit 310463 selects an intra prediction mode from among multiple angular modes derived from pixels in the target image based on gradients, and thus, an angular mode with higher accuracy can be obtained. As described above, the prediction mode selection unit 310463 selects the angular mode estimated from the gradients and outputs the weights corresponding to each angular mode. Prediction image generation unit 310464
[0170] The prediction image generation unit 310464 generates a prediction image using two or three intra prediction modes, where one is set to the planar mode and the other one or two intra prediction modes are the angular modes input from the angular mode selection unit 310463. First, prediction images (pred0, pred1, pred2) are generated according to each intra prediction mode. Second, these prediction images are synthesized using the corresponding weights (w0, w1, w2) and output as the prediction image q[x][y]. The prediction image q[x][y] is derived as follows. q[x][y] = (w0 * pred0[x][y] + w1 * pred1[x][y] + w2 * pred2[x][y]) >> 6 However, if the frequency of the third mode is 0 or it is not a directional prediction mode (such as the DC mode (represented by the number 1), etc.), then the prediction image q[x][y] is generated from the first mode and the second mode, as shown below. q[x][y] = (w0 * pred0[x][y] + w1 * pred1[x][y]) >> 6 Here, the first mode and the second mode are angular modes, and the third mode is the planar mode. w1 is 21, the sum of w0 and w2 is 43, and if w2 is 0, then w0 is 43. It can also be written as follows. The prediction image generation unit 310464 uses the intra prediction modes derived in the angular mode derivation device 310465 to generate a prediction image. The prediction image generation unit 310464 can use the intra prediction derived in the angular mode derivation device 310465 to generate the prediction image q[x][y]. q[x][y] = w0 * pred0[x][y] The prediction image generation unit 310464 can generate a prediction image q[x][y] using two intra prediction modes exported in the angular mode derivation device 310465. q[x][y] = (w0 * pred0[x][y] + w1 * pred1[x][y] + 32) >> 6 The prediction image generation unit 310464 can generate a prediction image q[x][y] using one or more intra prediction modes exported in the angular mode derivation device 310465. q[x][y] = (w0 * pred0[x][y] + w1 * pred1[x][y] + wP * predP[x][y] + 32) >> 6 where pred0, pred1, and predP are the prediction images with IntraPredModeY being dimdModeVal0, the prediction image with IntraPredModeY being dimdModeVal1, and the prediction image with IntraPredModeY being the planar mode, respectively.
[0171] The prediction image generation unit 31046 can derive the weights for dimdModeVal0, dimdModeVal1, and the planar mode as w0, w1, and wP, respectively. The weights w0 and w1 are derived using the occurrence ratios of the first mode and the second mode in the histograms HistMode[dimdModeVal0] and HistMode[dimdModeVal1]. wP can be set to be equal to 21. The weighting of the first mode, the second mode, and the third mode is not limited to this. For example, w1 can be increased or decreased.
[0172] Generally, the prediction image generation unit 310464 can generate a prediction image using one or more intra prediction modes. Here, the number of intra prediction modes N can be from 2 to 6. One of them can be set to the planar mode, and the other intra prediction modes are the angular modes exported in the angular mode derivation device 310465. First, prediction images (pred0, pred1, pred2, pred3, pred4, predP) are generated according to each intra prediction mode. Second, these prediction images are synthesized using the corresponding weights (w0, w1, w2, w3, w4, wP) and output as the prediction image q[x][y]. The prediction image q[x][y] is derived as follows. q[x][y] = (w0 * pred0[x][y] + w1 * pred1[x][y] + w2 * pred2[x][y] + w3 * pred3[x][y] + w4 * pred4[x][y] + wP * predP[x][y] + 32) >> 6 Here, the sum of weights is 64. Among them, pred2, pred3, and pred4 are prediction images with IntraPredModeY being dimdModeVal2, dimdModeVal3, and dimeModeVal4 respectively.
[0173] If the occurrence count of the dimd mode is 0 (dimdModeVal1 == -1 or dimdModeVal2 == -1 or dimdModeVal3 == -1 or dimeModeVal4 == -1), then the prediction image q[x][y] is generated using less than 5 derived intra prediction modes. Specifically, if the occurrence count of the dimd mode is 1 or dimdModeVal1 == -1, then N = 2 is used. Similarly, if the occurrence count of the dimd mode is 2, 3, 4 or dimdModeVal2 == -1, dimdModeVal1 == -3, dimdModeVal1 == -4, then N = 3, 4, 5 is used.
[0174] The prediction image can be generated by 2, 3, 4, or 5 intra prediction modes, and the prediction image q[x][y] is derived as follows. q[x][y] = (w0 * pred0[x][y] + wP * predP[x][y] + 32) >> 6 q[x][y] = (w0 * pred0[x][y] + w1 * pred1[x][y] + wP * predP[x][y] + 32) >> 6 q[x][y] = (w0 * pred0[x][y] + w1 * pred1[x][y] + w2 * pred2[x][y] + wP * predP[x][y] + 32) >> 6 q[x][y] = (w0 * pred0[x][y] + w1 * pred1[x][y] + w2 * pred2[x][y] + w3 * pred3[x][y] + wP * predP[x][y] + 32) >> 6 Here, the last mode can be the planar mode, and the other modes can be angular modes.
[0175] The following is a detailed explanation of the DIMD chrominance prediction unit 31047.
[0176] The DIMD chroma prediction unit 31047 generates a chroma prediction image. The ChromaFusionFlag is used to indicate whether the current target block uses the DIMD chroma method for intra-chroma prediction. If the ChromaFusionFlag is 1, the DIMD chroma prediction unit 31047 derives an angular mode from the pixel values to indicate the texture direction in the adjacent region. This angular mode will be used to generate the chroma prediction image. Reference sample derivation unit
[0177] The reference sample derivation unit 310470 derives the reference sample refUnit from the previously decoded pixels recSamples adjacent to the target block. In one embodiment, the reference samples for DIMD chroma may include the luminance component as well as the chroma components Cb or Cr.
[0178] The reference sample derivation unit 310470 may refer to one or more reference ranges in chroma: for example, the left, upper, upper-left, upper-right, and lower-left ranges of the target block in the chroma (Cb or Cr) component. For the Cr region, these regions are referred to as RTL_cr, RT_cr, RL_cr, RTR_cr, and RLB_cr; for the Cb region, these regions are referred to as RTL_cb, RT_cb, RL_cb, RTR_cb, and RLB_cb. In addition, the reference sample derivation unit 310470 may refer to one or more luminance reference ranges in luminance: for example, the adjacent ranges (RTL, RT, RL, RTR, and RLB) of the target block and the left, upper, upper-left, upper-right, and lower-left ranges of the target block (RLTAR) in the luminance component. The method for deriving the left, upper, upper-left, upper-right, and lower-left samples of the target block is the same for DIMD chroma and the DIMD method. The specific method is described in the DIMD prediction and will not be repeated here.
[0179] The reference sample derivation unit 310470 may refer to the prediction image of the luminance target block, denoted as RLTAR (region of the luminance target block). Note that when generating the chroma prediction image, the prediction image of the luminance target block has already been generated, so the pixels in the luminance target block can also be used to generate the chroma prediction image. To obtain the pixels of RLTAR, the following process may be performed at the (x, y) position located in the upper-left of the target luminance block: refUnit[x / SubWidthC][y / SubHeightC] = recSamples[(xC + x) / SubWidthC][(yC + y) / SubHeight] Here, x = 0..bW-1, y = 0..bH-1, where bW and bH are the width and height of the target luminance block. (xC, yC) is the upper-left coordinate of the luminance target block. The variables SubWidthC and SubHeightC are variables depending on the chroma format sampling structure, where chroma_format_idc syntax can be decoded or encoded from the bitstream. If chroma_format_idc is 0 representing 4:2:0, then (SubWidthC, SubHeightC) = (2, 2). If chroma_format_idc is 2 representing 4:2:2, then (SubWidthC, SubHeightC) = (2, 1). If chroma_format_idc is 3 representing 4:4:4, then (SubWidthC, SubHeightC) = (1, 1).
[0180] Note that the reference regions used in the DIMD chroma method may vary. In Embodiments A, B, and C, the reference sample derivation unit 310470 may refer to only the chroma region, the chroma region plus the luminance region, and the chroma region, the luminance region, and the luminance prediction region, respectively. The embodiments are utilized based on the performance-complexity balance. Embodiment A can be used for low complexity cases, Embodiment B for medium complexity cases, and Embodiment C for high complexity and high performance cases.
[0181] The specific selection of the reference region to be used can be determined based on the target block. For example, as described in 310460, the reference range can be selected according to the size of the target block. Angle mode derivation device 310475
[0182] The angle mode derivation device 310475 is used to derive the angle mode. It includes two parts: the gradient derivation unit 310471 and the angle mode derivation unit 310472. The gradient derivation unit 310471 is used to derive the gradient according to the pixels, and the angle mode derivation unit 310472 is used to derive the angle according to the gradient. In addition, the filter and reference range used in the gradient derivation unit 310471 can also be changed according to the size and shape of the target block. Gradient derivation unit
[0183] The gradient derivation unit 310471 is the same as the gradient derivation unit 310461, but is applied to the chroma reference region, the luminance reference region, and the luminance prediction image (luminance target region). It derives the gradient values in two or more specific directions (e.g., Dx and Dy) for each region based on the pixel values of the target block region, and derives the angle information representing the texture pattern direction from the gradient values of each region. Angle mode derivation unit
[0184] The angle mode derivation unit 310472 derives an angle mode (modeVal) based on the gradient information of the point P[x][y]. Figure 27 FIG. is a block diagram showing one configuration of the angle mode derivation unit 310472. Since the angle mode derivation unit 310472 is the same as the angle mode derivation unit 310462, it will not be described again. Angle mode selection unit
[0185] The prediction mode selection unit 310473 derives one or more angle modes dimdChromaModeVal (dimdchromamodeVal0, dimdchromamodeVal1, etc.) using a histogram derived from modeVal calculated by a plurality of points P belonging to the gradient derivation target image. Since the prediction mode derivation unit 310473 is the same as the prediction mode selection unit 310463, it will not be described again.
[0186] For the reference luminance component, the position range (x, y) for histogram counting to derive the gradient is located in the reference regions RL, RT, RTL, RTR, RLB, and RLTAR. The gradient derivation ranges corresponding to RL, RT, RTL, RTR, RLB, and RLTAR are denoted as RDL, RDT, RDTL, RDTR, RDLB, and RDLTAR, as Figure 28 shown.
[0187] For the reference chrominance (Cb) component, the position range (x, y) for histogram counting to derive the gradient and the angle mode is located in the reference regions RL_cb, RT_cb, RTL_cb, RTR_cb, and RLB_cb. The gradient derivation ranges corresponding to RL_cb, RT_cb, RTL_cb, RTR_cb, and RLB_cb are denoted as RDL_cb, RDT_cb, RDTL_cb, RDTR_cb, and RDLB_cb, as Figure 28 shown.
[0188] The chrominance (Cr) component is the same as the chrominance (Cb) component, and the gradient derivation ranges (x, y) corresponding to RL_cr, RT_cr, RTL_cr, RTR_cr, and RLB_cr are denoted as RDL_cr, RDT_cr, RDTL_cr, RDTR_cr, and RDLB_cr, as Figure 28 shown.
[0189] The correspondence between the gradient derivation range and the filter used is explained in the gradient derivation range.
[0190] The prediction mode selection unit 310473 has the following function: selecting one angular mode from a plurality of angular modes derived from the pixels in the target image based on the gradient, so that an angular mode with higher accuracy can be obtained. As described above, the prediction mode selection unit 310473 selects the angular mode estimated from the gradient and outputs the weights corresponding to each angular mode. Prediction image generation unit 310474
[0191] The prediction image generation unit 310474 generates a chrominance prediction image using the chrominance intra prediction mode derived in the angular mode derivation device 310475. The prediction image generation unit 310474 can generate the chrominance prediction image q[x][y] using the chrominance intra prediction derived in the angular mode derivation device 310475. q[x][y] = w0 * pred0[x][y] The prediction image generation unit 310474 can generate the chrominance prediction image q[x][y] using two chrominance intra prediction modes derived in the angular mode derivation device 310475. q[x][y] = (w0 * pred0[x][y] + w1 * pred1[x][y] + 32) >> 6 The prediction image generation unit 310474 can generate the chrominance prediction image q[x][y] using one or more chrominance intra prediction modes derived in the angular mode derivation device 310475. q[x][y] = (w0 * pred0[x][y] + w1 * pred1[x][y] + wP * predP[x][y] + 32) >> 6 Wherein, pred0, pred1, and predP are the chrominance prediction images with IntraPredModeC being dimdModeVal0, the chrominance prediction images with IntraPredModeC being dimdModeVal1, and the chrominance prediction images with IntraPredModeC being the planar mode, respectively. The prediction image generation unit 310474 can generate the chrominance prediction image q[x][y] using one or more chrominance intra prediction modes derived in the angular mode derivation device 310475. q[x][y] = (w0 * pred0[x][y] + w1 * predLM[x][y] + 32) >> 6 Among them, predLM is the predicted image when IntraPredModeC is CCLM (Cross-Component Linear Model) or MMLM (Multi-Mode Linear Model). CCLM is a prediction mode in which the chrominance prediction image is generated by performing a linear operation with the luminance component values. For example, predC = (a * luma + b) >> shift. Here, predC is the predicted image value of the target block, and a, b, and shift are integer variables. It should be noted that the linear model parameters (a, b, and shift) are derived from the adjacent chrominance pixel values and luminance pixel values of the target block. MMLM is a prediction mode that utilizes a combination of the CCLM mode. Specifically, in the MMLM mode, two sets of model parameters (a0, b0, shift0) and (a1, b1, shift1) are derived, and the model is selected by the corresponding luminance value. For example, predC = (a * luma + b) >> shift, where if luma < thred, then (a, b, shift) = (a0, b0, shift0), otherwise (a, b, shift) = (a1, b1, shift1).
[0192] The inverse quantization and inverse transformation processing unit 311 performs inverse quantization on the quantized transform coefficients input from the prediction parameter derivation unit 320 to calculate the transform coefficients. The quantized transform coefficients are coefficients obtained by performing a frequency transformation such as discrete cosine transform (DCT), discrete sine transform (DST), etc. on the prediction error for quantization in the encoding process. The inverse quantization and inverse transformation processing unit 311 performs an inverse frequency transformation, such as inverse DCT, inverse DST, etc., on the calculated transform coefficients to calculate the prediction error. The inverse quantization and inverse transformation processing unit 311 outputs the prediction error to the addition unit 312.
[0193] Figure 18 is a block diagram showing the configuration of the inverse quantization and inverse transformation unit 311 of this embodiment. The inverse quantization and inverse transformation processing unit 311 includes a scaling unit 31111, an inverse non-separable transformer 31121, and an inverse separable transformer 31123. The inverse quantization and inverse transformation processing unit 311 uses the angle mode derived by the angle mode derivation device 310465 to transform the transform coefficients decoded from the encoded data.
[0194] The inverse quantization and inverse transformation processing unit 311 obtains the transform coefficients d[][] by scaling the quantized transform coefficients qd[][] input from the prediction parameter derivation unit 320 using the scaling unit 31111. The quantized transform coefficients qd[][] are the coefficients obtained through the quantization process. The quantization operation is performed by performing a transform operation such as DCT (Discrete Cosine Transform) or DST (Discrete Sine Transform) on the prediction error during the encoding process. In some cases, a non-separable transform may be performed to obtain the quantized transform coefficients qd[][]. When the non-separable transform flag Ifnst_idx!= 0, the inverse quantization and inverse transformation processing unit 311 performs an inverse non-separable transform through the inverse non-separable transformer 31121. In addition, the inverse quantization and inverse transformation processing unit 311 refers to the transform coefficients to perform inverse frequency transforms such as inverse DCT and inverse DST, and calculates the prediction error. If the non-separable transform flag Ifnst_idx!= 0, the inverse frequency transforms such as inverse DCT and inverse DST are directly calculated, and the prediction error is calculated without using the inverse non-separable transformer 31121. The inverse quantization and inverse transformation processing unit 311 outputs the prediction error to the addition unit 312.
[0195] The addition unit 312 adds the predicted image of the block input from the intra-frame predicted image generation unit 310 and the prediction error input from the inverse quantization and inverse transformation processing unit 311 for each pixel, and generates the decoded image of the block. The addition unit 312 stores the decoded image of the block in the reference picture memory 306 and outputs it to the loop filter 305. Configuration of the video encoding device
[0196] Next, the configuration of the video encoding device 11 according to the present embodiment will be described. Figure 19 is a block diagram illustrating the configuration of the video encoding device 11 according to the present embodiment. The video encoding device 11 is configured to include a predicted image generation unit 101, a subtraction unit 102, a transform and quantization processing unit 103, an inverse quantization and inverse transformation processing unit 105, an addition unit 106, a loop filter 107, a prediction parameter memory (prediction parameter storage unit, frame memory) 108, a reference picture memory (reference image storage unit, frame memory) 109, an encoding parameter determination unit 110, a parameter encoding unit 111, a prediction parameter derivation unit 120, and an entropy encoding unit 104.
[0197] The predicted image generation unit 101 generates a predicted image for each CU, which is a region obtained by dividing each picture of the image T. The operation of the predicted image generation unit 101 is the same as the operation of the intra-frame predicted image generation unit 310 described above, and thus the description thereof will be omitted.
[0198] The subtraction unit 102 subtracts the pixel values of the predicted image of the block input from the predicted image generation unit 101 from the pixel values of the image T to generate a prediction error. The subtraction unit 102 outputs the prediction error to the transform and quantization processing unit 103.
[0199] The transform and quantization processing unit 103 performs a frequency transform on the prediction error input from the subtraction unit 102 to calculate transform coefficients, and derives quantized transform coefficients through quantization. The transform and quantization processing unit 103 outputs the quantized transform coefficients to the entropy encoding unit 104 and the inverse quantization and inverse transform processing unit 105.
[0200] The inverse quantization and inverse transform processing unit 105 is the same as the inverse quantization and inverse transform processing unit 311 of the video decoding device 31 ( Figure 4 ), and its description is omitted. The calculated prediction error is output to the addition unit 106.
[0201] For the entropy encoding unit 104, the quantized transform coefficients are input from the transform and quantization processing unit 103, and the encoding parameters are input from the parameter encoding unit 111. The entropy encoding unit 104 performs entropy encoding on the segmentation information, prediction parameters, quantized transform coefficients, etc., to generate and output an encoded stream Te.
[0202] The parameter encoding unit 111 instructs the entropy encoding unit 104 to encode the prediction parameters and quantization coefficients derived from the prediction parameter derivation unit 120.
[0203] The prediction parameter derivation unit 120 derives syntax elements from the parameters input from the encoding parameter determination unit 110. Some parts of the prediction parameter derivation unit 120 have the same structure as the prediction parameter derivation unit 320.
[0204] The addition unit 106 adds the pixel values of the predicted image of the block input from the predicted image generation unit 101 and the prediction error input from the inverse quantization and inverse transform processing unit 105 for each pixel, and generates a decoded image. The addition unit 106 stores the generated decoded image in the reference picture memory 109.
[0205] The loop filter 107 applies a deblocking filter, SAO, and ALF to the decoded image generated by the addition unit 106. Note that the loop filter 107 does not necessarily include the above three types of filters, and for example, it can have a configuration with only a deblocking filter.
[0206] The prediction parameter memory 108 stores the prediction parameters generated by the prediction parameter derivation unit 120 for each target picture and CU at a predetermined position. It can store the transform coefficients created by the transform and quantization processing unit 103.
[0207] The reference picture memory 109 stores the decoded pictures generated by the loop filter 107 for each target picture and CU at predetermined positions.
[0208] The encoding parameter determination unit 110 selects one set from multiple sets of encoding parameters. The encoding parameters refer to the above-mentioned QT, BT, or TT segmentation information, prediction parameters, or parameters to be encoded, which are generated in association therewith. The prediction picture generation unit 101 uses these encoding parameters to generate a prediction picture.
[0209] The encoding parameter determination unit 110 calculates an RD cost value for each set of multiple sets of encoding parameters. The RD cost value indicates the amount of information and the magnitude of the encoding error. The RD cost value is, for example, the sum of the amount of code and the value obtained by multiplying the coefficient λ by the squared error. The encoding parameter determination unit 110 selects the set of encoding parameters for which the calculated cost value is the minimum. With this configuration, the entropy encoding unit 104 outputs the selected set of encoding parameters as the encoded stream Te. The encoding parameter determination unit 110 outputs the determined encoding parameters to the parameter encoding unit 111, the prediction parameter derivation unit 120, and the prediction picture generation unit 101.
[0210] Note that some parts of the video encoding device 11 and the video decoding device 31 in the above embodiments, for example, the entropy decoding unit 301, the parameter decoding unit 302, the loop filter 305, the intra prediction image generation unit 310, the inverse quantization and inverse transformation processing unit 311, the addition unit 312, the prediction parameter derivation unit 320, the prediction image generation unit 101, the subtraction unit 102, the transformation and quantization processing unit 103, the entropy encoding unit 104, the inverse quantization and inverse transformation processing unit 105, the loop filter 107, the encoding parameter determination unit 110, and the parameter encoding unit 111, and the prediction parameter derivation unit 120 can be implemented by a computer. In this case, this configuration can be implemented by recording a program for implementing such control functions on a computer-readable recording medium and causing a computer system to read the program recorded on the recording medium for execution. Note that the "computer system" mentioned here refers to the computer system built into the video encoding device 11 or the video decoding device 31 and is assumed to include an OS and hardware components such as peripheral devices. In addition, the "computer-readable recording medium" refers to portable media such as floppy disks, magneto-optical disks, ROMs, CD-ROMs, etc., and storage devices such as hard disks built into the computer system. In addition, the "computer-readable recording medium" may include a medium that dynamically stores a program for a short period of time, such as a communication line in the case where the program is transmitted through a network (such as the Internet) or through a communication line (such as a telephone line), and may also include a medium that stores a program for a fixed period of time, such as a volatile memory included in a computer system that serves as a server or a client in such a case. In addition, the above program may be a program for implementing some of the above functions or a program that can implement the above functions in combination with a program already recorded in the computer system.
[0211] In addition, part or all of the video encoding device 11 and the video decoding device 31 in the above embodiments can also be implemented as an integrated circuit such as a large-scale integrated circuit (LSI). Each functional block of the video encoding device 11 and the video decoding device 31 can be implemented as a processor individually, or part or all of them can be integrated into a processor. The circuit integration technology is not limited to LSI, and the integrated circuit for the functional block can be implemented as a dedicated circuit or a general-purpose processor. In the case where a circuit integration technology alternative to LSI appears with the progress of semiconductor technology, an integrated circuit based on that technology can be used.
[0212] The embodiments of the present disclosure have been described in detail above with reference to the drawings, but the specific configuration is not limited to the above embodiments, and various modifications can be made to the design without departing from the gist of the present disclosure.
[0213] Embodiments of the present invention can be applied to a video decoding device that decodes encoded data of image data and a video encoding device that generates encoded data from image data. Additionally, the data structure of the encoded data is generated by a video encoding device and referred to by a video decoding device. List of Reference Numerals 31 Image decoding device 301 Entropy decoding unit 302 Parameter decoding unit 310 Predicted image generation unit 3104 Intra prediction unit 31046 Dimd prediction unit 310460 Reference sample derivation unit 310465 Angle mode derivation device 310461 Gradient derivation unit 310462 Angle mode derivation unit 310463 Prediction mode selection unit 310464 Predicted image generation unit 31047 DIMD chrominance prediction unit 310470 Reference sample derivation unit 310475 Angle mode derivation device 310471 Gradient derivation unit 310472 Angle mode derivation unit 310473 Prediction mode selection unit 310474 Predicted image generation unit 311 Inverse quantization and inverse transform processing unit 312 Addition unit 11 Image encoding device 101 Predicted image generation unit 102 Subtraction unit 103 Transform and quantization processing unit 104 Entropy encoding unit 105 Inverse quantization and inverse transform processing unit 107 Loop filter 110 Encoding parameter determination unit 111 Parameter encoding unit
Claims
1. An image decoding device, the image decoding device having: a decoding circuit configured to decode dimd_mode; A reference sample derivation circuit configured to select adjacent images of a target block according to the dimd_mode; a gradient derivation circuit configured to derive a gradient according to pixels of the selected adjacent images and a filter selected according to the size of the block. And an angle mode selection circuit configured to derive an intra prediction mode according to the gradient.
2. The image decoding device according to claim 1, wherein, The gradient derivation circuit changes the size of the filter used according to the width and height of the target block.
3. The image decoding device according to claim 1, wherein The gradient derivation circuit determines a smaller filter for a small target block.
4. The image decoding device according to claim 1, wherein, The gradient derivation circuit changes the size of the filter used according to the derived dimd_mode.
5. The image decoding apparatus according to claim 1, wherein, In the case where the dimd_mode shows the use of both the upper and left adjacent blocks, the gradient derivation circuit determines a smaller filter.
6. An image coding apparatus having: a reference sample derivation circuit configured to select adjacent images of a target block according to the dimd_mode; a gradient derivation circuit for deriving a gradient according to pixels of the selected adjacent images and a filter selected according to the dimd_mode; and an angle mode selection circuit configured to derive an intra prediction mode according to the gradient.
7. An image decoding device, the image decoding device being configured to include: A gradient derivation circuit configured to use pixels of an adjacent image and use different filters according to a target block to derive the gradient of the target block; a prediction mode selection circuit configured to derive an intra prediction mode according to the gradient; and an intra prediction circuit that uses the derived intra prediction mode to derive a predicted image.
8. The image decoding device according to claim 1, wherein, The gradient derivation circuit selects a filter for the target block, and the gradient derivation circuit is configured to use the pixels of the adjacent image and the derived filter to derive the gradient.
9. The image decoding device according to claim 1, wherein, The gradient derivation circuit determines different weights of the central element in the 3x3 filter according to the size of the target block.
10. The image decoding device according to claim 3, wherein, The gradient derivation circuit determines the weight of the central element in the 3x3 filter such that the weight first decreases and then increases as the size of the target block increases.
11. The image decoding apparatus according to claim 1, wherein, The gradient derivation circuit changes the size of the filter according to the size of the target block.
12. The image decoding apparatus according to claim 1, wherein, The gradient derivation circuit changes the shape of the filter according to the size of the target block.
13. The image decoding apparatus according to claim 1, wherein, The gradient derivation circuit changes the shape of the filter according to the shape of the target block.
14. An image coding apparatus having: a reference sample derivation circuit configured to select a filter according to the size or shape of a target block; a gradient derivation circuit that derives a gradient according to pixels of an adjacent image and the selected filter; and an angle mode selection circuit that derives an intra prediction mode according to the gradient.