Moving picture decoding device and moving picture encoding device
By deriveing the reference image and the weight matrix in the moving image decoding and encoding device and generating a predicted image, the problems of large memory size and large processing volume in the prior art are solved, and more efficient intra prediction is achieved.
Patent Information
- Application Number
- CN202010979485.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-20
- Filing Date
- 2020-09-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-09-17
AI Technical Summary
In the existing matrix intra prediction technology, there are problems such as large memory size and huge processing volume of the weight matrix.
A moving image decoding device and encoding device are designed to derive a reference image that downsamples the adjacent images, and derive a matrix of weight coefficients in combination with intra prediction mode and object block size, thereby generating a predicted image and interpolation is performed to obtain the final predicted image.
While reducing the memory size of the weight matrix, the intra prediction process is optimized, the processing volume is reduced, and the encoding and decoding efficiency is improved.
Smart Images

Figure CN112532976B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to a moving picture decoding device and a moving picture encoding device. Background Art
[0002] In order to efficiently transmit or record moving images, a moving image encoding device that generates encoded data by encoding the moving image and a moving image decoding device that generates a decoded image by decoding the encoded data are used.
[0003] As specific moving picture coding methods, for example, methods proposed in H.264 / AVC and HEVC (High-Efficiency Video Coding) can be cited.
[0004] In this motion picture encoding method, the image (picture: Picture) constituting the motion picture is managed by a hierarchical structure including slices, coding tree units (CTU: Coding Tree Unit), coding units (sometimes also called coding units (Coding Unit: CU)) and transform units (TU: Transform Unit), and encoding / decoding is performed on each CU. The above-mentioned slices are obtained by dividing the image, the above-mentioned coding tree units are obtained by dividing the slices, the above-mentioned coding units are obtained by dividing the coding tree units, and the above-mentioned transform units are obtained by dividing the coding units.
[0005] In addition, in such a moving picture coding method, a prediction image is usually generated based on a local decoded image obtained by encoding / decoding an input image, and a prediction error (sometimes also referred to as a "difference image" or "residual image") obtained by subtracting the prediction image from the input image (original image) is encoded. Examples of methods for generating a prediction image include inter-picture prediction (inter-frame prediction) and intra-picture prediction (intra-frame prediction).
[0006] In addition, as a recent moving picture encoding and decoding technology, Non-Patent Document 1 can be cited. Non-Patent Document 2 discloses a matrix-based intra prediction technology (MIP) that derives a predicted image by performing a product-sum operation on a reference image derived from adjacent images and a weight matrix.
[0007] Prior art literature
[0008] Non-patent literature
[0009] Non-patent document 1: "Versatile Video Coding (Draft 6)", JVET-O2001-vE, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11
[0010] Non-patent document 2: “CE3: Affine linear weighted intra prediction (CE3-4.1, CE3-4.2)”, JVET-N0217-v1, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11 Summary of the invention
[0011] Problem that the invention aims to solve
[0012] In matrix intra prediction such as Non-Patent Documents 1 and 2, different weight matrices are stored according to various block sizes and intra prediction modes, so there is a problem that the memory size for storing the weight matrices is large. In addition, there is a problem that the amount of processing for generating a predicted image is huge.
[0013] An object of the present invention is to perform optimal intra-frame prediction and reduce the amount of processing while reducing the memory size of the weight matrix.
[0014] Technical Solution
[0015] In order to solve the above-mentioned problem, a motion image decoding device of one scheme of the present invention is characterized in that it comprises: a matrix reference pixel derivation unit, which derives an image obtained by downsampling the image adjacent to the upper and left sides of the object block as a reference image; a weight matrix derivation unit, which derives a matrix of weight coefficients according to the intra-frame prediction mode and the object block size; a matrix prediction image derivation unit, which derives a prediction image by the product of the elements of the above-mentioned reference image and the elements of the matrix of the above-mentioned weight coefficients; and a matrix prediction image interpolation unit, which derives the above-mentioned prediction image or an image obtained by interpolating the above-mentioned prediction image as a prediction image, and the above-mentioned weight matrix derivation unit derives a matrix with a size less than the width and less than the height of the object block size.
[0016] The moving picture decoding device is characterized in that the weight matrix derivation unit derives a matrix of a size of 4×4 when a side of the target block is 4.
[0017] The above-mentioned moving picture decoding device is characterized in that the above-mentioned weight matrix derivation unit derives a matrix of a size of 4×4 when the target block size is 4×16 or 16×4.
[0018] The above-mentioned motion image decoding device is characterized in that the above-mentioned weight matrix derivation unit derives either a matrix of sizeId=0, 1 of size 4×4 or a matrix of sizeId=2 of size 8×8, and derives a matrix of sizeId=1 or 2 when one side of the object block is 4.
[0019] Furthermore, the weight matrix derivation unit may derive a matrix of a size of 4×4 when the product of the width and the height of the target block size is 64 or less.
[0020] The above-mentioned moving picture decoding device is characterized in that the above-mentioned matrix prediction image derivation unit derives the intermediate prediction image predMip[][] of a square with equal width and height.
[0021] A motion image encoding device, characterized in that it comprises: a matrix reference pixel derivation unit, which derives an image obtained by downsampling images adjacent to the upper and left sides of an object block as a reference image; a weight matrix derivation unit, which derives a matrix of weight coefficients based on an intra-frame prediction mode and an object block size; a matrix prediction image derivation unit, which derives a prediction image by the product of elements of the reference image and elements of the matrix of weight coefficients; and a matrix prediction image interpolation unit, which derives the prediction image or an image obtained by interpolating the prediction image as a prediction image, the weight matrix derivation unit deriving a matrix having a size less than the width and less than the height of the object block size.
[0022] The above-mentioned moving picture coding device is characterized in that the above-mentioned weight matrix derivation unit derives a matrix of a size of 4×4 when a side of the target block is 4.
[0023] The above-mentioned moving picture coding device is characterized in that the above-mentioned weight matrix derivation unit derives a matrix of a size of 4×4 when the target block size is 4×16 or 16×4.
[0024] The above-mentioned motion image encoding device is characterized in that the above-mentioned weight matrix derivation unit derives either a matrix of sizeId=0, 1 of size 4×4 or a matrix of sizeId=2 of size 8×8, and derives a matrix of sizeId=1 or 2 when one side of the object block is 4.
[0025] Furthermore, the weight matrix derivation unit may derive a matrix of a size of 4×4 when the product of the width and the height of the target block size is 64 or less.
[0026] The above-mentioned moving picture encoding device is characterized in that the above-mentioned matrix prediction image derivation unit derives the intermediate prediction image predMip[][] of a square with equal width and height.
[0027] Beneficial Effects
[0028] According to one aspect of the present invention, it is possible to perform optimal intra-frame prediction while reducing the memory size of the weight matrix or reducing the amount of processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 This is a schematic diagram showing the configuration of the image transmission system according to the present embodiment.
[0030] Figure 2 This is a diagram showing the configuration of a transmitting device equipped with a moving picture encoding device according to the present embodiment and a receiving device equipped with a moving picture decoding device. PROD_A shows a transmitting device equipped with a moving picture encoding device, and PROD_B shows a receiving device equipped with a moving picture decoding device.
[0031] Figure 3 This is a diagram showing the configuration of a recording device equipped with a moving picture encoding device and a playback device equipped with a moving picture decoding device according to the present embodiment. PROD_C represents a recording device equipped with a moving picture encoding device, and PROD_D represents a playback device equipped with a moving picture decoding device.
[0032] Figure 4 This is a diagram showing the hierarchical structure of the data of the coded stream.
[0033] Figure 5 This is a diagram showing an example of CTU segmentation.
[0034] Figure 6 This is a schematic diagram showing the types (mode numbers) of intra prediction modes.
[0035] Figure 7 This is a schematic diagram showing the structure of a moving picture decoding device.
[0036] Figure 8 This is a schematic diagram showing the structure of an intra-frame prediction parameter decoding unit.
[0037] Fig. 9 A diagram showing a reference area used in intra prediction.
[0038] Fig.10 This is a diagram showing the structure of an intra-frame prediction image generation unit.
[0039] Fig.11 is a diagram showing an example of MIP processing.
[0040] Fig.12 is a diagram showing an example of MIP processing.
[0041] Fig.13 This is a block diagram showing the structure of a moving picture encoding device.
[0042] Fig.14 This is a schematic diagram showing the structure of an intra-frame prediction parameter encoding unit.
[0043] Fig.15 This is a diagram showing the details of the MIP unit.
[0044] Fig.16 It is a diagram showing the MIP processing according to this embodiment.
[0045] Fig.17 This is a diagram showing parameters for generating a predicted image when MIP is derived including a non-square predMip.
[0046] Fig.18 This is a diagram showing a method of deriving sizeId according to one embodiment (MIP Example 1) of the present invention.
[0047] Fig.19 This is a diagram showing parameters for generating a predicted image when a square predMip is derived through MIP.
[0048] Fig. 20 This is a diagram showing a method of deriving sizeId according to one embodiment (MIP Example 2) of the present invention.
[0049] Fig.21 This is a diagram showing a method of deriving sizeId according to one embodiment (MIP Example 3) of the present invention.
[0050] Fig. 22 This is a diagram showing a method of deriving sizeId according to one embodiment (MIP Example 4) of the present invention. DETAILED DESCRIPTION
[0051] (First Embodiment)
[0052] Hereinafter, embodiments of the present invention will be described with reference to the drawings.
[0053] Figure 1 It is a schematic diagram showing the configuration of the image transmission system 1 according to the present embodiment.
[0054] The image transmission system 1 is a system for transmitting a coded stream obtained by encoding a coding target image and decoding the transmitted coded stream to display the image. The image transmission system 1 is configured to include: a moving image encoding device (image encoding device) 11, a network 21, a moving image decoding device (image decoding device) 31, and a moving image display device (image display device) 41.
[0055] An image T is input to the video encoding device 11 .
[0056] The network 21 transmits the coded stream Te generated by the motion image coding device 11 to the motion image decoding device 31. The network 21 is the Internet, a wide area network (WAN), a local area network (LAN), or a combination of these. The network 21 is not necessarily limited to a two-way communication network, and may also be a one-way communication network that transmits broadcast waves such as terrestrial digital broadcasting and satellite broadcasting. In addition, the network 21 may also be replaced by a storage medium such as a DVD (Digital Versatile Disc: registered trademark) or a BD (Blue-ray Disc: registered trademark) that records the coded stream Te.
[0057] The moving picture decoding device 31 decodes each coded stream Te transmitted through the network 21 and generates one or more decoded pictures Td after decoding.
[0058] The moving image display device 41 displays all or part of one or more decoded images Td generated by the moving image decoding device 31. The moving image display device 41 includes, for example, a liquid crystal display, an organic EL (Electro-luminescence) display, and other display devices. Examples of display schemes include fixed type, mobile type, HMD (Helmet Mounted Display), and the like. In addition, when the moving image decoding device 31 has high processing power, a high-quality image is displayed, and when it has only low processing power, an image that does not require high processing power or display power is displayed.
[0059] <operator>
[0060] The operators used in this specification are described below.
[0061] >> is a right shift, << is a left shift, & is a bitwise AND, | is a bitwise OR, |= is an OR substitution operator, and || represents a logical OR.
[0062] x?y:z is a three-term operator that takes y when x is true (other than 0) and takes z when x is false (0).
[0063] Clip3(a, b, c) is a function that limits c to a value above a and below b. It returns a when c < a, b when c > b, and c in other cases (where a <= b).
[0064] Clip1Y(c) is an operator where a = 0 and b = (1 << BitDepthY) - 1 are set in Clip3(a, b, c). BitDepthY is the bit depth of luminance.
[0065] abs(a) is a function that returns the absolute value of a.
[0066] Int(a) is a function that returns the integer value of a.
[0067] floor(a) is a function that returns the largest integer less than or equal to a.
[0068] ceil(a) is a function that returns the smallest integer greater than or equal to a.
[0069] a / d means a divided by d (discarding the decimal part).
[0070] min(a, b) represents a function that returns the smaller value of a and b.
[0071] <Structure of the encoded stream Te>
[0072] Before detailing the moving image encoding device 11 and the moving image decoding device 31 of this embodiment, the data structure of the encoded stream Te generated by the moving image encoding device 11 and decoded by the moving image decoding device 31 will be described.
[0073] Figure 4 is a diagram showing the hierarchical structure of the data in the encoded stream Te. The encoded stream Te exemplarily includes a sequence and multiple pictures constituting the sequence. Figure 4 is a diagram showing the encoded video sequence of a specified sequence SEQ, the encoded picture of a specified picture PICT, the encoded slice of a specified slice S, the encoded slice data of the specified slice data, the encoded tree units included in the encoded slice data, and the encoded units included in the encoded tree units, respectively.
[0074] (Encoded video sequence)
[0075] In the encoded video sequence, a set of data referred to by the moving image decoding device 31 is specified for decoding the sequence SEQ to be processed. As Figure 4As shown in the coded video sequence, the sequence SEQ includes: a video parameter set (VideoParameter Set), a sequence parameter set SPS (Sequence Parameter Set), a picture parameter set PPS (PictureParameter Set), a picture PICT and supplemental enhancement information SEI (Supplemental EnhancementInformation).
[0076] The video parameter set VPS defines a set of coding parameters common to a plurality of moving images in a moving image composed of multiple layers, and a set of coding parameters associated with the multiple layers included in the moving image and with each layer.
[0077] In the sequence parameter set SPS, a set of coding parameters referenced by the motion picture decoding device 31 is specified for decoding the object sequence. For example, the width and height of the picture are specified. It should be noted that there can be multiple SPSs. In this case, any one of the multiple SPSs is selected from the PPS.
[0078] In the picture parameter set PPS, a set of coding parameters referenced by the motion picture decoding device 31 is specified for decoding each picture in the target sequence. For example, it includes a reference value (pic_init_qp_minus26) of the quantization width used in decoding the picture and a flag (weighted_pred_flag) indicating the application of weighted prediction. It should be noted that there can be multiple PPSs. In this case, any one of the multiple PPSs is selected from each picture in the target sequence.
[0079] (Encoded image)
[0080] In the coded picture, a set of data referenced by the moving image decoding device 31 is specified for decoding the picture PICT to be processed. Figure 4 As shown in the coded picture, the picture PICT includes slices 0 to slice NS-1 (NS is the total number of slices included in the picture PICT).
[0081] It should be noted that, in the following, when there is no need to distinguish between slices 0 to NS-1, the coding subscript may be omitted and described. The same applies to other data included in the coded stream Te described below and to which the subscript is attached.
[0082] (Encoding slice)
[0083] In the coded slice, a set of data referred to by the video decoding device 31 is specified for decoding the slice S to be processed. Figure 4The encoded slice is shown as follows, including a slice header and slice data.
[0084] The slice header includes a coding parameter group referred to by the moving picture decoding device 31 to determine a decoding method of a target slice. Slice type designation information (slice_type) that designates a slice type is one example of the coding parameter included in the slice header.
[0085] Examples of slice types that can be specified by the slice type specification information include: (1) I slices that use only intra-frame prediction during encoding, (2) P slices that use unidirectional prediction or intra-frame prediction during encoding, and (3) B slices that use unidirectional prediction, bidirectional prediction, or intra-frame prediction during encoding. It should be noted that inter-frame prediction is not limited to unidirectional prediction and bidirectional prediction, and more reference pictures can be used to generate predicted images. In the following, when referred to as P or B slices, it refers to slices that include blocks that can use inter-frame prediction.
[0086] It should be noted that the slice header may also include a reference to the picture parameter set PPS (pic_parameter_set_id).
[0087] (Slice Encoded Data)
[0088] The coded slice data specifies a set of data that the video decoding device 31 refers to in order to decode the slice data to be processed. Figure 4 As shown in the coded slice header of , the slice data includes CTU. CTU is a fixed-size (for example, 64×64) block constituting a slice, and is sometimes also called the largest coding unit (LCU).
[0089] (Coding Tree Unit)
[0090] exist Figure 4 In the coding tree unit, a set of data referenced by the motion image decoding device 31 is specified for decoding the CTU of the processing object. The CTU is divided into the basic unit of coding processing, namely the coding unit CU, through recursive quadtree partitioning (QT (Quad Tree) partitioning), binary tree partitioning (BT (Binary Tree) partitioning) or ternary tree partitioning (TT (Ternary Tree) partitioning). BT partitioning and TT partitioning are collectively referred to as multi-tree partitioning (MT (Multi Tree) partitioning). The nodes of the tree structure obtained by recursive quadtree partitioning are called coding nodes (Coding Node). The intermediate nodes of the quadtree, binary tree and ternary tree are coding nodes, and the CTU itself is also specified as the highest-order coding node.
[0091] CT includes, as CT information, a QT split flag (cu_split_flag) indicating whether QT splitting is performed, an MT split flag (split_mt_flag) indicating whether MT splitting is performed, an MT split direction (split_mt_dir) indicating the splitting direction of MT splitting, and an MT split type (split_mt_type) indicating the splitting type of MT splitting. cu_split_flag, split_mt_flag, split_mt_dir, and split_mt_type are transmitted for each coding node.
[0092] When cu_split_flag is 1, the coding node is split into four coding nodes ( Figure 5 QT).
[0093] When cu_split_flag is 0, in the case where split_mt_flag is 0, the coding node has one CU as a node without being split ( Figure 5 CU is the terminal node of the coding node and is not further divided. CU is the basic unit of coding processing.
[0094] When split_mt_flag is 1, the coding node is MT split as follows. When split_mt_type is 0, and split_mt_dir is 1, the coding node is horizontally split into two coding nodes ( Figure 5 BT (horizontal split)), when split_mt_dir is 0, the coding node is split vertically into two coding nodes ( Figure 5 BT (vertical split)). In addition, when split_mt_type is 1, when split_mt_dir is 1, the coding node is horizontally split into three coding nodes ( Figure 5 TT (horizontal split)), when split_mt_dir is 0, the coding node is split vertically into three coding nodes ( Figure 5 TT (vertical split)). Figure 5 The CT information is shown in
[0095] In addition, when the size of CTU is 64×64 pixels, the size of CU can be any one of 64×64 pixels, 64×32 pixels, 32×64 pixels, 32×32 pixels, 64×16 pixels, 16×64 pixels, 32×16 pixels, 16×32 pixels, 16×16 pixels, 64×8 pixels, 8×64 pixels, 32×8 pixels, 8×32 pixels, 16×8 pixels, 8×16 pixels, 8×8 pixels, 64×4 pixels, 4×64 pixels, 32×4 pixels, 4×32 pixels, 16×4 pixels, 4×16 pixels, 8×4 pixels, 4×8 pixels and 4×4 pixels.
[0096] (Coding unit)
[0097] like Figure 4 As shown in the coding unit, a set of data referenced by the motion image decoding device 31 is specified for decoding the coding unit to be processed. Specifically, the CU is composed of a CU header CUH, prediction parameters, transformation parameters, quantized transformation coefficients, etc. The prediction mode, etc. are specified in the CU header.
[0098] The prediction process can be performed in units of CU or in units of sub-CUs formed by further dividing the CU. When the size of the CU and the sub-CU are equal, there is one sub-CU in the CU. When the size of the CU is larger than the size of the sub-CU, the CU is divided into sub-CUs. For example, when the CU is 8×8 and the sub-CU is 4×4, the CU is divided into four sub-blocks, including two parts divided horizontally and two parts divided vertically.
[0099] There are two types of prediction (prediction modes): intra prediction and inter prediction. Intra prediction is prediction within the same picture, while inter prediction refers to prediction processing between different pictures (for example, between display times, between layer images).
[0100] The transform / quantization process is performed on a CU basis, but the quantized transform coefficients may be entropy encoded on a 4×4 or other sub-block basis.
[0101] (Prediction parameters)
[0102] The predicted image is derived from prediction parameters attached to the block. The prediction parameters include prediction parameters for intra-frame prediction and inter-frame prediction.
[0103] The following describes prediction parameters for intra prediction. The intra prediction parameters are composed of a luma prediction mode IntraPredModeY and a chroma prediction mode IntraPredModeC. Figure 6 is a schematic diagram showing the types (mode numbers) of intra-frame prediction modes. Figure 6As shown, there are 67 intra prediction modes (0 to 66), for example, planar prediction (0), DC prediction (1), and Angular prediction (2 to 66). Furthermore, LM mode (67 to 72) may be added to chrominance.
[0104] The syntax elements used to derive intra prediction parameters include, for example, intra_luma_mpm_flag, intra_luma_mpm_idx, intra_luma_mpm_remainder, etc.
[0105] (MPM)
[0106] intra_luma_mpm_flag is a flag indicating whether the IntraPredModeY of the object block is consistent with MPM (Most ProbableMode). MPM is a prediction mode included in the MPM candidate list mpmCandList[]. The MPM candidate list is a list that stores candidates with a high probability of being applied to the object block based on the intra prediction mode of the adjacent block and the specified intra prediction mode. When intra_luma_mpm_flag is 1, the IntraPredModeY of the object block is derived using the MPM candidate list and the index intra_luma_mpm_idx.
[0107] IntraPredModeY=mpmCandList[intra_luma_mpm_idx]
[0108] (REM)
[0109] When intra_luma_mpm_flag is 0, an intra prediction mode is selected from the mode RemIntraPredMode remaining after removing the intra prediction mode included in the MPM candidate list from all intra prediction modes. The intra prediction mode that can be selected as RemIntraPredMode is called "non-MPM" or "REM". RemIntraPredMode is derived using intra_luma_mpm_remainder.
[0110] (Configuration of Moving Image Coding Device)
[0111] The video decoding device 31 ( Figure 7 ) is described below.
[0112] The moving image decoding device 31 is configured to include an entropy decoding unit 301, a parameter decoding unit (prediction image decoding device) 302, a loop filter 305, a reference picture memory 306, a prediction parameter memory 307, a prediction image generation unit (prediction image generation device) 308, an inverse quantization / inverse transformation unit 311, and an addition unit 312. It should be noted that there is also a configuration in which the loop filter 305 is not included in the moving image decoding device 31 according to the moving image encoding device 11 described later.
[0113] The parameter decoding unit 302 includes an inter prediction parameter decoding unit 303 and an intra prediction parameter decoding unit 304 (not shown). The predicted image generation unit 308 includes an inter prediction image generation unit 309 and an intra prediction image generation unit 310 .
[0114] In addition, the following describes an example using CTU and CU as processing units, but the present invention is not limited to this example, and processing can also be performed in sub-CU units. Alternatively, CTU and CU can be replaced by blocks, and sub-CU can be replaced by sub-blocks, and processing can be performed in blocks or sub-blocks.
[0115] The entropy decoding unit 301 performs entropy decoding on the coded stream Te input from the outside, and separates each code (syntax element) for decoding. There are the following methods in entropy coding: a method of variable-length coding the syntax element using a context (probability model) adaptively selected according to the type of syntax element or the surrounding conditions; and a method of variable-length coding the syntax element using a predetermined table or calculation formula. In the former CABAC (Context Adaptive Binary Arithmetic Coding), the probability model updated for each picture (slice) after encoding or decoding is stored in the memory. Then, from the probability models stored in the memory, the probability model of the picture using the same slice type and the same slice level quantization parameter is set as the initial state of the context of the P picture or B picture. This initial state is used for encoding and decoding processing. The separated code includes prediction information for generating a predicted image and a prediction error for generating a differential image.
[0116] The entropy decoding unit 301 outputs the separated codes to the parameter decoding unit 302. Based on the instruction of the parameter decoding unit 302, it is controlled which code is to be decoded.
[0117] (Configuration of Intra-frame Prediction Parameter Coding Unit 304)
[0118] The intra prediction parameter decoding unit 304 decodes the intra prediction parameters, such as the intra prediction mode IntraPredMode, based on the code input from the entropy decoding unit 301, with reference to the prediction parameters stored in the prediction parameter memory 307. The intra prediction parameter decoding unit 304 outputs the decoded intra prediction parameters to the prediction image generation unit 308, and stores them in the prediction parameter memory 307. The intra prediction parameter decoding unit 304 may also derive different intra prediction modes for brightness and color difference.
[0119] Figure 8 3 is a schematic diagram showing the configuration of the intra-frame prediction parameter decoding unit 304 of the parameter decoding unit 302. Figure 8 As shown, the intra prediction parameter decoding unit 304 is configured to include: a parameter decoding control unit 3041, a luma intra prediction parameter decoding unit 3042, and a chroma intra prediction parameter decoding unit 3043.
[0120] The parameter decoding control unit 3041 instructs the entropy decoding unit 301 to decode the syntax elements, and receives the syntax elements from the entropy decoding unit 301. When intra_luma_mpm_flag is 1, the parameter decoding control unit 3041 outputs intra_luma_mpm_idx to the MPM parameter decoding unit 30422 in the luma intra prediction parameter decoding unit 3042. In addition, when intra_luma_mpm_flag is 0, the parameter decoding control unit 3041 outputs intra_luma_mpm_remainder to the non-MPM parameter decoding unit 30423 of the luma intra prediction parameter decoding unit 3042. In addition, the parameter decoding control unit 3041 outputs the syntax elements of the chroma intra prediction parameters to the chroma intra prediction parameter decoding unit 3043.
[0121] The luma intra prediction parameter decoding unit 3042 is configured to include an MPM candidate list derivation unit 30421 , an MPM parameter decoding unit 30422 , and a non-MPM parameter decoding unit 30423 (decoding unit, derivation unit).
[0122] The MPM parameter decoding unit 30422 refers to mpmCandList[] and intra_luma_mpm_idx derived by the MPM candidate list deriving unit 30421 , derives IntraPredModeY, and outputs it to the intra-frame prediction image generating unit 310 .
[0123] The non-MPM parameter decoding unit 30423 derives RemIntraPredMode from mpmCandList[] and intra_luma_mpm_remainder, and outputs IntraPredModeY to the intra-frame prediction image generation unit 310 .
[0124] The chroma intra prediction parameter decoding unit 3043 derives IntraPredModeC from the syntax elements of the chroma intra prediction parameters, and outputs it to the intra prediction image generation unit 310 .
[0125] The loop filter 305 is a filter provided in the encoding loop and removes block distortion and ringing distortion to improve image quality. The loop filter 305 performs filtering such as deblocking filtering, sample adaptive offset (SAO), and adaptive loop filtering (ALF) on the decoded image of the CU generated by the adding unit 312 .
[0126] The reference picture memory 306 stores the decoded image of the CU generated by the adding unit 312 at a position predetermined for each target picture and each target CU.
[0127] The prediction parameter memory 307 stores the prediction parameters at a predetermined position for each CTU or CU to be decoded. Specifically, the prediction parameter memory 307 stores the parameters decoded by the parameter decoding unit 302 and the prediction mode predMode separated by the entropy decoding unit 301 .
[0128] The prediction mode predMode, prediction parameters, etc. are input to the prediction image generation unit 308. In addition, the prediction image generation unit 308 reads the reference picture from the reference picture memory 306. The prediction image generation unit 308 generates a prediction image of a block or a sub-block using the prediction parameters and the read reference picture (reference picture block) in the prediction mode indicated by the prediction mode predMode. Here, the reference picture block refers to a set of pixels on the reference picture (usually rectangular, so it is called a block), which is a region referenced to generate a prediction image.
[0129] (Intra-frame prediction image generation unit 310)
[0130] When the prediction mode predMode indicates the intra prediction mode, the intra prediction image generation unit 310 performs intra prediction using the intra prediction parameters input from the intra prediction parameter decoding unit 304 and the reference pixels read from the reference picture memory 306 .
[0131] Specifically, the intra-frame prediction image generator 310 reads out adjacent blocks on the target picture that are within a predetermined range from the target block from the reference picture memory 306. The predetermined range refers to adjacent blocks to the left, upper left, upper, and upper right of the target block, and the reference area varies depending on the intra-frame prediction mode.
[0132] The intra-frame prediction image generation unit 310 generates a prediction image of the target block by referring to the read decoded pixel value and the prediction mode indicated by IntraPredMode. The intra-frame prediction image generation unit 310 outputs the generated block prediction image to the addition unit 312.
[0133] The following describes the generation of a predicted image based on an intra-frame prediction mode. In planar prediction, DC prediction, and angular prediction, a decoded surrounding area adjacent to (close to) the prediction target block is set as a reference area R. Then, the predicted image is generated by extrapolating pixels in the reference area R in a specific direction. For example, the reference area R can also be set to an L-shaped area (e.g., a region including the left and top (or further including the top left, top right, and bottom left) of the prediction target block. Fig. 9 The reference area of Example 1 is the area shown by the pixels marked with diagonal circles).
[0134] (Details of Prediction Image Generator)
[0135] Next, use Fig.10 The following describes the details of the configuration of the intra-frame prediction image generation unit 310. The intra-frame prediction image generation unit 310 includes a reference sampling filter unit 3103 (second reference image setting unit), a prediction unit 3104, and a prediction image correction unit 3105 (prediction image correction unit, filter switching unit, weight coefficient changing unit).
[0136] The prediction unit 3104 generates a temporary prediction image (pre-correction prediction image) of the prediction target block based on each reference pixel (reference image) in the reference region R, a filtered reference image generated by applying the reference pixel filter (first filter), and the intra prediction mode, and outputs it to the prediction image correction unit 3105. The prediction image correction unit 3105 corrects the temporary prediction image according to the intra prediction mode, generates a prediction image (corrected prediction image), and outputs it.
[0137] Hereinafter, each unit included in the intra-frame prediction image generation unit 310 will be described.
[0138] (See sampling filter unit 3103)
[0139] The reference sampling filter unit 3103 refers to the reference image to derive the reference samples s[x][y] at each position (x, y) on the reference region R. In addition, the reference sampling filter unit 3103 applies a reference pixel filter (first filter) to the reference samples s[x][y] according to the intra-frame prediction mode, and updates the reference samples s[x][y] at each position (x, y) on the reference region R (derives the filtered reference image s[x][y]). Specifically, a low-pass filter is applied to the reference image at the position (x, y) and its surroundings, and the filtered reference image ( Fig. 9 Example 2 of a reference area of the intra-frame prediction mode). It should be noted that it is not necessary to apply a low-pass filter to all intra-frame prediction modes, and a low-pass filter may be applied to a part of the intra-frame prediction modes. It should be noted that the filter applied to the reference image on the reference area R in the reference sampling filter unit 3103 is called a "reference pixel filter (first filter)", and in contrast, the filter for correcting the temporary predicted image in the predicted image correction unit 3105 described later is called a "position-dependent filter (second filter)".
[0140] (Configuration of Intra-frame Prediction Unit 3104)
[0141] The intra prediction unit 3104 generates a temporary prediction image (temporary prediction pixel value, prediction image before correction) of the prediction target block based on the intra prediction mode, the reference image, and the filtered reference pixel value, and outputs it to the prediction image correction unit 3105. The prediction unit 3104 internally includes: a Planar prediction unit 31041, a DC prediction unit 31042, an Angular prediction unit 31043, an LM prediction unit 31044, and a MIP unit 31045. The prediction unit 3104 selects a specific prediction unit according to the intra prediction mode, and inputs the reference image and the filtered reference image. The relationship between the intra prediction mode and the corresponding prediction unit is as follows.
[0142] ·Planar forecast···Planar forecast department 31041
[0143] ·DC Prediction···DC Prediction Department 31042
[0144] ·Angular prediction···Angular prediction department 31043
[0145] LM Forecast LM Forecast Department 31044
[0146] ·Matrix Intra Prediction···MIP Section 31045
[0147] (Planar forecast)
[0148] The planar prediction unit 31041 generates a temporary prediction image by linearly adding the reference samples s[x][y] according to the distance between the prediction target pixel position and the reference pixel position, and outputs the temporary prediction image to the prediction image correction unit 3105 .
[0149] (DC Prediction)
[0150] The DC prediction unit 31042 derives a DC prediction value equivalent to the average value of the reference samples s[x][y], and outputs a temporary prediction image q[x][y] having the DC prediction value as a pixel value.
[0151] (Angular prediction)
[0152] The Angular prediction unit 31043 generates a temporary prediction image q[x][y] using the reference sample s[x][y] in the prediction direction (reference direction) indicated by the intra prediction mode, and outputs it to the prediction image correction unit 3105 .
[0153] (LM forecast)
[0154] The LM prediction unit 31044 predicts the color difference pixel value based on the brightness pixel value. Specifically, it is a method of generating a predicted image of the color difference image (Cb, Cr) using a linear model based on the decoded brightness image. As one of the LM predictions, there is CCLM (Cross-Component Linear Model prediction). CCLM prediction is a prediction method that uses a linear model for predicting color difference based on brightness for a block.
[0155] (MIP Example 1)
[0156] Below, use Figure 11 to Figure 22 An example of MIP processing (Matrix-based intraprediction) performed by the MIP unit 31045 is described. MIP is a technique for deriving a predicted image by performing a product-sum operation on a reference image derived from an adjacent image and a weight matrix. In the figure, the width of the target block is nTbW and the height is nTbH.
[0157] (1) Boundary reference pixel extraction
[0158] The MIP unit uses the following formula to derive the variable sizeId related to the object block size: Fig.18 ).
[0159] sizeId=(nTbW<=4&&nTbH<=4)? 0: (nTbW<=8&&nTbH<=8)? 1:2(MIP-1)
[0160] like Fig.18 As shown, when the target block size (nTbWxnTbH) is 4×4, 8×8, and 16×16, sizeId is 0, 1, and 2, respectively. When it is 4×16 or 16×4, sizeId=2.
[0161] Next, the MIP unit 31045 uses the sizeId to derive the number of MIP modes used numModes, the size boundarySize of the downsampled reference regions redT[] and redL[], the width and height predW and predH of the intermediate prediction image predMip[][], and the size predC of one side of the prediction image obtained during the prediction process for the weight matrix mWeight[predC*predC][inSize].
[0162] numModes = (sizeId == 0)? 35 : (sizeId == 1)? 19 : 11 (MIP-2)
[0163] boundarySize = (sizeId == 0)? 2 : 4
[0164] predW = (sizeId <= 1)? 4 : Min(nTbW, 8)
[0165] predH = (sizeId <= 1)? 4 : Min(nTbH, 8)
[0166] predC = (sizeId <= 1)? 4 : 8
[0167] In Fig.17 shows the relationship between the sizeId and the values of these variables.
[0168] The weight matrix is square (predC*predC), 4×4 when sizeId = 0 or sizeId = 1, and 8×8 when sizeId = 2. When the size of the weight matrix is different from the output size predW*predH of the intermediate prediction image (especially when predC > predW and predC > predH), the weight matrix is referred to by interval rejection as described later. For example, in this embodiment, when the output size is 4×16 or 16×4, the weight matrix with size predC = 8 indicated by sizeId = 2 is selected, so cases where predW = 4 (< predC = 8) and predH = 4 (< predC = 8) are generated respectively. Since the size (predW*predH) of the intermediate prediction image needs to be less than or equal to the object block size nTbW*nTbH, when the object block size is small and a larger weight matrix (predC*predC) is selected, processing to make the weight matrix conform to the intermediate prediction image size is required.
[0169] In addition, the MIP unit 31045 uses IntraPredMode to derive the transpose processing flag isTransposed. IntraPredMode is, for example, Figure 6 Intra-prediction modes 0 to 66 are shown.
[0170] isTransposed=(IntraPredMode>(numModes / 2))? 1:0
[0171] In addition, the number of reference pixels inSize used in the prediction using the weight matrix mWeight[predC*predC][inSize] and the width and height mipW and mipH of the converted intermediate prediction image predMip[][] are derived.
[0172] inSize=2*boundarySize-((sizeId==2)?1:0)
[0173] mipW=isTransposed? predH:predW
[0174] mipH=isTransposed? predW:predH
[0175] The matrix reference pixel derivation unit of the MIP unit 31045 sets the pixel value predSamples[x][-1] (x=0..nTbW-1) of the block adjacent to the upper side of the object block in the first reference area refT[x] (x=0..nTbW-1). In addition, the pixel value predSamples[-1][y] (y=0..nTbH-1) of the block adjacent to the left side of the object block is set in the first reference area refL[y] (y=0..nTbH-1). Next, the MIP unit 31045 downsamples the first reference areas refT[x] and refL[y] to derive the second reference areas redT[x] (x=0..boundarySize-1) and redL[y] (y=0..boundarySize-1). During downsampling, refT[] and refL[] are processed in the same manner, and are therefore referred to as refS[i] (i=0..nTbX-1) and redS[i] (i=0..boundarySize-1) below.
[0176] The matrix reference pixel derivation unit performs the following processing on refT[] or refS[] substituted with refL[] to derive redS[]. When refT is substituted with refS, nTbS=nTbW, and when refL is substituted with refS, nTbS=nTbH.
[0177]
[0178] Here, Σ is the sum of i=0 to i=bDwn-1.
[0179] Next, the matrix reference pixel derivation unit combines the second reference areas redL[] and redT[] and derives p[i] (i=0..2*boundarySize-1).
[0180]
[0181]
[0182] bitDepthY is the bit depth of brightness, for example, it can be 10 bits.
[0183] It should be noted that, when the above reference pixels cannot be referenced, the available reference pixel values are used in the same way as the existing intra prediction. When all reference pixels cannot be referenced, 1<<(bitDepthY-1) is used as the pixel value. isTransposed indicates whether the prediction direction is close to the vertical prediction, so when switching whether to store redL or redT in the first half of p[] according to isTransposed, the pattern of mWeight[][] can be halved.
[0184] (2) Predicted pixel derivation (matrix operation)
[0185] MIP Department 31045 Fig.11 In STEP 2 predicted pixel derivation (matrix operation), the intermediate predicted image predMip[][] of size predW*predH is derived by matrix operation on p[].
[0186] The weight matrix derivation unit of the MIP unit 31045 selects a weight matrix mWeight[predC*predC][inSize] from the matrix set with reference to sizeId and modeId.
[0187] First, the weight matrix derivation unit derives modeId using IntraPredMode. modeId is the intra prediction mode used in MIP.
[0188] modeId=IntraPredMode-((isTransposed==1)?(numModes / 2):0)
[0189] When sizeId=0, the weight matrix derivation unit selects mWeight
[16] [4] from the array WeightS0
[18]
[16] [4] storing the weight matrix with reference to modeId. When sizeId=1, mWeight
[16] [8] is selected from the array WeightS1
[10]
[16] [8] storing the weight matrix with reference to modeId. When sizeId=2, mWeight
[64] [7] is selected from the array WeightS2[6]
[64] [7] storing the weight matrix with reference to modeId. These are expressed by the following formula.
[0190]
[0191] Next, the weight matrix derivation unit refers to sizeId and modeId to derive the shift value sW and offset coefficient fO used in (MIP-7). ShiftS0
[18] , ShiftS1
[10] , and ShiftS2[6] are arrays for storing shift values, and OffsetS0
[18] , OffsetS1
[10] , and OffsetS2[6] are arrays for storing offset coefficients.
[0192]
[0193] The matrix prediction image derivation unit of the MIP unit 31045 performs a product-sum operation (MIP-7) on p[], thereby deriving predMip[][] of the size of mipW*mipH. Here, the elements of the weight matrix mWeight[][] are referenced for each position corresponding to predMip[][] to derive the intermediate prediction image. It should be noted that in this embodiment, when sizeId=2, sometimes the size of the weight matrix predC is larger than the size of predMip mipW or mipH. Therefore, the variables incW and incH are used to eliminate the weight matrix for reference.
[0194]
[0195]
[0196] Σ is the sum of i=0 to i=inSize-1.
[0197] When isTransposed=1, the positions of the upper reference pixel and the left reference pixel are replaced and stored in the input p[] of the product-sum operation, and the transformation is performed before the output predMip[][] of the product-sum operation is output to (3).
[0198]
[0199] (3) Predicted pixel derivation (linear interpolation)
[0200] When nTbW=predW and nTbH=predH, the matrix prediction image interpolation unit of the MIP unit 31045 copies predMip[][] to predsamples[][].
[0201] for(x=0;x <nTbW;x++)
[0202] for(y=0;y <nTbH;y++)
[0203] predSamples[x][y]=predMip[x][y]
[0204] In other cases (nTbW>predW or nTbH>predH), the matrix prediction image interpolation unit is Fig.11 In 3-1 of STEP 3 (step 3) prediction pixel derivation (linear interpolation), predMip[][] is stored in the prediction image predSamples[][] of size nTbW*nTbH. When predW, predH are different from nTbW, nTbH, the prediction pixel value is interpolated in 3-2.
[0205] (3-1) The matrix prediction image interpolation unit stores predMip[][] in predSamples[][]. That is, Fig.12 In the pre-interpolation image, predMip[][] is stored at the shadow pixel position in the upper right and lower left directions.
[0206]
[0207] (3-2) When nTbH>nTbW, the pixels not stored in (3-1) are complemented using the pixel values of the adjacent blocks in the horizontal and vertical order to generate a predicted image.
[0208] To implement horizontal interpolation, use predSamples[xHor][yHor] and predSamples[xHor+upHor][yHor]( Fig.12 The pixel value at the position indicated by “○” is derived from the shadow pixels of the horizontally interpolated image.
[0209]
[0210] After interpolation in the horizontal direction, use predSamples[xVer][yVer] and predSamples[xVer][yVer+upVer]( Fig.12 The pixel value at the position indicated by “○” is derived from the shadow pixels of the vertically interpolated image.
[0211]
[0212]
[0213] When nTbH<=nTbW, the pixel values of the adjacent blocks are interpolated in the order of the vertical direction and the horizontal direction to generate a predicted image. The vertical and horizontal interpolation processes are the same as those in the case of nTbH>nTbW.
[0214] (MIP Example 2)
[0215] In this embodiment, an example is described in which the processing is simplified without reducing the coding efficiency compared with MIP Embodiment 1. The following description will focus on the changes, so the parts not described are the same processing as MIP Embodiment 1.
[0216] Fig.16 The structure of the MIP unit 31045 that represents a square matrix mWeight having a size not larger than the width nTbW and the height bTbH of the reference target block and derives a square intermediate prediction image predMip having the same size.
[0217] In this embodiment, when sizeId = 2, it is set to predW = predH = predC. The definition of sizeId is changed accordingly. Hereinafter, predW, predH, and predC are described as predSize.
[0218] (1) Boundary reference pixel extraction
[0219] The MIP unit uses the following formula to derive the variable sizeId related to the object block size: Fig. 20 ).
[0220] sizeId=(nTbW<=4&&nTbH<=4)? 0: ((nTbW<=4||nTbH<=4)||(nTbW==8&&nTbH==8))? 1:2(MIP-21)
[0221] For example, when the object block size is 4xN, Nx4 (N>4), or 8x8, sizeId is 1. If they are of the same category, the formula (MIP-21) may also be expressed in other ways, for example, as shown below.
[0222] sizeId=(nTbW<=4&&nTbH<=4)? 0: ((nTbW<=8&&nTbH<=8)||nTbW<=4||nTbH<=4)? 1:2(MIP-21)
[0223] As another example,
[0224] It may also be sizeId=(nTbW<=4&&nTbH<=4)? 0:((nTbW==8&&nTbH==8)||nTbW<=4||nTbH<=4)? 1:2(MIP-21). In addition, when the minimum size of the input block is 4×4, nTbW<=4 and nTbH<=4 may be replaced by nTbW==4 and nTbH==4, respectively.
[0225] Furthermore, when the block size to which the MIP is applied is limited, the MIP unit may also derive sizeId by other deriving methods. Fig. 20 As shown, when MIP is applied only to blocks whose aspect ratio of the input block size is 4 times or less (Abs(Log2(nTbW)-Log2(nTbH))<=2), sizeId may be derived as follows instead of (MIP-21).
[0226] sizeId=(nTbW<=4&&nTbH<=4)? 0: (nTbW*nTbH<=64)? 1:2(MIP-21a)
[0227] Alternatively, it can be derived as follows using a logarithmic expression.
[0228] sizeId=(nTbW<=4&&nTbH<=4)? 0: (Log2(nTbW)+Log2(nTbH)<=6)? 1:2(MIP-21b)
[0229] When the block size to which the MIP is applied is limited, sizeId is derived using (MIP-21a) and (MIP-21b), thereby simplifying the process.
[0230] like Fig. 20 As shown, in the present embodiment, in the case of 4×16 and 16×4, a matrix of size (predC) of 4 indicated by sizeId=1 is selected, so that the situation where predW and predH are smaller than the matrix size predC (=predSize) does not occur. The MIP unit 31045 of the present embodiment selects a matrix of size less than nTbW and bTbH (predC=predSize), that is, a matrix that satisfies the following formula.
[0231] predSize=predC<=min(nTbW,nTbH)
[0232] In this embodiment, the matrix size is 4×4 when sizeId=0 or 1, and 8×8 when sizeId=2. Therefore, the MIP unit 31045 selects "a matrix with sizeId=0 or sizeId=1 when one of nTbW and bTbH is 4". Such selection is limited to the following Fig.21 , Fig. 22 The same is true in Chinese.
[0233] That is, the weight matrix derivation unit of the MIP unit 31045 derives a matrix of a size less than the width and less than the height of the object block size. In addition, the weight matrix derivation unit derives a matrix of a size of 4×4 when one side of the object block is 4. In addition, the weight matrix derivation unit derives a matrix of a size of 4×4 when the size of the object block is 4×16 and 16×4. In addition, the weight matrix derivation unit derives any one of the matrices indicated by sizeId=0, 1 of a size of 4×4 and the matrix indicated by sizeId=2 of a size of 8×8, and derives a matrix of sizeId=1 or 2 when one side of the object block is 4.
[0234] Next, the MIP unit 31045 uses sizeId to derive the number of MIP modes numModes, the size boundarySize of the downsampled reference areas redT[], redL[], the weight matrix mWeight, and the width and height predSize of the intermediate prediction image predMip[][].
[0235] numModes=(sizeId==0)? 35: (sizeId==1)? 19:11(MIP-22)
[0236] boundarySize=(sizeId==0)? 2:4
[0237] predSize=(sizeId<=1)? 4:8
[0238] exist Fig.19 The relationship between sizeId and the values of these variables is shown in FIG.
[0239] The derivation of isTransposed and inSize is the same as that in MIP embodiment 1.
[0240] The derivation of p[] and pTemp[] required for the derivation of the first reference region refT[], refL[], the second reference region redT[], redL[], and predMip is also the same as that in the first MIP embodiment.
[0241] (2) Predicted pixel derivation (matrix operation)
[0242] MIP Department 31045 Fig.11 In STEP 2 predicted pixel derivation (matrix operation), predMip[][] of size predSize*predSize is derived through matrix operation on p[].
[0243] The weight matrix derivation unit of the MIP unit 31045 selects a weight matrix mWeight[predSize*predSize][inSize] from the set of matrices with reference to sizeId and modeId.
[0244] The selection method of modeId and mWeight[][], and the derivation method of the shift value sW and the offset coefficient fO are the same as those in MIP embodiment 1.
[0245] The matrix prediction image derivation unit of the MIP unit 31045 derives predMip[][] of size predSize*predSize by performing a product-sum operation (MIP-23) on p[]. Here, in the classification of sizeId in this embodiment, mipW and mipH are always greater than predSize(predC). Therefore, incW and incH of Embodiment 1 are always 1, and the calculation process is omitted.
[0246]
[0247] Σ is the sum of i=0 to i=inSize-1.
[0248]
[0249]
[0250] (3) Predicted pixel derivation (linear interpolation)
[0251] When nTbW=predSize and nTbH=predSize, the matrix prediction image interpolation unit of the MIP unit 31045 copies predMip[][] to predsamples[][].
[0252] for(x=0;x <nTbW;x++)
[0253] for(y=0;y <nTbH;y++)
[0254] predSamples[x][y]=predMip[x][y]
[0255] In other cases (nTbW>predSize or nTbH>predSize), the matrix prediction image interpolation unit is Fig.11 In STEP3 prediction pixel derivation (linear interpolation), enlarge predMip[][] of predSize*predSize to prediction image predSamples[][] of size nTbW*nTbH. Copy the pixels at corresponding positions in 3-1, and derive the pixels at non-corresponding positions by interpolation in 3-2.
[0256] (3-1) The matrix prediction image interpolation unit stores predMip[][] in the corresponding position of predSamples[][]. That is, Fig.12 In the pre-interpolation image, predMip[][] is stored at the shadow pixel position of predSamples[][] at 3-1.
[0257]
[0258] (3-2) When nTbH>nTbW, the pixels not stored in (3-1) are interpolated using the pixel values of the adjacent blocks in the horizontal and vertical directions to generate a predicted image. The interpolation is performed in the order of the horizontal direction and the vertical direction, but it can also be performed in the order of the vertical direction and the horizontal direction.
[0259] To implement horizontal interpolation, use predSamples[xHor][yHor] and predSamples[xHor+upHor][yHor]( Fig.12 The pixel value at the position indicated by “○” is derived from the shadow pixels of the horizontally interpolated image.
[0260]
[0261] After interpolation in the horizontal direction, use predSamples[xVer][yVer] and predSamples[xVer][yVer+upVer]( Fig.12 The pixel value at the position indicated by “○” is derived from the shadow pixels of the vertically interpolated image.
[0262]
[0263] When nTbH<=nTbW, the pixel values of the adjacent blocks are interpolated in the order of the vertical direction and the horizontal direction to generate a predicted image. The vertical and horizontal interpolation processes are the same as those in the case of nTbH>nTbW.
[0264] The MIP unit 31045 of the second MIP embodiment is characterized in that it derives a square (predW=predH=predSize) intermediate prediction image predMip[][]. The derivation process is simplified to facilitate address calculation of the prediction image.
[0265] The MIP unit 31045 of the second MIP embodiment selects predSize that is smaller than the width nTbW and height nTbH of the target block, so that the matrix size predC (=predSize) selected by sizeId is equal to predW and predH, thereby facilitating reference of matrix elements in predMip derivation.
[0266] In MIP Example 2, by limiting the width and height of the prediction image classified as sizeId=2, the amount of calculation can be significantly reduced compared with MIP Example 1. Simulations have confirmed that there is almost no reduction in encoding efficiency due to these changes.
[0267] (MIP Example 3)
[0268] In this embodiment, another example is described in which the processing is simplified without reducing the coding efficiency compared with MIP Embodiment 1. The following description will focus on the changes, and the parts not described are the same processing as MIP Embodiment 2.
[0269] In this embodiment, when sizeId = 2, it is set to predW = predH = predC. Accordingly, the definition of sizeId is changed. Hereinafter, predW, predH, and predC are described as predSize.
[0270] (1) Boundary reference pixel extraction
[0271] The MIP unit uses the following formula to derive the variable sizeId related to the object block size: Fig.21 Figure above).
[0272] sizeId=(nTbW<=4&&nTbH<=4)? 0: (nTbW<=4||nTbH<=4)? 1:2(MIP-28)
[0273] Or you can use other conditions to determine sizeId( Fig.21 Figure below).
[0274] sizeId=(nTbW<=4&&nTbH<=4)? 0: (nTbW<=8||nTbH<=8)? 1:2(MIP-29)
[0275] (2) Predicted pixel derivation (matrix operation)
[0276] Same as MIP implementation example 2.
[0277] (3) Predicted pixel derivation (linear interpolation)
[0278] Same as MIP embodiment 2.
[0279] As described above, in the third embodiment of MIP, the determination of sizeId is further simplified compared with the second embodiment of MIP, thereby making it possible to further reduce the amount of calculation compared with the second embodiment of MIP.
[0280] It should be noted that MIP embodiment 3 is also the same as MIP embodiment 2. It derives a square (predW=predH=predSize) intermediate prediction image predMip[][], selects a predSize that is less than the width nTbW and height nTbH of the object block, and limits the width and height of the prediction image classified as sizeId=2, thereby achieving the same effect as MIP embodiment 2.
[0281] (MIP Example 4)
[0282] In this embodiment, another example of reducing the memory required for storing the weight matrix is described compared with MIP Embodiment 1. The following description will focus on the changes, and the parts not described are the same processing as MIP Embodiment 2.
[0283] In this embodiment, when sizeId = 2, it is set to predW = predH = predC. Accordingly, the definition of sizeId is changed. Hereinafter, predW, predH, and predC are described as predSize.
[0284] (1) Boundary reference pixel extraction
[0285] The MIP unit derives a variable sizeId related to the target block size using the following formula.
[0286] sizeId=(nTbW<=4||nTbH<=4)? 0:1(MIP-30)
[0287] In the above example, the value of sizeId is set to 0 or 1. Fig. 22 As shown in the figure above, sizeId = (nTbW <= 4 || nTbH <= 4)? 0: 2 (MIP-34) or if Fig. 22 As shown in the figure below, sizeId = (nTbW <= 4 || nTbH <= 4) ? 1:2 (MIP-34), sizeId can be expressed as a combination of 0, 2 or 1, 2. It should be noted that (nTbW <= 8 || nTbH <= 8) can also be used instead of the conditional expression (nTbW <= 4 || nTbH <= 4).
[0288] In the example of formula MIP-30, the value of sizeId is 0 and 1. Therefore, the processing of sizeId=2 in MIP embodiment 2 can be completely omitted. For example, the following formula is only needed to derive p[i] (i=0..2*boundarySize-1) from the second reference area redL[] and redT[].
[0289]
[0290] (2) Predicted pixel derivation (matrix operation)
[0291] It can also be set to be the same as MIP Example 2, but since sizeId=2 is not used, the process of selecting the weight matrix mWeight[predSize*predSize][inSize] from the matrix set with reference to sizeId and modeId omits the case of sizeId=2 and is expressed as the following formula.
[0292]
[0293]
[0294] Similarly, the process of deriving the shift value sW and the offset coefficient fO with reference to sizeId and modeId is expressed by the following formula.
[0295]
[0296] (3) Predicted pixel derivation (linear interpolation)
[0297] Same as MIP embodiment 2.
[0298] As described above, in MIP Example 3, the types of sizeId are reduced compared to MIP Example 2, so the memory required for storing the weight matrix can be reduced compared to MIP Example 2.
[0299] It should be noted that MIP embodiment 3 can also be the same as MIP embodiment 2, deriving a square (predW=predH=predSize) intermediate prediction image predMip[][], selecting a matrix (predSize) less than the object block size nTbWxnTbH, and limiting the width and height of the prediction image classified as sizeId=2, thereby achieving the same effect as MIP embodiment 2.
[0300] (Configuration of the Prediction Image Correction Unit 3105)
[0301] The predicted image correction unit 3105 corrects the temporary predicted image output from the prediction unit 3104 according to the intra-frame prediction mode. Specifically, the predicted image correction unit 3105 derives a position-dependent weight coefficient for each pixel of the temporary predicted image according to the reference area R and the position of the object prediction pixel. Then, by weighted addition (weighted averaging) of the reference sample s[][] and the temporary predicted image, a predicted image (corrected predicted image) Pred[][] that has been corrected for the temporary predicted image is derived. It should be noted that in some intra-frame prediction modes, the output of the prediction unit 3104 may be directly used as the predicted image without correcting the temporary predicted image using the predicted image correction unit 3105.
[0302] The inverse quantization / inverse transformation unit 311 inversely quantizes the quantized transform coefficients input from the entropy decoding unit 301 to obtain transform coefficients. The quantized transform coefficients are coefficients obtained by performing frequency transformation such as DCT (Discrete Cosine Transform) and DST (Discrete Sine Transform) on the prediction error in the encoding process and quantizing them. The inverse quantization / inverse transformation unit 311 performs inverse frequency transformation such as inverse DCT and inverse DST on the obtained transform coefficients to calculate the prediction error. The inverse quantization / inverse transformation unit 311 outputs the prediction error to the addition unit 312.
[0303] The adder 312 adds the block prediction image input from the prediction image generator 308 and the prediction error input from the inverse quantization / inverse transform unit 311 for each pixel to generate a block decoded image. The adder 312 stores the block decoded image in the reference picture memory 306 and outputs it to the loop filter 305.
[0304] (Configuration of Moving Image Coding Device)
[0305] Next, the configuration of the moving picture encoding device 11 according to the present embodiment will be described. Fig.131 is a block diagram showing the structure of a moving picture coding apparatus 11 according to an example of the present embodiment. The moving picture coding apparatus 11 is configured to include a prediction image generation unit 101, a subtraction unit 102, a transformation / quantization unit 103, an inverse quantization / inverse transformation unit 105, an addition unit 106, a loop filter 107, a prediction parameter memory (prediction parameter storage unit, frame memory) 108, a reference image memory (reference image storage unit, frame memory) 109, a coding parameter determination unit 110, a parameter coding unit 111, and an entropy coding unit 104.
[0306] The predicted image generation unit 101 generates a predicted image for each region obtained by dividing each picture of the image T, that is, for each CU. The predicted image generation unit 101 operates in the same manner as the predicted image generation unit 308 described above, and the description thereof is omitted here.
[0307] The subtraction unit 102 generates a prediction error by subtracting the pixel value of the block prediction image input from the prediction image generation unit 101 from the pixel value of the image T. The subtraction unit 102 outputs the prediction error to the transformation / quantization unit 103 .
[0308] The transform / quantization unit 103 calculates transform coefficients by frequency transforming the prediction error input from the subtraction unit 102 , and derives quantized transform coefficients by quantization. The transform / quantization unit 103 outputs the quantized transform coefficients to the entropy coding unit 104 and the inverse quantization / inverse transform unit 105 .
[0309] The inverse quantization / inverse transformation unit 105 and the inverse quantization / inverse transformation unit 311 ( Figure 7 ) is the same as that of FIG. 106, and its description is omitted here. The calculated prediction error is output to the adding unit 106.
[0310] The entropy coding unit 104 receives quantized transform coefficients from the transform / quantization unit 103 and receives coding parameters from the parameter coding unit 111. The entropy coding unit 104 entropy codes partition information, prediction parameters, quantized transform coefficients, etc., generates a coded stream Te, and outputs it.
[0311] The parameter coding unit 111 includes a header coding unit 1110 (not shown), a CT information coding unit 1111 , a CU coding unit 1112 (prediction mode coding unit), an inter prediction parameter coding unit 112 , and an intra prediction parameter coding unit 113 . The CU coding unit 1112 further includes a TU coding unit 1114 .
[0312] (Configuration of Intra-frame Prediction Parameter Coding Unit 113)
[0313] The intra prediction parameter encoder 113 derives a format for encoding (eg, intra_luma_mpm_idx, intra_luma_mpm_remmainder, etc.) based on the IntraPredMode input from the encoding parameter determiner 110. The intra prediction parameter encoder 113 includes a configuration that is partially identical to the configuration in which the intra prediction parameter decoder 304 derives intra prediction parameters.
[0314] Fig.14 1 is a schematic diagram showing the configuration of the intra prediction parameter coding unit 113 of the parameter coding unit 111. The intra prediction parameter coding unit 113 includes a parameter coding control unit 1131, a luma intra prediction parameter derivation unit 1132, and a chroma intra prediction parameter derivation unit 1133.
[0315] IntraPredModeY and IntraPredModeC are input from the encoding parameter determination unit 110 to the parameter encoding control unit 1131. The parameter encoding control unit 1131 determines intra_luma_mpm_flag with reference to mpmCandList[] of the MPM candidate list derivation unit 30421. Then, intra_luma_mpm_flag and IntraPredModeY are output to the luma intra prediction parameter derivation unit 1132. IntraPredModeC is output to the chroma intra prediction parameter derivation unit 1133.
[0316] The luma intra prediction parameter derivation unit 1132 is configured to include an MPM candidate list derivation unit 30421 (candidate list derivation unit), an MPM parameter derivation unit 11322 and a non-MPM parameter derivation unit 11323 (encoding unit, derivation unit).
[0317] The MPM candidate list derivation unit 30421 derives mpmCandList[] with reference to the intra prediction mode of the adjacent block stored in the prediction parameter memory 108. The MPM parameter derivation unit 11322 derives intra_luma_mpm_idx from IntraPredModeY and mpmCandList[] when intra_luma_mpm_flag is 1, and outputs it to the entropy coding unit 104. The non-MPM parameter derivation unit 11323 derives RemIntraPredMode from IntraPredModeY and mpmCandList[] when intra_luma_mpm_flag is 0, and outputs intra_luma_mpm_remainder to the entropy coding unit 104.
[0318] The chroma intra prediction parameter derivation unit 1133 derives intra_chroma_pred_mode from IntraPredModeY and IntraPredModeC, and outputs it.
[0319] The adder 106 generates a decoded image by adding the pixel value of the block prediction image input from the prediction image generator 101 and the prediction error input from the inverse quantization / inverse transform unit 105 for each pixel. The adder 106 stores the generated decoded image in the reference picture memory 109.
[0320] The loop filter 107 performs deblocking filtering, SAO, and ALF on the decoded image generated by the adding unit 106. It should be noted that the loop filter 107 does not necessarily include the above three filters, and may be composed of only a deblocking filter, for example.
[0321] The prediction parameter memory 108 stores the prediction parameter generated by the encoding parameter determination unit 110 in a position predetermined for each target picture and each CU.
[0322] The reference picture memory 109 stores the decoded image generated by the loop filter 107 at a position predetermined for each target picture and each CU.
[0323] The coding parameter determination unit 110 selects one of a plurality of sets of coding parameters. The coding parameters refer to the QT, BT or TT split information, prediction parameters or parameters generated in association with these as the coding target. The predicted image generation unit 101 generates a predicted image using these coding parameters.
[0324] The coding parameter determination unit 110 calculates the RD cost value indicating the amount of information and the coding error for each of the plurality of sets. The coding parameter determination unit 110 selects the coding parameter set with the smallest calculated cost value. Thus, the entropy coding unit 104 outputs the selected coding parameter set as the coding stream Te. The coding parameter determination unit 110 stores the determined coding parameters in the prediction parameter memory 108.
[0325] It should be noted that a part of the moving picture encoding device 11 and the moving picture decoding device 31 in the above-mentioned embodiment, such as the entropy decoding unit 301, the parameter decoding unit 302, the loop filter 305, the predicted image generation unit 308, the inverse quantization / inverse transformation unit 311, the addition unit 312, the predicted image generation unit 101, the subtraction unit 102, the transformation / quantization unit 103, the entropy encoding unit 104, the inverse quantization / inverse transformation unit 105, the loop filter 107, the encoding parameter determination unit 110, and the parameter encoding unit 111, can also be implemented by a computer. In this case, the program for implementing the above-mentioned control function can be recorded in a computer-readable recording medium, and the computer system can read the program recorded in the recording medium and execute it. It should be noted that the "computer system" mentioned here refers to a computer system built into any one of the moving picture encoding device 11 and the moving picture decoding device 31, and a computer system including hardware such as an OS and peripheral devices is used. In addition, "computer-readable recording medium" refers to removable media such as floppy disks, magneto-optical disks, ROMs, CD-ROMs, and storage devices such as hard disks built into computer systems. Moreover, "computer-readable recording medium" may also include: recording media that dynamically store programs for a short period of time, such as communication lines in the case of sending programs via networks such as the Internet or communication lines such as telephone lines; and recording media that store programs for a fixed period of time, such as volatile memories inside computer systems that serve as servers or clients in this case. In addition, the above-mentioned program may be a program for realizing a part of the aforementioned functions, or a program that can realize the aforementioned functions by combining with a program already recorded in a computer system.
[0326] In addition, part or all of the moving picture encoding device 11 and the moving picture decoding device 31 in the above-mentioned embodiment may be implemented as an integrated circuit such as LSI (Large Scale Integration). Each functional block of the moving picture encoding device 11 and the moving picture decoding device 31 may be individually processor-based, or part or all may be integrated to form a processor. In addition, the method of integrated circuitization is not limited to LSI, and may be implemented by a dedicated circuit or a general-purpose processor. In addition, when a technology for integrated circuitization that replaces LSI appears with the advancement of semiconductor technology, the integrated circuit of this technology may also be used.
[0327] As mentioned above, one embodiment of the present invention has been described in detail with reference to the drawings, but the specific configuration is not limited to the above embodiment, and various design changes and the like can be made within the scope not departing from the gist of the present invention.
[0328] [Application Examples]
[0329] The above-mentioned moving picture encoding device 11 and moving picture decoding device 31 can be mounted on various devices for sending, receiving, recording, and reproducing moving pictures. It should be noted that the moving pictures can be natural moving pictures captured by a camera or the like, or artificial moving pictures (including CG and GUI) generated by a computer or the like.
[0330] First, refer to Figure 2 A case where the above-described moving picture encoding device 11 and moving picture decoding device 31 can be used for transmission and reception of moving pictures will be described.
[0331] exist Figure 2 2 is a block diagram showing the structure of a transmitting device PROD_A equipped with a motion picture encoding device 11. Figure 2 As shown, the transmitting device PROD_A comprises: a coding unit PROD_A1 for obtaining coded data by coding a moving picture; a modulating unit PROD_A2 for obtaining a modulated signal by modulating a carrier wave using the coded data obtained in the coding unit PROD_A1; and a transmitting unit PROD_A3 for transmitting the modulated signal obtained by the modulating unit PROD_A2. The moving picture coding device 11 is used as the coding unit PROD_A1.
[0332] The transmitting device PROD_A may further include a camera PROD_A4 for shooting moving images as a supply source of moving images to be input to the encoding unit PROD_A1, a recording medium PROD_A5 on which moving images are recorded, an input terminal PROD_A6 for inputting moving images from the outside, and an image processing unit A7 for generating or processing images. The figure illustrates that the transmitting device PROD_A includes all of these configurations, but some of them may be omitted.
[0333] It should be noted that the recording medium PROD_A5 may be a medium on which uncoded moving images are recorded, or may be a medium on which moving images are recorded after being coded using a coding method for recording that is different from the coding method for transmission. In the latter case, it is preferred that a decoding unit (not shown) that decodes coded data read from the recording medium PROD_A5 using the coding method for recording is interposed between the recording medium PROD_A5 and the coding unit PROD_A1.
[0334] In addition, Figure 2 2 is a block diagram showing the structure of a receiving device PROD_B equipped with a motion picture decoding device 31. Figure 2As shown, the receiving device PROD_B includes: a receiving unit PROD_B1 that receives a modulated signal; a demodulating unit PROD_B2 that demodulates the modulated signal received by the receiving unit PROD_B1 to obtain coded data; and a decoding unit PROD_B3 that decodes the coded data obtained by the demodulating unit PROD_B2 to obtain a moving image. The moving image decoding device 31 is used as the decoding unit PROD_B3.
[0335] The receiving device PROD_B may also include a display PROD_B4 for displaying moving images as a supply destination of the moving images output by the decoding unit PROD_B3, a recording medium PROD_B5 for recording the moving images, and an output terminal PROD_B6 for outputting the moving images to the outside. Figure 2 In the example, the receiving device PROD_B is shown to have all of these structures, but some of them may be omitted.
[0336] It should be noted that the recording medium PROD_B5 may be a medium for recording uncoded moving images, or may be a medium for recording moving images coded in a recording coding method different from a transmission coding method. In the latter case, it is preferred that an encoding unit (not shown) that encodes the moving images obtained from the decoding unit PROD_B3 in the recording coding method is interposed between the decoding unit PROD_B3 and the recording medium PROD_B5.
[0337] It should be noted that the transmission medium for transmitting the modulated signal can be wireless or wired. In addition, the transmission scheme for transmitting the modulated signal can be broadcast (here refers to a transmission scheme in which the transmission destination is not predetermined) or communication (here refers to a transmission scheme in which the transmission destination is predetermined). That is, the transmission of the modulated signal can be achieved through any one of wireless broadcasting, wired broadcasting, wireless communication and wired communication.
[0338] For example, a broadcasting station (broadcasting equipment, etc.) / receiving station (television receiver, etc.) of terrestrial digital broadcasting is an example of a transmitting device PROD_A / receiving device PROD_B that transmits and receives modulated signals via wireless broadcasting. In addition, a broadcasting station (broadcasting equipment, etc.) / receiving station (television receiver, etc.) of cable television broadcasting is an example of a transmitting device PROD_A / receiving device PROD_B that transmits and receives modulated signals via wired broadcasting.
[0339] In addition, a server (workstation, etc.) / client (TV receiver, personal computer, smart phone, etc.) of a VOD (Video On Demand) service or a motion picture sharing service using the Internet is an example of a transmitter PROD_A / receiver PROD_B that transmits and receives a modulated signal through communication (usually, either wireless or wired is used as a transmission medium in a LAN, and wired is used as a transmission medium in a WAN). Here, a personal computer includes a desktop PC, a laptop PC, and a tablet PC. In addition, a smart phone also includes a multi-function portable phone terminal.
[0340] It should be noted that the client of the moving image sharing service has the function of decoding the encoded data downloaded from the server and displaying it on the display, and also has the function of encoding the moving images captured by the camera and uploading them to the server. That is, the client of the moving image sharing service functions as both the sending device PROD_A and the receiving device PROD_B.
[0341] Next, refer to Figure 3 A case where the above-described moving image encoding device 11 and moving image decoding device 31 can be used for recording and reproducing moving images will be described.
[0342] exist Figure 3 2 is a block diagram showing the structure of a recording device PROD_C equipped with the above-mentioned motion picture encoding device 11. Figure 3 As shown, the recording device PROD_C includes: an encoding unit PROD_C1 that obtains encoded data by encoding a moving image; and a writing unit PROD_C2 that writes the encoded data obtained by the encoding unit PROD_C1 to the recording medium PROD_M. The moving image encoding device 11 described above is used as the encoding unit PROD_C1.
[0343] It should be noted that the recording medium PROD_M can be (1) a type of recording medium built into the recording device PROD_C such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive), or (2) a type of recording medium connected to the recording device PROD_C such as an SD memory card or a USB (Universal Serial Bus) flash memory, or (3) a recording medium loaded into a drive device (not shown) built into the recording device PROD_C such as a DVD (Digital Versatile Disc: registered trademark) or a BD (Blu-ray Disc: registered trademark).
[0344] In addition, the recording device PROD_C may further include a camera PROD_C3 for shooting moving images as a supply source of moving images input to the encoding unit PROD_C1, an input terminal PROD_C4 for inputting moving images from the outside, a receiving unit PROD_C5 for receiving moving images, and an image processing unit PROD_C6 for generating or processing images. Figure 3 In the example, the recording device PROD_C is shown to have all of these structures, but some of them may be omitted.
[0345] It should be noted that the receiving unit PROD_C5 can receive uncoded moving images, or can receive coded data coded in a transmission coding method different from the recording coding method. In the latter case, it is preferable to place a transmission decoding unit (not shown) that decodes the coded data coded in the transmission coding method between the receiving unit PROD_C5 and the coding unit PROD_C1.
[0346] Examples of such a recording device PROD_C include a DVD recorder, a BD recorder, and a HDD (Hard Disk Drive) recorder (in this case, the input terminal PROD_C4 or the receiving unit PROD_C5 is the main source of motion images). In addition, a portable camera (in this case, the camera PROD_C3 is the main source of motion images), a personal computer (in this case, the receiving unit PROD_C5 or the image processing unit C6 is the main source of motion images), and a smart phone (in this case, the camera PROD_C3 or the receiving unit PROD_C5 is the main source of motion images) are also examples of such a recording device PROD_C.
[0347] In addition, Figure 3 2 is a block diagram showing the structure of a playback device PROD_D equipped with the above-mentioned motion picture decoding device 31. Figure 3 As shown, the playback device PROD_D includes a reading unit PROD_D1 for reading the coded data written in the recording medium PROD_M, and a decoding unit PROD_D2 for decoding the coded data read by the reading unit PROD_D1 to obtain a moving picture. The moving picture decoding device 31 described above is used as the decoding unit PROD_D2.
[0348] It should be noted that the recording medium PROD_M can be (1) a type of recording medium built into the reproduction device PROD_D such as an HDD or SSD, or (2) a type of recording medium connected to the reproduction device PROD_D such as an SD memory card or USB flash memory, or (3) a recording medium loaded into a drive device (not shown) built into the reproduction device PROD_D such as a DVD or BD.
[0349] In addition, the playback device PROD_D may further include a display PROD_D3 for displaying moving images as a supply destination of the moving images output by the decoding unit PROD_D2, an output terminal PROD_D4 for outputting the moving images to the outside, and a transmission unit PROD_D5 for transmitting the moving images. Figure 3 In the example, the reproduction device PROD_D is shown to have all of these structures, but some of them may be omitted.
[0350] It should be noted that the sending unit PROD_D5 can send uncoded moving images or coded data coded in a transmission coding method different from the recording coding method. In the latter case, it is better to place a coding unit (not shown) that encodes moving images in a transmission coding method between the decoding unit PROD_D2 and the sending unit PROD_D5.
[0351] Examples of such a reproduction device PROD_D include a DVD player, a BD player, and an HDD player (in this case, the output terminal PROD_D4 connected to a television receiver or the like is the main supply destination of motion images). In addition, a television receiver (in this case, the display PROD_D3 is the main supply destination of motion images), a digital signage (also called an electronic billboard, an electronic bulletin board, etc., the display PROD_D3 or the transmission unit PROD_D5 is the main supply destination of motion images), a desktop PC (in this case, the output terminal PROD_D4 or the transmission unit PROD_D5 is the main supply destination of motion images), a laptop or tablet PC (in this case, the display PROD_D3 or the transmission unit PROD_D5 is the main supply destination of motion images), and a smartphone (in this case, the display PROD_D3 or the transmission unit PROD_D5 is the main supply destination of motion images) are also examples of such a reproduction device PROD_D. (Implemented in hardware and implemented in software)
[0352] Furthermore, each block of the video decoding device 31 and the video encoding device 11 may be implemented in hardware by a logic circuit formed on an integrated circuit (IC chip), or may be implemented in software by a CPU (Central Processing Unit).
[0353] In the latter case, each of the above-mentioned devices has: a CPU that executes commands of a program that realizes each function, a ROM (Read Only Memory) storing the above-mentioned program, a RAM (Random Access Memory) that expands the above-mentioned program, and a storage device (recording medium) such as a memory that stores the above-mentioned program and various data. Moreover, the purpose of the embodiment of the present invention can also be achieved in the following manner: a recording medium that records the program code (executable form program, intermediate code program, source program) of the control program of the above-mentioned each device in a computer-readable manner is supplied to each of the above-mentioned devices, and the computer (or CPU, MPU) reads and executes the program code recorded in the recording medium.
[0354] As the above-mentioned recording medium, for example, the following can be used: magnetic tapes such as magnetic tapes and cassette tapes; magnetic disks including floppy disks (registered trademark) / hard disks, CD-ROMs (Compact Disc Read-Only Memory) / MO disks (Magneto-Optical disc) / MD (Mini Disc) / DVD (Digital Versatile Disc) / CD-R (CD Recordable) / Blu-ray Discs (registered trademark) and the like; cards such as IC cards (including memory cards) / optical cards; semiconductor memories such as mask ROM / EPROM (Erasable Programmable Read-Only Memory) / EEPROM (Electrically Erasable and Programmable Read-Only Memory: registered trademark) / flash ROM; or logic circuits such as PLDs (Programmable logic device) and FPGAs (Field Programmable Gate Array), etc.
[0355] In addition, each of the above-mentioned devices can also be configured to be connected to a communication network, and the above-mentioned program code can be supplied via the communication network. The communication network is not particularly limited as long as it can transmit the program code. For example, the following can be used: the Internet, intranet, extranet, LAN (Local Area Network), ISDN (Integrated Services Digital Network), VAN (Value-Added Network), CATV (Community Antennalevision / Cable Television) communication network, virtual private network (Virtual Private Network), telephone line network, mobile communication network, satellite communication network, etc. In addition, the transmission medium constituting the communication network can be any medium capable of transmitting the program code, and is not limited to a specific structure or type. For example, it can be used for wired communication such as IEEE (Institute of Electrical and Electronic Engineers) 1394, USB, power line transmission, cable TV line, telephone line, ADSL (Asymmetric Digital Subscriber Line) line, etc., and can also be used for wireless communication such as IrDA (Infrared Data Association), infrared such as remote control, BlueTooth (registered trademark), IEEE802.11 wireless, HDR (High Data Rate), NFC (Near Field Communication), DLNA (Digital Living Network Alliance: registered trademark), mobile phone network, satellite line, terrestrial digital broadcasting network, etc. It should be noted that the embodiments of the present invention can also be implemented in the form of a computer data signal embedded in a carrier wave that embodies the above-mentioned program code through electronic transmission.
[0356] The embodiments of the present invention are not limited to the above-mentioned embodiments, and various modifications can be made within the scope of the claims. That is, embodiments obtained by combining technical solutions appropriately modified within the scope of the claims are also included in the technical scope of the present invention.
[0357] Industrial Applicability
[0358] The embodiments of the present invention can be preferably applied to a motion picture decoding device that decodes coded data after encoding image data and a motion picture encoding device that generates coded data after encoding image data. In addition, it can be preferably applied to the data structure of coded data generated by the motion picture encoding device and referred to by the motion picture decoding device.
[0359] Explanation of symbols
[0360] 31: Image decoding device
[0361] 301: Entropy decoding unit
[0362] 302: Parameter decoding unit
[0363] 3020: Header decoding unit
[0364] 303: Inter-frame prediction parameter decoding unit
[0365] 304: Intra-frame prediction parameter decoding unit
[0366] 308: Prediction image generation unit
[0367] 309: Inter-frame prediction image generation unit
[0368] 310: Intra-frame prediction image generation unit
[0369] 311: Inverse quantization / inverse transformation unit
[0370] 312: Addition Department
[0371] 11: Image coding device
[0372] 101: Prediction image generation unit
[0373] 102: Subtraction Department
[0374] 103: Transformation / Quantization Unit
[0375] 104: Entropy coding unit
[0376] 105: Inverse quantization / inverse transformation unit
[0377] 107: Loop filter
[0378] 110: Coding parameter determination unit
[0379] 111: Parameter encoding unit
[0380] 112: Inter-frame prediction parameter encoding unit
[0381] 113: Intra-frame prediction parameter encoding unit
[0382] 1110: Header encoding unit
[0383] 1111: CT Information Coding Department
[0384] 1112: CU encoding unit (prediction mode encoding unit)
[0385] 1114: TU encoding unit
Claims
1. A moving picture decoding device for decoding coded data, the moving picture decoding device comprising: a weight matrix deriving unit configured to derive a weight matrix by using (i) an intra prediction mode and (ii) a size variable set according to a width of a transform block and a height of the transform block; a matrix prediction image deriving unit configured to derive an intermediate prediction image defined by a predetermined size by using (i) a reference image obtained by downsampling neighboring pixel values adjacent to the target block; and (ii) a product-sum operation of the weight matrix; and a matrix prediction image interpolation section configured to derive a predicted image by using the intermediate prediction image, in: When a first condition is true, the first condition being that the value of the width of the transform block and the value of the height of the transform block are both equal to 4, the size variable is set equal to 0 and the predetermined size is set equal to 4; If the first condition is false, but a second condition is true, the second condition being (i) the value of the width of the transform block or the value of the height of the transform block is equal to 4, or (ii) the value of the width of the transform block and the value of the height of the transform block are both equal to 8, the size variable is set equal to 1, and the predetermined size is set equal to 4; and In the event that both the first condition and the second condition are negative, the size variable is set equal to 2 and the predefined size is set to 8.
2. The moving picture decoding device according to claim 1, characterized in that The intermediate prediction images are defined as square matrices.
3. A motion picture encoding device for encoding image data, the motion picture encoding device comprising: a weight matrix deriving unit configured to derive a weight matrix by using (i) an intra prediction mode and (ii) a size variable set according to a width of a transform block and a height of the transform block; a matrix prediction image deriving unit configured to derive an intermediate prediction image defined by a predetermined size by using (i) a reference image obtained by downsampling neighboring pixel values adjacent to the target block; and (ii) a product-sum operation of the weight matrix; and a matrix prediction image interpolation section configured to derive a predicted image by using the intermediate prediction image, in: When a first condition is true, the first condition being that the value of the width of the transform block and the value of the height of the transform block are both equal to 4, the size variable is set equal to 0 and the predetermined size is set equal to 4; If the first condition is false, but a second condition is true, the second condition being (i) the value of the width of the transform block or the value of the height of the transform block is equal to 4, or (ii) the value of the width of the transform block and the value of the height of the transform block are both equal to 8, the size variable is set equal to 1, and the predetermined size is set equal to 4; and In the event that both the first condition and the second condition are negative, the size variable is set equal to 2 and the predefined size is set to 8.
4. A moving picture decoding method for decoding coded data, the moving picture decoding method comprising: deriving a weight matrix by using (i) an intra prediction mode and (ii) a size variable set according to a width of a transform block and a height of the transform block; An intermediate prediction image defined by a predetermined size is derived by using (i) a reference image obtained by downsampling neighboring pixel values adjacent to the target block; and (ii) a product-sum operation of the weight matrix; and deriving a predicted image by using the intermediate prediction image, in: When a first condition is true, the first condition being that the value of the width of the transform block and the value of the height of the transform block are both equal to 4, the size variable is set equal to 0 and the predetermined size is set equal to 4; When the first condition is false, but a second condition is true, the second condition being (i) the value of the width of the transform block or the value of the height of the transform block is equal to 4, or (ii) the value of the width of the transform block and the value of the height of the transform block are both equal to 8, the size variable is set equal to 1, and the predetermined size is set equal to 4; as well as In the event that both the first condition and the second condition are negative, the size variable is set equal to 2 and the predefined size is set to 8.
5. A moving picture encoding method for encoding image data, the moving picture encoding method comprising: deriving a weight matrix by using (i) an intra prediction mode and (ii) a size variable set according to a width of a transform block and a height of the transform block; An intermediate prediction image defined by a predetermined size is derived by using (i) a reference image obtained by downsampling neighboring pixel values adjacent to the target block; and (ii) a product-sum operation of the weight matrix; and deriving a predicted image by using the intermediate prediction image, in: When a first condition is true, the first condition being that the value of the width of the transform block and the value of the height of the transform block are both equal to 4, the size variable is set equal to 0 and the predetermined size is set equal to 4; When the first condition is false, but a second condition is true, the second condition being (i) the value of the width of the transform block or the value of the height of the transform block is equal to 4, or (ii) the value of the width of the transform block and the value of the height of the transform block are both equal to 8, the size variable is set equal to 1, and the predetermined size is set equal to 4; as well as In the event that both the first condition and the second condition are negative, the size variable is set equal to 2 and the predefined size is set to 8.
Citation Information
Patent Citations
Image decoding apparatus, image encoding apparatus, and data structure of encoded data
CN103493494A
Video coding using hybrid intra prediction
CN108781283A