Moving image decoding device and moving image encoding device
By designing a matrix reference pixel derivation unit, a weight matrix derivation unit, a matrix predicted image derivation unit and a matrix predicted image interpolation unit in the moving image decoding and encoding device, the intra prediction process is optimized, and the problems of large size and large processing volume in the prior art are solved, thereby achieving efficient intra prediction.
Patent Information
- Application Number
- CN202510548933.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-20
- Filing Date
- 2020-09-17
- Publication Date
- 2025-06-10
AI Technical Summary
In the existing matrix intra prediction technology, there are problems such as large memory size and huge processing volume of the weight matrix.
A moving image decoding device and an encoding device are designed to optimize the intra prediction process and reduce the memory size and processing amount of the weight matrix through the matrix reference pixel derivation unit, the weight matrix derivation unit, the matrix prediction image derivation unit, and the matrix prediction image interpolation unit.
It is realized that the preferred intra prediction is performed while reducing the weight matrix memory size, reducing the processing amount and improving the encoding and decoding efficiency.
Smart Images

Figure CN120128705A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to a moving image decoding apparatus and a moving image encoding apparatus. Background Art
[0002] In order to efficiently transmit or record moving images, a moving image encoding apparatus that generates encoded data by encoding a moving image and a moving image decoding apparatus that generates a decoded image by decoding the encoded data are used.
[0003] As a specific moving image encoding method, for example, methods proposed in H.264 / AVC, HEVC
[0004] (High-Efficiency Video Coding) can be cited.
[0005] In such a moving image encoding method, images (pictures) constituting a moving image are managed by a hierarchical structure including slices, coding tree units (CTUs), coding units (sometimes also referred to as coding units (CUs)), and transform units (TUs). Encoding / decoding is performed for each CU. The above-mentioned slices are obtained by dividing an image, the above-mentioned coding tree units are obtained by dividing a slice, the above-mentioned coding units are obtained by dividing a coding tree unit, and the above-mentioned transform units are obtained by dividing a coding unit.
[0006] In addition, in such a moving image encoding method, a prediction image is usually generated based on a locally decoded image obtained by encoding / decoding an input image, and a prediction error (sometimes also referred to as a "difference image" or "residual image") obtained by subtracting the prediction image from the input image (original image) is encoded. As methods for generating a prediction image, inter-picture prediction (inter-frame prediction) and intra-picture prediction (intra-frame prediction) can be cited.
[0007] In addition, as recent moving image encoding and decoding techniques, Non-Patent Document 1 can be cited. In Non-Patent Document 2, a matrix-based intra prediction technique (MIP) for deriving a prediction image by product-sum operation of a reference image derived from an adjacent image and a weight matrix is disclosed.
[0008] Prior Art Documents
[0009] Non-Patent Documents
[0010] Non-Patent Document 1: "Versatile Video Coding (Draft 6)", JVET-O2001-vE, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11
[0011] Non-Patent Document 2: "CE3: Affine linear weighted intra prediction (CE3-4.1, CE3-4.2)", JVET-N0217-v1, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11 Summary of the Invention
[0012] Problems to be Solved by the Invention
[0013] In the matrix intra prediction such as Non-Patent Document 1 and Non-Patent Document 2, different weight matrices are stored according to various block sizes and intra prediction modes, so there is a problem that the memory size for storing the weight matrices is large. In addition, there is a problem that the processing amount for generating a predicted image is huge.
[0014] An object of the present invention is to perform preferable intra prediction and reduce the processing amount while reducing the memory size of the weight matrix.
[0015] Technical Solution
[0016] In order to solve the above problems, a moving image decoding device according to one aspect of the present invention is characterized by including: a matrix reference pixel derivation unit that derives, as a reference image, an image obtained by downsampling an image adjacent to the upper side and the left side of an object block; a weight matrix derivation unit that derives a matrix of weight coefficients according to an intra prediction mode and an object block size; a matrix predicted image derivation unit that derives a predicted image from the product of elements of the reference image and elements of the matrix of the weight coefficients; and a matrix predicted image interpolation unit that derives, as a predicted image, the predicted image or an image obtained by interpolating the predicted image, wherein the weight matrix derivation unit derives a matrix having a size equal to or smaller than the width and the height of the object block size.
[0017] The moving image decoding device is characterized in that the weight matrix derivation unit derives a 4×4 matrix when one side of the object block is 4.
[0018] The above-described moving image decoding device is characterized in that the weight matrix derivation unit derives a 4×4 matrix when the object block size is 4×16 and 16×4.
[0019] The above-described moving image decoding device is characterized in that the weight matrix derivation unit derives either a matrix of size 4×4 with sizeId = 0, 1 or a matrix of size 8×8 with sizeId = 2, and derives a matrix of sizeId = 1 or 2 when one side of the object block is 4.
[0020] In addition, the weight matrix derivation unit may derive a 4×4 matrix when the product of the width and height of the object block size is 64 or less.
[0021] The above-described moving image decoding device is characterized in that the matrix prediction image derivation unit derives a square intermediate prediction image predMip[][] with equal width and height.
[0022] A moving image encoding device, characterized by comprising: a matrix reference pixel derivation unit that derives, as a reference image, an image obtained by downsampling an image adjacent to the upper side and the left side of an object block; a weight matrix derivation unit that derives a matrix of weight coefficients according to an intra prediction mode and an object block size; a matrix prediction image derivation unit that derives a prediction image from the product of elements of the reference image and elements of the matrix of weight coefficients; and a matrix prediction image interpolation unit that derives the prediction image or an image obtained by interpolating the prediction image as a prediction image, wherein the weight matrix derivation unit derives a matrix with a size less than or equal to the width and less than or equal to the height of the object block size.
[0023] The above-described moving image encoding device is characterized in that the weight matrix derivation unit derives a 4×4 matrix when one side of the object block is 4.
[0024] The above-described moving image encoding device is characterized in that the weight matrix derivation unit derives a 4×4 matrix when the object block size is 4×16 and 16×4.
[0025] The above-described moving image encoding device is characterized in that the weight matrix derivation unit derives either a matrix of size 4×4 with sizeId = 0, 1 or a matrix of size 8×8 with sizeId = 2, and derives a matrix of sizeId = 1 or 2 when one side of the object block is 4.
[0026] In addition, the weight matrix derivation unit may derive a 4×4 matrix when the product of the width and height of the object block size is 64 or less.
[0027] The above-described moving image encoding apparatus is characterized in that the matrix prediction image derivation unit derives an intermediate prediction image predMip[][], which is a square with equal width and height.
[0028] Advantageous Effects
[0029] According to one aspect of the present invention, it is possible to perform preferable intra prediction while reducing the memory size of the weight matrix or reducing the processing amount. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a schematic diagram showing the configuration of the image transmission system of the present embodiment.
[0031] Figure 2 It is a diagram showing the configurations of a transmitting device equipped with the moving image encoding apparatus of the present embodiment and a receiving device equipped with a moving image decoding apparatus. PROD_A represents the transmitting device equipped with the moving image encoding apparatus, and PROD_B represents the receiving device equipped with the moving image decoding apparatus.
[0032] Figure 3 It is a diagram showing the configurations of a recording device equipped with the moving image encoding apparatus of the present embodiment and a reproducing device equipped with a moving image decoding apparatus. PROD_C represents the recording device equipped with the moving image encoding apparatus, and PROD_D represents the reproducing device equipped with the moving image decoding apparatus.
[0033] Figure 4 It is a diagram showing the hierarchical structure of the data of the encoded stream.
[0034] Figure 5 It is a diagram showing an example of the division of CTUs.
[0035] Figure 6 It is a schematic diagram showing the types (mode numbers) of intra prediction modes.
[0036] Figure 7 It is a schematic diagram showing the configuration of the moving image decoding apparatus.
[0037] Figure 8 It is a schematic diagram showing the configuration of the intra prediction parameter decoding unit.
[0038] Figure 9 It is a diagram showing the reference region used in intra prediction.
[0039] Figure 10 It is a diagram showing the configuration of the intra prediction image generation unit.
[0040] Figure 11 It is a diagram showing an example of MIP processing.
[0041] Figure 12 It is a diagram showing an example of MIP processing.
[0042] Figure 13 It is a block diagram showing the configuration of a moving image encoding device.
[0043] Figure 14 It is a schematic diagram showing the configuration of an intra prediction parameter encoding unit.
[0044] Figure 15 It is a diagram showing the details of the MIP unit.
[0045] Figure 16 It is a diagram showing the MIP processing of the present embodiment.
[0046] Figure 17 It is a diagram showing the parameters for generating a prediction image in the case of deriving a predMip including a non-square through MIP.
[0047] Figure 18 It is a diagram showing a method for deriving sizeId in an embodiment of the present invention (MIP Example 1).
[0048] Figure 19 It is a diagram showing the parameters for generating a prediction image in the case of deriving a square predMip through MIP.
[0049] Figure 20 It is a diagram showing a method for deriving sizeId in an embodiment of the present invention (MIP Example 2).
[0050] Figure 21 It is a diagram showing a method for deriving sizeId in an embodiment of the present invention (MIP Example 3).
[0051] Figure 22 It is a diagram showing a method for deriving sizeId in an embodiment of the present invention (MIP Example 4). Detailed Embodiments
[0052] (First Embodiment)
[0053] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings.
[0054] Figure 1 It is a schematic diagram showing the configuration of the image transmission system 1 of the present embodiment.
[0055] The image transmission system 1 is a system that transmits an encoded stream obtained by encoding an encoded object image and decodes the transmitted encoded stream to display the image. The image transmission system 1 is configured to include: a moving image encoding device (image encoding device) 11, a network 21, a moving image decoding device (image decoding device) 31, and a moving image display device (image display device) 41.
[0056] An image T is input to the moving image encoding device 11.
[0057] The network 21 transmits the encoded stream Te generated by the moving image encoding device 11 to the moving image decoding device 31. The network 21 is the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. The network 21 is not necessarily limited to a two-way communication network and may also be a one-way communication network that transmits broadcast waves such as terrestrial digital broadcasts and satellite broadcasts. In addition, the network 21 may be replaced by a storage medium such as a DVD (Digital Versatile Disc: registered trademark) or a BD (Blue-ray Disc: registered trademark) on which the encoded stream Te is recorded.
[0058] The moving image decoding device 31 decodes each encoded stream Te transmitted by the network 21 and generates one or more decoded images Td after decoding.
[0059] The moving image display device 41 displays all or part of the one or more decoded images Td generated by the moving image decoding device 31. The moving image display device 41 includes, for example, a display device such as a liquid crystal display or an organic EL (Electro-luminescence) display. Examples of the display include a fixed type, a mobile type, and an HMD (Helmet Mounted Display). In addition, when the moving image decoding device 31 has high processing power, an image with high display quality is displayed, and when it has only low processing power, an image that does not require high processing power and display ability is displayed.
[0060] <Operator>
[0061] The following describes the operators used in this specification.
[0062] >> is a right shift, << is a left shift, & is a bitwise AND, | is a bitwise OR, |= is an OR assignment operator, and || represents a logical OR.
[0063] x? y : z is a ternary operator that takes y when x is true (other than 0) and takes z when x is false (0).
[0064] Clip3(a, b, c) is a function that limits c to a value above a and below b, and returns a when c < a, returns b when c > b, and returns c in other cases (where a <= b).
[0065] Clip1Y(c) is an operator where a = 0 and b = (1 << BitDepthY) - 1 are set in Clip3(a, b, c). BitDepthY is the bit depth of luminance.
[0066] abs(a) is a function that returns the absolute value of a.
[0067] Int(a) is a function that returns the integer value of a.
[0068] floor(a) is a function that returns the largest integer less than or equal to a.
[0069] ceil(a) is a function that returns the smallest integer greater than or equal to a.
[0070] a / d means a divided by d (truncating the decimal part).
[0071] min(a, b) represents a function that returns the smaller value of a and b.
[0072] <Structure of Encoded Stream Te>
[0073] Before explaining the moving image encoding device 11 and the moving image decoding device 31 of the present embodiment in detail, the data structure of the encoded stream Te generated by the moving image encoding device 11 and decoded by the moving image decoding device 31 will be described.
[0074] Figure 4 is a diagram showing the hierarchical structure of the data in the encoded stream Te. The encoded stream Te exemplarily includes a sequence and multiple pictures constituting the sequence. Figure 4 shows diagrams representing the encoded video sequence of a specified sequence SEQ, the encoded picture of a specified picture PICT, the encoded slice of a specified slice S, the encoded slice data of the specified slice data, the encoded tree units included in the encoded slice data, and the encoded units included in the encoded tree units, respectively.
[0075] (Encoded Video Sequence)
[0076] In the encoded video sequence, a set of data referred to by the moving image decoding device 31 is specified for decoding the sequence SEQ to be processed. As shown in the encoded video sequence of Figure 4 , the sequence SEQ includes: a Video Parameter Set, a Sequence Parameter Set SPS
[0077] (Sequence Parameter Set), Picture Parameter Set (PPS), Picture (PICT), and Supplemental Enhancement Information (SEI).
[0078] In a moving picture composed of multiple layers, the Video Parameter Set (VPS) defines a set of encoding parameters common to multiple moving pictures and a set of encoding parameters associated with each layer included in the moving picture and the multiple layers.
[0079] In the Sequence Parameter Set (SPS), a set of encoding parameters that the moving picture decoding device 31 refers to for decoding an object sequence is defined. For example, the width and height of a picture are defined. It should be noted that multiple SPSs can exist. In this case, any one of the multiple SPSs is selected from the PPS.
[0080] In the Picture Parameter Set (PPS), a set of encoding parameters that the moving picture decoding device 31 refers to for decoding each picture within an object sequence is defined. For example, it includes the reference value of the quantization width (pic_init_qp_minus26) used in picture decoding and a flag (weighted_pred_flag) indicating the application of weighted prediction. It should be noted that multiple PPSs can exist. In this case, any one of the multiple PPSs is selected from each picture within the object sequence.
[0081] (Encoded Picture)
[0082] In an encoded picture, a set of data that the moving picture decoding device 31 refers to for decoding the picture PICT to be processed is defined. As shown in the encoded picture of Figure 4 , the picture PICT includes slices 0 to NS - 1 (NS is the total number of slices included in the picture PICT).
[0083] It should be noted that hereinafter, when it is not necessary to distinguish each of the slices 0 to NS - 1, the encoding subscript may sometimes be omitted in the description. The same applies to other data with subscripts included in the encoding stream Te described below.
[0084] (Encoded Slice)
[0085] In an encoded slice, a set of data that the moving picture decoding device 31 refers to for decoding the slice S to be processed is defined. As shown in the encoded slice of Figure 4 , it includes a slice header and slice data.
[0086] The slice header includes a set of coding parameters referred to by the moving image decoding device 31 for determining the decoding method of the object slice. The slice type specifying information (slice_type) that specifies the slice type is an example of the coding parameters included in the slice header.
[0087] As the slice types that can be specified by the slice type specifying information, the following can be cited: (1) I slice that uses only intra prediction during encoding, (2) P slice that uses unidirectional prediction or intra prediction during encoding, and (3) B slice that uses unidirectional prediction, bidirectional prediction, or intra prediction during encoding, etc. It should be noted that inter prediction is not limited to unidirectional prediction and bidirectional prediction, and more reference pictures can also be used to generate the predicted image. Hereinafter, in the case of P and B slices, it means a slice including blocks that can use inter prediction.
[0088] It should be noted that the slice header may also include a reference to the picture parameter set PPS (pic_parameter_set_id).
[0089] (Slice encoded data)
[0090] In the encoded slice data, a set of data referred to by the moving image decoding device 31 is defined for decoding the slice data to be processed. As Figure 4 shown in the encoded slice header, the slice data includes CTUs. A CTU is a block of a fixed size (e.g., 64×64) that constitutes a slice, and is sometimes also called the largest coding unit (LCU: Largest Coding Unit).
[0091] (Coding tree unit)
[0092] In Figure 4 the coding tree unit, a set of data referred to by the moving image decoding device 31 is defined for decoding the CTU to be processed. The CTU is divided into coding units CU, which are the basic units of coding processing, by recursive quadtree partitioning (QT (Quad Tree) partitioning), binary tree partitioning (BT (Binary Tree) partitioning), or ternary tree partitioning (TT (Ternary Tree) partitioning). The BT partitioning and TT partitioning are collectively called multi-tree partitioning (MT (Multi Tree) partitioning). The nodes of the tree structure obtained by recursive quadtree partitioning are called coding nodes. The intermediate nodes of the quadtree, binary tree, and ternary tree are coding nodes, and the CTU itself is also defined as the highest-level coding node.
[0093] The CT includes a QT segmentation flag (cu_split_flag) indicating whether QT segmentation is performed, an MT segmentation flag (split_mt_flag) indicating the presence or absence of MT segmentation, an MT segmentation direction (split_mt_dir) indicating the segmentation direction of MT segmentation, and an MT segmentation type (split_mt_type) indicating the segmentation type of MT segmentation as CT information. The cu_split_flag, split_mt_flag, split_mt_dir, and split_mt_type are transmitted for each coding node.
[0094] When the cu_split_flag is 1, the coding node is split into four coding nodes ( Figure 5 of QT).
[0095] When the cu_split_flag is 0, when the split_mt_flag is 0, the coding node has one CU as a node without being split ( Figure 5 of non - split). The CU is the end node of the coding node and is not further split. The CU is the basic unit of coding processing.
[0096] When the split_mt_flag is 1, the coding node is MT - segmented as follows. When the split_mt_type is 0, when the split_mt_dir is 1, the coding node is horizontally split into two coding nodes ( Figure 5 of BT (horizontal split)), and when the split_mt_dir is 0, the coding node is vertically split into two coding nodes ( Figure 5 of BT (vertical split)). In addition, when the split_mt_type is 1, when the split_mt_dir is 1, the coding node is horizontally split into three coding nodes ( Figure 5 of TT (horizontal split)), and when the split_mt_dir is 0, the coding node is vertically split into three coding nodes ( Figure 5 of TT (vertical split)). This content is shown in the Figure 5 CT information.
[0097] In addition, when the size of the CTU is 64×64 pixels, the size of the CU can be any one of 64×64 pixels, 64×32 pixels, 32×64 pixels, 32×32 pixels, 64×16 pixels, 16×64 pixels, 32×16 pixels, 16×32 pixels, 16×16 pixels, 64×8 pixels, 8×64 pixels, 32×8 pixels, 8×32 pixels, 16×8 pixels, 8×16 pixels, 8×8 pixels, 64×4 pixels, 4×64 pixels, 32×4 pixels, 4×32 pixels, 16×4 pixels, 4×16 pixels, 8×4 pixels, 4×8 pixels, and 4×4 pixels.
[0098] (Coding Unit)
[0099] As Figure 4 shown by the coding unit of
[0100] it is stipulated that a set of data referred to by the moving image decoding device 31 is used to decode the coding unit to be processed. Specifically, the CU is composed of a CU header CUH, prediction parameters, transform parameters, quantized transform coefficients, etc. The prediction mode, etc. is stipulated in the CU header.
[0101] The prediction process may be performed in units of CU or in units of sub-CUs obtained by further dividing the CU. When the size of the CU is equal to the size of the sub-CU, there is one sub-CU in the CU. When the size of the CU is larger than the size of the sub-CU, the CU is divided into sub-CUs. For example, when the CU is 8×8 and the sub-CU is 4×4, the CU is divided into four sub-blocks, including two horizontally divided parts and two vertically divided parts.
[0101] There are two types of prediction types (prediction modes): intra prediction and inter prediction. Intra prediction is prediction within the same picture, and inter prediction refers to prediction processing between different pictures (for example, between display times, between layer images).
[0102] The transform / quantization process is performed in units of CU, but the quantized transform coefficients can also be entropy-coded in units of sub-blocks such as 4×4.
[0103] (Prediction Parameters)
[0104] The predicted image is derived from the prediction parameters attached to the block. The prediction parameters include prediction parameters for intra prediction and inter prediction.
[0105] Hereinafter, the prediction parameters for intra prediction will be described. The intra prediction parameters are composed of a luminance prediction mode IntraPredModeY and a chrominance prediction mode IntraPredModeC. Figure 6 is a schematic diagram showing the types (mode numbers) of intra prediction modes. As Figure 6As shown, there are, for example, 67 intra prediction modes (0 to 66). For example, they are planar prediction (0), DC prediction (1), and Angular prediction (2 to 66). Furthermore, an LM mode (67 to 72) can be added to the chrominance difference.
[0106] Among the syntax elements for deriving intra prediction parameters, there are, for example, intra_luma_mpm_flag, intra_luma_mpm_idx, intra_luma_mpm_remainder, etc.
[0107] (MPM)
[0108] intra_luma_mpm_flag is a flag indicating whether the IntraPredModeY of the target block is the same as the MPM (Most Probable Mode). MPM is the prediction mode included in the MPM candidate list mpmCandList[]. The MPM candidate list is a list that stores candidates that are speculated to be highly likely to be applied to the target block based on the intra prediction mode of adjacent blocks and a specified intra prediction mode. When intra_luma_mpm_flag is 1, the IntraPredModeY of the target block is derived using the MPM candidate list and the index intra_luma_mpm_idx.
[0109] IntraPredModeY = mpmCandList[intra_luma_mpm_idx]
[0110] (REM)
[0111] When intra_luma_mpm_flag is 0, the intra prediction mode is selected from the modes RemIntraPredMode that remain after removing the intra prediction modes included in the MPM candidate list from all intra prediction modes. The intra prediction modes that can be selected as RemIntraPredMode are called "non - MPM" or "REM". intra_luma_mpm_remainder is used to derive RemIntraPredMode.
[0112] (Configuration of the Moving Picture Encoding Device)
[0113] The configuration of the moving picture decoding device 31 of the present embodiment ( Figure 7 ) will be described.
[0114] The moving image decoding device 31 is configured to include: an entropy decoding unit 301, a parameter decoding unit (predicted image decoding device) 302, a loop filter 305, a reference picture memory 306, a prediction parameter memory 307, a predicted image generation unit (predicted image generation device) 308, an inverse quantization / inverse transformation unit 311, and an addition unit 312. It should be noted that there is also a configuration in which, according to the moving image encoding device 11 described later, the loop filter 305 is not included in the moving image decoding device 31.
[0115] In addition, the parameter decoding unit 302 is configured to include an inter prediction parameter decoding unit 303 and an intra prediction parameter decoding unit 304 (not shown). The predicted image generation unit 308 is configured to include an inter predicted image generation unit 309 and an intra predicted image generation unit 310.
[0116] In addition, the following description uses examples in which CTU and CU are used as processing units, but is not limited to this example, and processing may also be performed in units of sub-CUs. Alternatively, CTU and CU may be replaced with blocks, and sub-CU may be replaced with sub-blocks, and processing may be performed in units of blocks or sub-blocks.
[0117] The entropy decoding unit 301 performs entropy decoding on the encoded stream Te input from the outside, and separates and decodes each encoded (syntax element). In entropy encoding, there are the following methods: a method of performing variable-length encoding on a syntax element using a context (probability model) adaptively selected according to the syntax element type or surrounding conditions; and a method of performing variable-length encoding on a syntax element using a pre-specified table or calculation formula. In the former CABAC (Context Adaptive Binary Arithmetic Coding), the probability model updated for each picture (slice) after encoding or decoding is stored in a memory. Then, among the probability models stored in the memory, the probability model of a picture using the quantization parameter of the same slice type and the same slice level is set as the initial state of the context of the P picture or B picture. This initial state is used for encoding and decoding processing. Among the separated encodings, there are prediction information for generating a predicted image and prediction errors for generating a differential image, etc.
[0118] The entropy decoding unit 301 outputs the separated encoding to the parameter decoding unit 302. It controls which encoding to decode based on the instruction of the parameter decoding unit 302.
[0119] (Configuration of the intra prediction parameter encoding unit 304)
[0120] The intra prediction parameter decoding unit 304 decodes intra prediction parameters such as the intra prediction mode IntraPredMode based on the coding input from the entropy decoding unit 301, with reference to the prediction parameters stored in the prediction parameter memory 307. The intra prediction parameter decoding unit 304 outputs the decoded intra prediction parameters to the predicted image generation unit 308 and stores them in the prediction parameter memory 307. The intra prediction parameter decoding unit 304 may also derive intra prediction modes with different luminances and color differences.
[0121] Figure 8 It is a schematic diagram showing the configuration of the intra prediction parameter decoding unit 304 of the parameter decoding unit 302. As Figure 8 shown, the intra prediction parameter decoding unit 304 is configured to include: a parameter decoding control unit 3041, a luminance intra prediction parameter decoding unit 3042, and a color difference intra prediction parameter decoding unit 3043.
[0122] The parameter decoding control unit 3041 instructs the entropy decoding unit 301 to decode syntax elements and receives the syntax elements from the entropy decoding unit 301. When intra_luma_mpm_flag is 1, the parameter decoding control unit 3041 outputs intra_luma_mpm_idx to the MPM parameter decoding unit 30422 in the luminance intra prediction parameter decoding unit 3042. In addition, when intra_luma_mpm_flag is 0, the parameter decoding control unit 3041 outputs intra_luma_mpm_remainder to the non-MPM parameter decoding unit 30423 of the luminance intra prediction parameter decoding unit 3042. In addition, the parameter decoding control unit 3041 outputs the syntax elements of the color difference intra prediction parameters to the color difference intra prediction parameter decoding unit 3043.
[0123] The luminance intra prediction parameter decoding unit 3042 is configured to include: an MPM candidate list derivation unit 30421, an MPM parameter decoding unit 30422, and a non-MPM parameter decoding unit 30423 (decoding unit, derivation unit).
[0124] The MPM parameter decoding unit 30422 derives IntraPredModeY with reference to mpmCandList[] derived by the MPM candidate list derivation unit 30421 and intra_luma_mpm_idx, and outputs it to the intra prediction image generation unit 310.
[0125] The non-MPM parameter decoding unit 30423 derives RemIntraPredMode from mpmCandList[] and intra_luma_mpm_remainder, and outputs IntraPredModeY to the intra prediction image generation unit 310.
[0126] The chrominance intra prediction parameter decoding unit 3043 derives IntraPredModeC from the syntax elements of the chrominance intra prediction parameters and outputs it to the intra prediction image generation unit 310.
[0127] The loop filter 305 is a filter provided within the encoding loop that removes block distortion and ringing distortion to improve the image quality. The loop filter 305 performs filtering such as deblocking filtering, sample adaptive offset (SAO), and adaptive loop filtering (ALF) on the decoded image of the CU generated by the adder 312.
[0128] The reference picture memory 306 stores the decoded image of the CU generated by the adder 312 at positions predefined for each target picture and each target CU.
[0129] The prediction parameter memory 307 stores the prediction parameters at positions predefined for each CTU or each CU to be decoded. Specifically, the prediction parameter memory 307 stores the parameters decoded by the parameter decoding unit 302 and the prediction mode predMode separated by the entropy decoding unit 301, etc.
[0130] The prediction mode predMode, prediction parameters, etc. are input to the prediction image generation unit 308. In addition, the prediction image generation unit 308 reads the reference picture from the reference picture memory 306. The prediction image generation unit 308 generates a prediction image of a block or sub-block in the prediction mode indicated by the prediction mode predMode, using the prediction parameters and the read reference picture (reference picture block). Here, the reference picture block refers to a set of pixels on the reference picture (since it is usually rectangular, it is called a block), and is the area referred to for generating the prediction image.
[0131] (Intra prediction image generation unit 310)
[0132] When the prediction mode predMode indicates the intra prediction mode, the intra prediction image generation unit 310 performs intra prediction using the intra prediction parameters input from the intra prediction parameter decoding unit 304 and the reference pixels read from the reference picture memory 306.
[0133] Specifically, the intra prediction image generation unit 310 reads the adjacent blocks on the target picture that are within a predefined range from the target block from the reference picture memory 306. The predefined range refers to the adjacent blocks to the left, upper left, upper, and upper right of the target block, and the reference area varies according to the intra prediction mode.
[0134] The intra-frame prediction image generation unit 310 generates a prediction image of the target block with reference to the read decoded pixel values and the prediction mode indicated by IntraPredMode. The intra-frame prediction image generation unit 310 outputs the generated block prediction image to the addition unit 312.
[0135] Hereinafter, the generation of the prediction image based on the intra-frame prediction mode will be described. In Planar prediction, DC prediction, and Angular prediction, the decoded peripheral region adjacent (close) to the target prediction block is set as the reference region R. Then, the prediction image is generated by extrapolating the pixels on the reference region R in a specific direction. For example, the reference region R may also be set as an L-shaped region including the left and top of the target prediction block (or further including the upper left, upper right, and lower left) (e.g., Figure 9 the region indicated by the slant circularly marked pixels in Example 1 of the reference region).
[0136] (Details of the prediction image generation unit)
[0137] Next, Figure 10 Details of the configuration of the intra-frame prediction image generation unit 310 will be described. The intra-frame prediction image generation unit 310 includes: a reference sampling filter unit 3103 (second reference image setting unit), a prediction unit 3104, and a prediction image correction unit 3105 (prediction image correction unit, filter switching unit, weight coefficient change unit).
[0138] The prediction unit 3104 generates a temporary prediction image (pre-correction prediction image) of the target prediction block based on each reference pixel (reference image) on the reference region R, the filtered reference image generated by applying the reference pixel filter (first filter), and the intra-frame prediction mode, and outputs it to the prediction image correction unit 3105. The prediction image correction unit 3105 corrects the temporary prediction image according to the intra-frame prediction mode, generates a prediction image (post-correction prediction image), and outputs it.
[0139] Hereinafter, each unit included in the intra-frame prediction image generation unit 310 will be described.
[0140] (Reference sampling filter unit 3103)
[0141] The reference sampling filter unit 3103 derives the reference sampling s[x][y] at each position (x, y) on the reference region R with reference to the reference image. In addition, the reference sampling filter unit 3103 applies the reference pixel filter (first filter) to the reference sampling s[x][y] according to the intra-frame prediction mode, and updates the reference sampling s[x][y] at each position (x, y) on the reference region R (derives the filtered reference image s[x][y]). Specifically, a low-pass filter is applied to the position (x, y) and its surrounding reference images to derive the filtered reference image (Figure 9 Example 2 of the reference area). It should be noted that it is not necessary to apply the low-pass filter to all intra prediction modes, and the low-pass filter can also be applied to a part of the intra prediction modes. It should be noted that the filter applied to the reference image on the reference area R in the reference sampling filter unit 3103 is called "reference pixel filter (first filter)", and in contrast, the filter that corrects the temporary prediction image in the prediction image correction unit 3105 described later is called "position-dependent filter (second filter)".
[0142] (Configuration of Intra Prediction Unit 3104)
[0143] The intra prediction unit 3104 generates a temporary prediction image (temporary prediction pixel values, pre-correction prediction image) of the prediction target block based on the intra prediction mode, the reference image, and the filtered reference pixel values, and outputs it to the prediction image correction unit 3105. The prediction unit 3104 internally includes: a Planar prediction unit 31041, a DC prediction unit 31042, an Angular prediction unit 31043, an LM prediction unit 31044, and an MIP unit 31045. The prediction unit 3104 selects a specific prediction unit according to the intra prediction mode, and inputs the reference image and the filtered reference image. The relationship between the intra prediction mode and the corresponding prediction unit is as follows.
[0144] ·Planar prediction ··· Planar prediction unit 31041
[0145] ·DC prediction ··· DC prediction unit 31042
[0146] ·Angular prediction ··· Angular prediction unit 31043
[0147] ·LM prediction ··· LM prediction unit 31044
[0148] ·Matrix intra prediction ··· MIP unit 31045
[0149] (Planar Prediction)
[0150] The Planar prediction unit 31041 linearly adds the reference samples s[x][y] according to the distance between the prediction target pixel position and the reference pixel position to generate a temporary prediction image, and outputs it to the prediction image correction unit 3105.
[0151] (DC Prediction)
[0152] The DC prediction unit 31042 derives a DC prediction value equivalent to the average value of the reference samples s[x][y], and outputs a temporary prediction image q[x][y] with the DC prediction value as the pixel value.
[0153] (Angular Prediction)
[0154] The Angular prediction unit 31043 generates a temporary prediction image q[x][y] using the reference sample s[x][y] in the prediction direction (reference direction) indicated by the intra prediction mode, and outputs it to the prediction image correction unit 3105.
[0155] (LM Prediction)
[0156] The LM prediction unit 31044 predicts the chrominance pixel values based on the luminance pixel values. Specifically, it is a method of generating a prediction image of the chrominance image (Cb, Cr) using a linear model based on the decoded luminance image. As one of the LM predictions, there is CCLM (Cross-Component Linear Model prediction). CCLM prediction is a prediction method that uses a linear model for predicting chrominance from luminance for a block.
[0157] (MIP Embodiment 1)
[0158] Hereinafter, Figures 11 to 22 An example of the MIP process (Matrix-based intraprediction) performed by the MIP unit 31045 will be described. MIP is a technique for deriving a prediction image through the product-sum operation of a reference image derived from an adjacent image and a weight matrix. In the figure, the width of the target block is nTbW and the height is nTbH.
[0159] (1) Derivation of boundary reference pixels
[0160] The MIP unit derives a variable sizeId related to the size of the target block using the following formula ( Figure 18 ).
[0161] sizeId = (nTbW <= 4 && nTbH <= 4)? 0 : (nTbW <= 8 && nTbH <= 8)? 1 : 2 (MIP-1)
[0162] As Figure 18 shown, when the size of the target block (nTbW x nTbH) is 4×4, 8×8, 16×16, sizeId is 0, 1, 2 respectively. When it is 4×16 or 16×4, sizeId = 2.
[0163] Next, the MIP unit 31045 uses the sizeId to derive the number of MIP modes used numModes, the size boundarySize of the downsampled reference regions redT[] and redL[], the width and height predW and predH of the intermediate prediction image predMip[][], and the size predC of one side of the prediction image obtained during the prediction process for the weight matrix mWeight[predC*predC][inSize].
[0164] numModes = (sizeId == 0)? 35 : (sizeId == 1)? 19 : 11 (MIP-2)
[0165] boundarySize = (sizeId == 0)? 2 : 4
[0166] predW = (sizeId <= 1)? 4 : Min(nTbW, 8)
[0167] predH = (sizeId <= 1)? 4 : Min(nTbH, 8)
[0168] predC = (sizeId <= 1)? 4 : 8
[0169] In Figure 17 shows the relationship between the sizeId and the values of these variables.
[0170] The weight matrix is square (predC*predC), 4×4 when sizeId = 0 or sizeId = 1, and 8×8 when sizeId = 2. When the size of the weight matrix is different from the output size predW*predH of the intermediate prediction image (especially when predC > predW or predC > predH), as described later, the weight matrix is thinned out at intervals for reference. For example, in this embodiment, when the output size is 4×16 or 16×4, the weight matrix with size (predC) of 8 indicated by sizeId = 2 is selected, so cases where predW = 4 (< predC = 8) and predH = 4 (< predC = 8) are generated respectively. Since the size (predW*predH) of the intermediate prediction image needs to be less than or equal to the object block size nTbW*nTbH, when the object block size is small, when a larger weight matrix (predC*predC) is selected, processing to make the weight matrix conform to the intermediate prediction image size is required.
[0171] In addition, the MIP unit 31045 uses IntraPredMode to derive the transpose processing flag isTransposed. IntraPredMode is, for example,Figure 6 The in-frame prediction modes 0 to 66 shown.
[0172] isTransposed = (IntraPredMode > (numModes / 2))? 1 : 0
[0173] In addition, the number of reference pixels inSize used in the prediction using the weight matrix mWeight[predC * predC][inSize], and the width mipW and height mipH of the transformed intermediate prediction image predMip[][] are derived.
[0174] inSize = 2 * boundarySize - ((sizeId == 2)? 1 : 0)
[0175] mipW = isTransposed? predH : predW
[0176] mipH = isTransposed? predW : predH
[0177] The matrix reference pixel derivation unit of the MIP unit 31045 sets the pixel values predSamples[x][-1] (x = 0..nTbW - 1) of the block adjacent to the upper side of the target block in the first reference area refT[x]
[0178] (x = 0..nTbW - 1). In addition, the pixel values predSamples[-1][y] (y = 0..nTbH - 1) of the block adjacent to the left side of the target block are set in the first reference area refL[y]
[0179] (y = 0..nTbH - 1). Then, the MIP unit 31045 downsamples the first reference areas refT[x] and refL[y] to derive the second reference areas redT[x] (x = 0..boundarySize - 1) and redL[y]
[0180] (y = 0..boundarySize - 1). Since the same processing is performed on refT[] and refL[] during downsampling, they are hereinafter referred to as refS[i] (i = 0..nTbX - 1) and redS[i] (i = 0..boundarySize - 1).
[0181] The matrix reference pixel derivation unit performs the following processing on refT[] or refS[] substituting refL[] to derive redS[]. When refT is substituted into refS, nTbS = nTbW, and when refL is substituted into refS, nTbS = nTbH.
[0182]
[0183] Here, Σ is the sum of i=0 to i=bDwn-1.
[0184] Next, the matrix reference pixel derivation unit combines the second reference areas redL[] and redT[] and derives p[i] (i=0..2*boundarySize-1).
[0185]
[0186]
[0187] bitDepthY is the bit depth of brightness, for example, it can be 10 bits.
[0188] It should be noted that, when the above reference pixels cannot be referenced, the available reference pixel values are used in the same way as the existing intra prediction. When all reference pixels cannot be referenced, 1<<(bitDepthY-1) is used as the pixel value. isTransposed indicates whether the prediction direction is close to the vertical prediction, so when switching whether to store redL or redT in the first half of p[] according to isTransposed, the pattern of mWeight[][] can be halved.
[0189] (2) Predicted pixel derivation (matrix operation)
[0190] MIP Department 31045 Figure 11 In STEP 2 predicted pixel derivation (matrix operation), the intermediate predicted image predMip[][] of size predW*predH is derived by matrix operation on p[].
[0191] The weight matrix derivation unit of the MIP unit 31045 selects a weight matrix mWeight[predC*predC][inSize] from the matrix set with reference to sizeId and modeId.
[0192] First, the weight matrix derivation unit derives modeId using IntraPredMode. modeId is the intra prediction mode used in MIP.
[0193] modeId=IntraPredMode-((isTransposed==1)?(numModes / 2):0)
[0194] When sizeId=0, the weight matrix derivation unit selects mWeight
[16] [4] from the array WeightS0
[18]
[16] [4] storing the weight matrix with reference to modeId. When sizeId=1, mWeight
[16] [8] is selected from the array WeightS1
[10]
[16] [8] storing the weight matrix with reference to modeId. When sizeId=2, mWeight
[64] [7] is selected from the array WeightS2[6]
[64] [7] storing the weight matrix with reference to modeId. These are expressed by the following formula.
[0195]
[0196] Next, the weight matrix derivation unit refers to sizeId and modeId to derive the shift value sW and offset coefficient fO used in (MIP-7). ShiftS0
[18] , ShiftS1
[10] , and ShiftS2[6] are arrays for storing shift values, and OffsetS0
[18] , OffsetS1
[10] , and OffsetS2[6] are arrays for storing offset coefficients.
[0197]
[0198] The matrix prediction image derivation unit of the MIP unit 31045 performs a product-sum operation (MIP-7) on p[], thereby deriving predMip[][] of the size of mipW*mipH. Here, the elements of the weight matrix mWeight[][] are referenced for each position corresponding to predMip[][] to derive the intermediate prediction image. It should be noted that in this embodiment, when sizeId=2, sometimes the size of the weight matrix predC is larger than the size of predMip mipW or mipH. Therefore, the variables incW and incH are used to eliminate the weight matrix for reference.
[0199]
[0200]
[0201] Σ is the sum of i=0 to i=inSize-1.
[0202] When isTransposed=1, the positions of the upper reference pixel and the left reference pixel are replaced and stored in the input p[] of the product-sum operation, and the transformation is performed before the output predMip[][] of the product-sum operation is output to (3).
[0203]
[0204] (3) Predicted pixel derivation (linear interpolation)
[0205] When nTbW=predW and nTbH=predH, the matrix prediction image interpolation unit of the MIP unit 31045 copies predMip[][] to predsamples[][].
[0206] for(x=0;x <nTbW;x++)
[0207] for(y=0;y <nTbH;y++)
[0208] predSamples[x][y]=predMip[x][y]
[0209] In other cases (nTbW>predW or nTbH>predH), the matrix prediction image interpolation unit is Figure 11 In 3-1 of STEP 3 (step 3) prediction pixel derivation (linear interpolation), predMip[][] is stored in the prediction image predSamples[][] of size nTbW*nTbH. When predW, predH are different from nTbW, nTbH, the prediction pixel value is interpolated in 3-2.
[0210] (3-1) The matrix prediction image interpolation unit stores predMip[][] in predSamples[][]. That is, Figure 12 In the pre-interpolation image, predMip[][] is stored at the shadow pixel position in the upper right and lower left directions.
[0211]
[0212] (3-2) When nTbH>nTbW, the pixels not stored in (3-1) are complemented using the pixel values of the adjacent blocks in the horizontal and vertical order to generate a predicted image.
[0213] To implement horizontal interpolation, use predSamples[xHor][yHor] and predSamples[xHor+upHor][yHor]( Figure 12 The pixel value at the position indicated by “○” is derived from the shadow pixels of the horizontally interpolated image.
[0214]
[0215] After interpolation in the horizontal direction, use predSamples[xVer][yVer] and predSamples[xVer][yVer+upVer](Figure 12 The pixel value at the position indicated by “○” is derived from the shadow pixels of the vertically interpolated image.
[0216]
[0217]
[0218] When nTbH<=nTbW, the pixel values of the adjacent blocks are interpolated in the order of the vertical direction and the horizontal direction to generate a predicted image. The vertical and horizontal interpolation processes are the same as those in the case of nTbH>nTbW.
[0219] (MIP Example 2)
[0220] In this embodiment, an example is described in which the processing is simplified without reducing the coding efficiency compared with MIP Embodiment 1. The following description will focus on the changes, so the parts not described are the same processing as MIP Embodiment 1.
[0221] Figure 16 The structure of the MIP unit 31045 that represents a square matrix mWeight having a size not larger than the width nTbW and the height bTbH of the reference target block and derives a square intermediate prediction image predMip having the same size.
[0222] In this embodiment, when sizeId = 2, it is set to predW = predH = predC. The definition of sizeId is changed accordingly. Hereinafter, predW, predH, and predC are described as predSize.
[0223] (1) Boundary reference pixel extraction
[0224] The MIP unit uses the following formula to derive the variable sizeId related to the object block size: Figure 20 ).
[0225] sizeId=(nTbW<=4&&nTbH<=4)? 0: ((nTbW<=4||nTbH<=4)||
[0226] (nTbW==8&&nTbH==8))? 1:2(MIP-21)
[0227] For example, when the object block size is 4xN, Nx4 (N>4), or 8x8, sizeId is 1. If they are of the same category, the formula (MIP-21) may also be expressed in other ways, for example, as shown below.
[0228] sizeId=(nTbW<=4&&nTbH<=4)? 0: ((nTbW<=8&&nTbH<=8)||nTbW<=4||nTbH<=4)? 1:2(MIP-21)
[0229] As another example,
[0230] It may also be sizeId=(nTbW<=4&&nTbH<=4)? 0:((nTbW==8&&nTbH==8)||nTbW<=4||nTbH<=4)? 1:2(MIP-21). In addition, when the minimum size of the input block is 4×4, nTbW<=4 and nTbH<=4 may be replaced by nTbW==4 and nTbH==4, respectively.
[0231] Furthermore, when the block size to which the MIP is applied is limited, the MIP unit may also derive sizeId by other deriving methods. Figure 20 As shown, when MIP is applied only to blocks whose aspect ratio of the input block size is 4 times or less (Abs(Log2(nTbW)-Log2(nTbH))<=2), sizeId may be derived as follows instead of (MIP-21).
[0232] sizeId=(nTbW<=4&&nTbH<=4)? 0: (nTbW*nTbH<=64)? 1:2(MIP-21a)
[0233] Alternatively, it can be derived as follows using a logarithmic expression.
[0234] sizeId=(nTbW<=4&&nTbH<=4)? 0: (Log2(nTbW)+Log2(nTbH)<=6)? 1:2(MIP-21b)
[0235] When the block size to which the MIP is applied is limited, sizeId is derived using (MIP-21a) and (MIP-21b), thereby simplifying the process.
[0236] like Figure 20 As shown, in the present embodiment, in the case of 4×16 and 16×4, a matrix of size (predC) of 4 indicated by sizeId=1 is selected, so that the situation where predW and predH are smaller than the matrix size predC (=predSize) does not occur. The MIP unit 31045 of the present embodiment selects a matrix of size less than nTbW and bTbH (predC=predSize), that is, a matrix that satisfies the following formula.
[0237] predSize=predC<=min(nTbW,nTbH)
[0238] In this embodiment, the matrix size is 4×4 when sizeId=0 or 1, and 8×8 when sizeId=2. Therefore, the MIP unit 31045 selects "a matrix with sizeId=0 or sizeId=1 when one of nTbW and bTbH is 4". Such selection is limited to the following Figure 21 , Figure 22 The same is true in Chinese.
[0239] That is, the weight matrix derivation unit of the MIP unit 31045 derives a matrix of a size less than the width and less than the height of the object block size. In addition, the weight matrix derivation unit derives a matrix of a size of 4×4 when one side of the object block is 4. In addition, the weight matrix derivation unit derives a matrix of a size of 4×4 when the size of the object block is 4×16 and 16×4. In addition, the weight matrix derivation unit derives any one of the matrices indicated by sizeId=0, 1 of a size of 4×4 and the matrix indicated by sizeId=2 of a size of 8×8, and derives a matrix of sizeId=1 or 2 when one side of the object block is 4.
[0240] Next, the MIP unit 31045 uses sizeId to derive the number of MIP modes numModes, the size boundarySize of the downsampled reference areas redT[], redL[], the weight matrix mWeight, and the width and height predSize of the intermediate prediction image predMip[][].
[0241] numModes=(sizeId==0)? 35: (sizeId==1)? 19:11(MIP-22)
[0242] boundarySize=(sizeId==0)? 2:4
[0243] predSize=(sizeId<=1)? 4:8
[0244] exist Figure 19 The relationship between sizeId and the values of these variables is shown in FIG.
[0245] The derivation of isTransposed and inSize is the same as that in MIP embodiment 1.
[0246] The derivation of p[] and pTemp[] required for the derivation of the first reference region refT[], refL[], the second reference region redT[], redL[], and predMip is also the same as that in the first MIP embodiment.
[0247] (2) Predicted pixel derivation (matrix operation)
[0248] MIP Department 31045 Figure 11 In STEP 2 predicted pixel derivation (matrix operation), predMip[][] of size predSize*predSize is derived through matrix operation on p[].
[0249] The weight matrix derivation unit of the MIP unit 31045 selects a weight matrix mWeight[predSize*predSize][inSize] from the set of matrices with reference to sizeId and modeId.
[0250] The selection method of modeId and mWeight[][], and the derivation method of the shift value sW and the offset coefficient fO are the same as those in MIP embodiment 1.
[0251] The matrix prediction image derivation unit of the MIP unit 31045 derives predMip[][] of size predSize*predSize by performing a product-sum operation (MIP-23) on p[]. Here, in the classification of sizeId in this embodiment, mipW and mipH are always greater than predSize(predC). Therefore, incW and incH of Embodiment 1 are always 1, and the calculation process is omitted.
[0252]
[0253] Σ is the sum of i=0 to i=inSize-1.
[0254]
[0255]
[0256] (3) Predicted pixel derivation (linear interpolation)
[0257] When nTbW=predSize and nTbH=predSize, the matrix prediction image interpolation unit of the MIP unit 31045 copies predMip[][] to predsamples[][].
[0258] for(x=0;x <nTbW;x++)
[0259] for(y=0;y <nTbH;y++)
[0260] predSamples[x][y]=predMip[x][y]
[0261] In other cases (nTbW>predSize or nTbH>predSize), the matrix prediction image interpolation unit is Figure 11 In STEP3 prediction pixel derivation (linear interpolation), enlarge predMip[][] of predSize*predSize to prediction image predSamples[][] of size nTbW*nTbH. Copy the pixels at corresponding positions in 3-1, and derive the pixels at non-corresponding positions by interpolation in 3-2.
[0262] (3-1) The matrix prediction image interpolation unit stores predMip[][] in the corresponding position of predSamples[][]. That is, Figure 12 In the pre-interpolation image, predMip[][] is stored at the shadow pixel position of predSamples[][] at 3-1.
[0263]
[0264] (3-2) When nTbH>nTbW, the pixels not stored in (3-1) are interpolated using the pixel values of the adjacent blocks in the horizontal and vertical directions to generate a predicted image. The interpolation is performed in the order of the horizontal direction and the vertical direction, but it can also be performed in the order of the vertical direction and the horizontal direction.
[0265] To implement horizontal interpolation, use predSamples[xHor][yHor] and predSamples[xHor+upHor][yHor]( Figure 12 The pixel value at the position indicated by “○” is derived from the shadow pixels of the horizontally interpolated image.
[0266]
[0267] After interpolation in the horizontal direction, use predSamples[xVer][yVer] and predSamples[xVer][yVer+upVer]( Figure 12 The pixel value at the position indicated by “○” is derived from the shadow pixels of the vertically interpolated image.
[0268]
[0269] When nTbH<=nTbW, the pixel values of the adjacent blocks are interpolated in the order of the vertical direction and the horizontal direction to generate a predicted image. The vertical and horizontal interpolation processes are the same as those in the case of nTbH>nTbW.
[0270] The MIP unit 31045 of the second MIP embodiment is characterized in that it derives a square (predW=predH=predSize) intermediate prediction image predMip[][]. The derivation process is simplified to facilitate address calculation of the prediction image.
[0271] The MIP unit 31045 of the second MIP embodiment selects predSize that is smaller than the width nTbW and height nTbH of the target block, so that the matrix size predC (=predSize) selected by sizeId is equal to predW and predH, thereby facilitating reference of matrix elements in predMip derivation.
[0272] In MIP Example 2, by limiting the width and height of the prediction image classified as sizeId=2, the amount of calculation can be significantly reduced compared with MIP Example 1. Simulations have confirmed that there is almost no reduction in encoding efficiency due to these changes.
[0273] (MIP Example 3)
[0274] In this embodiment, another example is described in which the processing is simplified without reducing the coding efficiency compared with MIP Embodiment 1. The following description will focus on the changes, and the parts not described are the same processing as MIP Embodiment 2.
[0275] In this embodiment, when sizeId = 2, it is set to predW = predH = predC. Accordingly, the definition of sizeId is changed. Hereinafter, predW, predH, and predC are described as predSize.
[0276] (1) Boundary reference pixel extraction
[0277] The MIP unit uses the following formula to derive the variable sizeId related to the object block size: Figure 21 Figure above).
[0278] sizeId=(nTbW<=4&&nTbH<=4)? 0: (nTbW<=4||nTbH<=4)? 1:2(MIP-28)
[0279] Or you can use other conditions to determine sizeId( Figure 21 Figure below).
[0280] sizeId=(nTbW<=4&&nTbH<=4)? 0: (nTbW<=8||nTbH<=8)? 1:2(MIP-29)
[0281] (2) Predicted pixel derivation (matrix operation)
[0282] Same as MIP implementation example 2.
[0283] (3) Predicted pixel derivation (linear interpolation)
[0284] Same as MIP embodiment 2.
[0285] As described above, in the third embodiment of MIP, the determination of sizeId is further simplified compared with the second embodiment of MIP, thereby making it possible to further reduce the amount of calculation compared with the second embodiment of MIP.
[0286] It should be noted that MIP embodiment 3 is also the same as MIP embodiment 2. It derives a square (predW=predH=predSize) intermediate prediction image predMip[][], selects a predSize that is less than the width nTbW and height nTbH of the object block, and limits the width and height of the prediction image classified as sizeId=2, thereby achieving the same effect as MIP embodiment 2.
[0287] (MIP Example 4)
[0288] In this embodiment, another example of reducing the memory required for storing the weight matrix is described compared with MIP Embodiment 1. The following description will focus on the changes, and the parts not described are the same processing as MIP Embodiment 2.
[0289] In this embodiment, when sizeId = 2, it is set to predW = predH = predC. Accordingly, the definition of sizeId is changed. Hereinafter, predW, predH, and predC are described as predSize.
[0290] (1) Boundary reference pixel extraction
[0291] The MIP unit derives a variable sizeId related to the target block size using the following formula.
[0292] sizeId=(nTbW<=4||nTbH<=4)? 0:1(MIP-30)
[0293] In the above example, the value of sizeId is set to 0 or 1. Figure 22 As shown in the figure above, sizeId = (nTbW <= 4 || nTbH <= 4)? 0: 2 (MIP-34) or ifFigure 22 As shown in the figure below, sizeId = (nTbW <= 4 || nTbH <= 4)? 1:2 (MIP-34), then sizeId can be expressed as a combination of 0, 2 or 1, 2. It should be noted that it can also be set to
[0294] (nTbW<=8||nTbH<=8) replaces the conditional expression (nTbW<=4||nTbH<=4).
[0295] In the example of formula MIP-30, the value of sizeId is 0 and 1. Therefore, the processing of sizeId=2 in MIP embodiment 2 can be completely omitted. For example, the following formula is only needed to derive p[i] (i=0..2*boundarySize-1) from the second reference area redL[] and redT[].
[0296]
[0297] (2) Predicted pixel derivation (matrix operation)
[0298] It can also be set to be the same as MIP Example 2, but since sizeId=2 is not used, the process of selecting the weight matrix mWeight[predSize*predSize][inSize] from the matrix set with reference to sizeId and modeId omits the case of sizeId=2 and is expressed as the following formula.
[0299]
[0300]
[0301] Similarly, the process of deriving the shift value sW and the offset coefficient fO with reference to sizeId and modeId is expressed by the following formula.
[0302]
[0303] (3) Predicted pixel derivation (linear interpolation)
[0304] Same as MIP embodiment 2.
[0305] As described above, in MIP Example 3, the types of sizeId are reduced compared to MIP Example 2, so the memory required for storing the weight matrix can be reduced compared to MIP Example 2.
[0306] It should be noted that MIP embodiment 3 can also be the same as MIP embodiment 2, deriving a square (predW=predH=predSize) intermediate prediction image predMip[][], selecting a matrix (predSize) less than the object block size nTbWxnTbH, and limiting the width and height of the prediction image classified as sizeId=2, thereby achieving the same effect as MIP embodiment 2.
[0307] (Configuration of the Prediction Image Correction Unit 3105)
[0308] The predicted image correction unit 3105 corrects the temporary predicted image output from the prediction unit 3104 according to the intra-frame prediction mode. Specifically, the predicted image correction unit 3105 derives a position-dependent weight coefficient for each pixel of the temporary predicted image according to the reference area R and the position of the object prediction pixel. Then, by weighted addition (weighted averaging) of the reference sample s[][] and the temporary predicted image, a predicted image (corrected predicted image) Pred[][] that has been corrected for the temporary predicted image is derived. It should be noted that in some intra-frame prediction modes, the output of the prediction unit 3104 may be directly used as the predicted image without correcting the temporary predicted image using the predicted image correction unit 3105.
[0309] The inverse quantization / inverse transformation unit 311 inversely quantizes the quantized transform coefficients input from the entropy decoding unit 301 to obtain transform coefficients. The quantized transform coefficients are obtained by performing DCT on the prediction error in the encoding process.
[0310] The coefficients are obtained by frequency transformation such as DCT (Discrete Cosine Transform) and DST (Discrete Sine Transform) and quantization. The inverse quantization / inverse transformation unit 311 performs inverse frequency transformation such as inverse DCT and inverse DST on the obtained transform coefficients to calculate the prediction error. The inverse quantization / inverse transformation unit 311 outputs the prediction error to the addition unit 312.
[0311] The adder 312 adds the block prediction image input from the prediction image generator 308 and the prediction error input from the inverse quantization / inverse transform unit 311 for each pixel to generate a block decoded image. The adder 312 stores the block decoded image in the reference picture memory 306 and outputs it to the loop filter 305.
[0312] (Configuration of Moving Image Coding Device)
[0313] Next, the configuration of the moving picture encoding device 11 according to the present embodiment will be described. Figure 131 is a block diagram showing the structure of a moving picture coding apparatus 11 according to an example of the present embodiment. The moving picture coding apparatus 11 is configured to include a prediction image generation unit 101, a subtraction unit 102, a transformation / quantization unit 103, an inverse quantization / inverse transformation unit 105, an addition unit 106, a loop filter 107, a prediction parameter memory (prediction parameter storage unit, frame memory) 108, a reference image memory (reference image storage unit, frame memory) 109, a coding parameter determination unit 110, a parameter coding unit 111, and an entropy coding unit 104.
[0314] The predicted image generation unit 101 generates a predicted image for each region obtained by dividing each picture of the image T, that is, for each CU. The predicted image generation unit 101 operates in the same manner as the predicted image generation unit 308 described above, and the description thereof is omitted here.
[0315] The subtraction unit 102 generates a prediction error by subtracting the pixel value of the block prediction image input from the prediction image generation unit 101 from the pixel value of the image T. The subtraction unit 102 outputs the prediction error to the transformation / quantization unit 103 .
[0316] The transform / quantization unit 103 calculates transform coefficients by frequency transforming the prediction error input from the subtraction unit 102 , and derives quantized transform coefficients by quantization. The transform / quantization unit 103 outputs the quantized transform coefficients to the entropy coding unit 104 and the inverse quantization / inverse transform unit 105 .
[0317] The inverse quantization / inverse transformation unit 105 and the inverse quantization / inverse transformation unit 311 ( Figure 7 ) is the same as that of FIG. 106, and its description is omitted here. The calculated prediction error is output to the adding unit 106.
[0318] The entropy coding unit 104 receives quantized transform coefficients from the transform / quantization unit 103 and receives coding parameters from the parameter coding unit 111. The entropy coding unit 104 entropy codes partition information, prediction parameters, quantized transform coefficients, etc., generates a coded stream Te, and outputs it.
[0319] The parameter coding unit 111 includes a header coding unit 1110 (not shown), a CT information coding unit 1111 , a CU coding unit 1112 (prediction mode coding unit), an inter prediction parameter coding unit 112 , and an intra prediction parameter coding unit 113 . The CU coding unit 1112 further includes a TU coding unit 1114 .
[0320] (Configuration of Intra-frame Prediction Parameter Coding Unit 113)
[0321] The intra prediction parameter encoder 113 derives a format for encoding (eg, intra_luma_mpm_idx, intra_luma_mpm_remmainder, etc.) based on the IntraPredMode input from the encoding parameter determiner 110. The intra prediction parameter encoder 113 includes a configuration that is partially identical to the configuration in which the intra prediction parameter decoder 304 derives intra prediction parameters.
[0322] Figure 14 1 is a schematic diagram showing the configuration of the intra prediction parameter coding unit 113 of the parameter coding unit 111. The intra prediction parameter coding unit 113 includes a parameter coding control unit 1131, a luma intra prediction parameter derivation unit 1132, and a chroma intra prediction parameter derivation unit 1133.
[0323] IntraPredModeY and IntraPredModeC are input from the encoding parameter determination unit 110 to the parameter encoding control unit 1131. The parameter encoding control unit 1131 determines intra_luma_mpm_flag with reference to mpmCandList[] of the MPM candidate list derivation unit 30421. Then, intra_luma_mpm_flag and IntraPredModeY are output to the luma intra prediction parameter derivation unit 1132. IntraPredModeC is output to the chroma intra prediction parameter derivation unit 1133.
[0324] The luma intra prediction parameter derivation unit 1132 is configured to include an MPM candidate list derivation unit 30421 (candidate list derivation unit), an MPM parameter derivation unit 11322 and a non-MPM parameter derivation unit 11323 (encoding unit, derivation unit).
[0325] The MPM candidate list derivation unit 30421 derives mpmCandList[] with reference to the intra prediction mode of the adjacent block stored in the prediction parameter memory 108. The MPM parameter derivation unit 11322 derives intra_luma_mpm_idx from IntraPredModeY and mpmCandList[] when intra_luma_mpm_flag is 1, and outputs it to the entropy coding unit 104. The non-MPM parameter derivation unit 11323 derives RemIntraPredMode from IntraPredModeY and mpmCandList[] when intra_luma_mpm_flag is 0, and outputs intra_luma_mpm_remainder to the entropy coding unit 104.
[0326] The chroma intra prediction parameter derivation unit 1133 derives intra_chroma_pred_mode from IntraPredModeY and IntraPredModeC, and outputs it.
[0327] The adder 106 generates a decoded image by adding the pixel value of the block prediction image input from the prediction image generator 101 and the prediction error input from the inverse quantization / inverse transform unit 105 for each pixel. The adder 106 stores the generated decoded image in the reference picture memory 109.
[0328] The loop filter 107 performs deblocking filtering, SAO, and ALF on the decoded image generated by the adding unit 106. It should be noted that the loop filter 107 does not necessarily include the above three filters, and may be composed of only a deblocking filter, for example.
[0329] The prediction parameter memory 108 stores the prediction parameter generated by the encoding parameter determination unit 110 in a position predetermined for each target picture and each CU.
[0330] The reference picture memory 109 stores the decoded image generated by the loop filter 107 at a position predetermined for each target picture and each CU.
[0331] The coding parameter determination unit 110 selects one of a plurality of sets of coding parameters. The coding parameters refer to the QT, BT or TT split information, prediction parameters or parameters generated in association with these as the coding target. The predicted image generation unit 101 generates a predicted image using these coding parameters.
[0332] The coding parameter determination unit 110 calculates the RD cost value indicating the amount of information and the coding error for each of the plurality of sets. The coding parameter determination unit 110 selects the coding parameter set with the smallest calculated cost value. Thus, the entropy coding unit 104 outputs the selected coding parameter set as the coding stream Te. The coding parameter determination unit 110 stores the determined coding parameters in the prediction parameter memory 108.
[0333] It should be noted that a part of the moving picture encoding device 11 and the moving picture decoding device 31 in the above-mentioned embodiment, such as the entropy decoding unit 301, the parameter decoding unit 302, the loop filter 305, the predicted image generation unit 308, the inverse quantization / inverse transformation unit 311, the addition unit 312, the predicted image generation unit 101, the subtraction unit 102, the transformation / quantization unit 103, the entropy encoding unit 104, the inverse quantization / inverse transformation unit 105, the loop filter 107, the encoding parameter determination unit 110, and the parameter encoding unit 111, can also be implemented by a computer. In this case, the program for implementing the above-mentioned control function can be recorded in a computer-readable recording medium, and the computer system can read the program recorded in the recording medium and execute it. It should be noted that the "computer system" mentioned here refers to a computer system built into any one of the moving picture encoding device 11 and the moving picture decoding device 31, and a computer system including hardware such as an OS and peripheral devices is used. In addition, "computer-readable recording medium" refers to removable media such as floppy disks, magneto-optical disks, ROMs, CD-ROMs, and storage devices such as hard disks built into computer systems. Moreover, "computer-readable recording medium" may also include: recording media that dynamically store programs for a short period of time, such as communication lines in the case of sending programs via networks such as the Internet or communication lines such as telephone lines; and recording media that store programs for a fixed period of time, such as volatile memories inside computer systems that serve as servers or clients in this case. In addition, the above-mentioned program may be a program for realizing a part of the aforementioned functions, or a program that can realize the aforementioned functions by combining with a program already recorded in a computer system.
[0334] In addition, part or all of the moving picture encoding device 11 and the moving picture decoding device 31 in the above-mentioned embodiment may be implemented as an integrated circuit such as LSI (Large Scale Integration). Each functional block of the moving picture encoding device 11 and the moving picture decoding device 31 may be individually processor-based, or part or all may be integrated to form a processor. In addition, the method of integrated circuitization is not limited to LSI, and may be implemented by a dedicated circuit or a general-purpose processor. In addition, when a technology for integrated circuitization that replaces LSI appears with the advancement of semiconductor technology, the integrated circuit of this technology may also be used.
[0335] As mentioned above, one embodiment of the present invention has been described in detail with reference to the drawings, but the specific configuration is not limited to the above embodiment, and various design changes and the like can be made within the scope not departing from the gist of the present invention.
[0336] [Application Examples]
[0337] The above-mentioned moving picture encoding device 11 and moving picture decoding device 31 can be mounted on various devices for sending, receiving, recording, and reproducing moving pictures. It should be noted that the moving pictures can be natural moving pictures captured by a camera or the like, or artificial moving pictures (including CG and GUI) generated by a computer or the like.
[0338] First, refer to Figure 2 A case where the above-described moving picture encoding device 11 and moving picture decoding device 31 can be used for transmission and reception of moving pictures will be described.
[0339] exist Figure 2 2 is a block diagram showing the structure of a transmitting device PROD_A equipped with a motion picture encoding device 11. Figure 2 As shown, the transmitting device PROD_A comprises: a coding unit PROD_A1 for obtaining coded data by coding a moving picture; a modulating unit PROD_A2 for obtaining a modulated signal by modulating a carrier wave using the coded data obtained in the coding unit PROD_A1; and a transmitting unit PROD_A3 for transmitting the modulated signal obtained by the modulating unit PROD_A2. The moving picture coding device 11 is used as the coding unit PROD_A1.
[0340] The transmitting device PROD_A may further include a camera PROD_A4 for shooting moving images as a supply source of moving images to be input to the encoding unit PROD_A1, a recording medium PROD_A5 on which moving images are recorded, an input terminal PROD_A6 for inputting moving images from the outside, and an image processing unit A7 for generating or processing images. The figure illustrates that the transmitting device PROD_A includes all of these configurations, but some of them may be omitted.
[0341] It should be noted that the recording medium PROD_A5 may be a medium on which uncoded moving images are recorded, or may be a medium on which moving images are recorded after being coded using a coding method for recording that is different from the coding method for transmission. In the latter case, it is preferred that a decoding unit (not shown) that decodes coded data read from the recording medium PROD_A5 using the coding method for recording is interposed between the recording medium PROD_A5 and the coding unit PROD_A1.
[0342] In addition, Figure 2 2 is a block diagram showing the structure of a receiving device PROD_B equipped with a motion picture decoding device 31. Figure 2As shown, the receiving device PROD_B includes: a receiving unit PROD_B1 that receives a modulated signal; a demodulating unit PROD_B2 that demodulates the modulated signal received by the receiving unit PROD_B1 to obtain coded data; and a decoding unit PROD_B3 that decodes the coded data obtained by the demodulating unit PROD_B2 to obtain a moving image. The moving image decoding device 31 is used as the decoding unit PROD_B3.
[0343] The receiving device PROD_B may also include a display PROD_B4 for displaying moving images as a supply destination of the moving images output by the decoding unit PROD_B3, a recording medium PROD_B5 for recording the moving images, and an output terminal PROD_B6 for outputting the moving images to the outside. Figure 2 In the example, the receiving device PROD_B is shown to have all of these structures, but some of them may be omitted.
[0344] It should be noted that the recording medium PROD_B5 may be a medium for recording uncoded moving images, or may be a medium for recording moving images coded in a recording coding method different from a transmission coding method. In the latter case, it is preferred that an encoding unit (not shown) that encodes the moving images obtained from the decoding unit PROD_B3 in the recording coding method is interposed between the decoding unit PROD_B3 and the recording medium PROD_B5.
[0345] It should be noted that the transmission medium for transmitting the modulated signal can be wireless or wired. In addition, the transmission scheme for transmitting the modulated signal can be broadcast (here refers to a transmission scheme in which the transmission destination is not predetermined) or communication (here refers to a transmission scheme in which the transmission destination is predetermined). That is, the transmission of the modulated signal can be achieved through any one of wireless broadcasting, wired broadcasting, wireless communication and wired communication.
[0346] For example, a broadcasting station (broadcasting equipment, etc.) / receiving station (television receiver, etc.) of terrestrial digital broadcasting is an example of a transmitting device PROD_A / receiving device PROD_B that transmits and receives modulated signals via wireless broadcasting. In addition, a broadcasting station (broadcasting equipment, etc.) / receiving station (television receiver, etc.) of cable television broadcasting is an example of a transmitting device PROD_A / receiving device PROD_B that transmits and receives modulated signals via wired broadcasting.
[0347] In addition, a server (workstation, etc.) / client (TV receiver, personal computer, smart phone, etc.) of a VOD (Video On Demand) service or a motion picture sharing service using the Internet is an example of a transmitter PROD_A / receiver PROD_B that transmits and receives a modulated signal through communication (usually, either wireless or wired is used as a transmission medium in a LAN, and wired is used as a transmission medium in a WAN). Here, a personal computer includes a desktop PC, a laptop PC, and a tablet PC. In addition, a smart phone also includes a multi-function portable phone terminal.
[0348] It should be noted that the client of the moving image sharing service has the function of decoding the encoded data downloaded from the server and displaying it on the display, and also has the function of encoding the moving images captured by the camera and uploading them to the server. That is, the client of the moving image sharing service functions as both the sending device PROD_A and the receiving device PROD_B.
[0349] Next, refer to Figure 3 A case where the above-described moving image encoding device 11 and moving image decoding device 31 can be used for recording and reproducing moving images will be described.
[0350] exist Figure 3 2 is a block diagram showing the structure of a recording device PROD_C equipped with the above-mentioned motion picture encoding device 11. Figure 3 As shown, the recording device PROD_C includes: an encoding unit PROD_C1 that obtains encoded data by encoding a moving image; and a writing unit PROD_C2 that writes the encoded data obtained by the encoding unit PROD_C1 to the recording medium PROD_M. The moving image encoding device 11 described above is used as the encoding unit PROD_C1.
[0351] It should be noted that the recording medium PROD_M can be (1) a type of recording medium built into the recording device PROD_C such as a HDD (Hard Disk Drive) or SSD (Solid State Drive), or (2) a type of recording medium such as an SD memory card or a USB flash drive.
[0352] The recording medium may be a type of recording medium connected to the recording device PROD_C, such as a Universal Serial Bus (Universal Serial Bus) flash memory, or (3) a recording medium loaded into a drive device (not shown) built into the recording device PROD_C, such as a DVD (Digital Versatile Disc: registered trademark) or BD (Blu-ray Disc: registered trademark).
[0353] In addition, the recording device PROD_C may further include a camera PROD_C3 for shooting moving images as a supply source of moving images input to the encoding unit PROD_C1, an input terminal PROD_C4 for inputting moving images from the outside, a receiving unit PROD_C5 for receiving moving images, and an image processing unit PROD_C6 for generating or processing images. Figure 3 In the example, the recording device PROD_C is shown to have all of these structures, but some of them may be omitted.
[0354] It should be noted that the receiving unit PROD_C5 can receive uncoded moving images, or can receive coded data coded in a transmission coding method different from the recording coding method. In the latter case, it is preferable to place a transmission decoding unit (not shown) that decodes the coded data coded in the transmission coding method between the receiving unit PROD_C5 and the coding unit PROD_C1.
[0355] Examples of such a recording device PROD_C include a DVD recorder, a BD recorder, and a HDD (Hard Disk Drive) recorder (in this case, the input terminal PROD_C4 or the receiving unit PROD_C5 is the main source of motion images). In addition, a portable camera (in this case, the camera PROD_C3 is the main source of motion images), a personal computer (in this case, the receiving unit PROD_C5 or the image processing unit C6 is the main source of motion images), and a smart phone (in this case, the camera PROD_C3 or the receiving unit PROD_C5 is the main source of motion images) are also examples of such a recording device PROD_C.
[0356] In addition, Figure 3 2 is a block diagram showing the structure of a playback device PROD_D equipped with the above-mentioned motion picture decoding device 31. Figure 3 As shown, the playback device PROD_D includes a reading unit PROD_D1 for reading the coded data written in the recording medium PROD_M, and a decoding unit PROD_D2 for decoding the coded data read by the reading unit PROD_D1 to obtain a moving picture. The moving picture decoding device 31 described above is used as the decoding unit PROD_D2.
[0357] It should be noted that the recording medium PROD_M can be (1) a type of recording medium built into the reproduction device PROD_D such as an HDD or SSD, or (2) a type of recording medium connected to the reproduction device PROD_D such as an SD memory card or USB flash memory, or (3) a recording medium loaded into a drive device (not shown) built into the reproduction device PROD_D such as a DVD or BD.
[0358] In addition, the playback device PROD_D may further include a display PROD_D3 for displaying moving images as a supply destination of the moving images output by the decoding unit PROD_D2, an output terminal PROD_D4 for outputting the moving images to the outside, and a transmission unit PROD_D5 for transmitting the moving images. Figure 3 In the example, the reproduction device PROD_D is shown to have all of these structures, but some of them may be omitted.
[0359] It should be noted that the sending unit PROD_D5 can send uncoded moving images or coded data coded in a transmission coding method different from the recording coding method. In the latter case, it is better to place a coding unit (not shown) that encodes moving images in a transmission coding method between the decoding unit PROD_D2 and the sending unit PROD_D5.
[0360] Examples of such a reproduction device PROD_D include a DVD player, a BD player, and an HDD player (in this case, the output terminal PROD_D4 connected to a television receiver or the like is the main supply destination of the motion image). In addition, a television receiver (in this case, the display PROD_D3 is the main supply destination of the motion image), a digital signage (also called an electronic billboard, an electronic bulletin board, etc., the display PROD_D3 or the transmission unit PROD_D5 is the main supply destination of the motion image), a desktop PC (in this case, the output terminal PROD_D4 or the transmission unit PROD_D5 is the main supply destination of the motion image), a laptop or tablet PC (in this case, the display PROD_D3 or the transmission unit PROD_D5 is the main supply destination of the motion image), and a smartphone (in this case, the display PROD_D3 or the transmission unit PROD_D5 is the main supply destination of the motion image) are also examples of such a reproduction device PROD_D.
[0361] (Implemented in hardware and implemented in software)
[0362] Furthermore, each block of the video decoding device 31 and the video encoding device 11 may be implemented in hardware by a logic circuit formed on an integrated circuit (IC chip), or may be implemented in software by a CPU (Central Processing Unit).
[0363] In the latter case, each of the above-mentioned devices has: a CPU that executes commands of a program that realizes each function, a ROM (Read Only Memory) storing the above-mentioned program, a RAM (Random Access Memory) that expands the above-mentioned program, and a storage device (recording medium) such as a memory that stores the above-mentioned program and various data. Moreover, the purpose of the embodiment of the present invention can also be achieved in the following manner: a recording medium that records the program code (executable form program, intermediate code program, source program) of the control program of the above-mentioned each device in a computer-readable manner is supplied to each of the above-mentioned devices, and the computer (or CPU, MPU) reads and executes the program code recorded in the recording medium.
[0364] As the recording medium, for example, magnetic tapes such as magnetic tapes and cassette tapes, magnetic disks such as floppy disks (registered trademark) / hard disks, CD-ROMs (Compact Disc Read-Only Memory: optical disk read-only memory) / MO disks (Magneto-Optical disc: magneto-optical disc) / MD
[0365] (Mini Disc: Mini Disc) / DVD (Digital Versatile Disc: registered trademark) / CD-R
[0366] CD Recordable / Blu-ray Disc (Blu-ray Disc: registered trademark) and other optical disks; IC cards (including memory cards) / optical cards and other cards; Mask ROM / EPROM (Erasable Programmable Read-Only Memory: Erasable Programmable Read-Only Memory) / EEPROM
[0367] (Electrically Erasable and Programmable Read-Only Memory: Registered Trademark) / Flash ROM and other semiconductor memories; or PLD
[0368] Logic circuits such as FPGA (Field Programmable Gate Array) and Programmable Logic Device (PLD).
[0369] In addition, each of the above-mentioned devices may be configured to be connected to a communication network, and the program code may be supplied via the communication network. The communication network is not particularly limited as long as it can transmit the program code. For example, the Internet, an intranet, an extranet, a LAN (Local Area Network), an ISDN (Integrated Services Digital Network), a VAN
[0370] (Value-Added Network), CATV (Community Antenna television / Cable Television) communication network, virtual private network (Virtual Private Network), telephone line network, mobile communication network, satellite communication network, etc. In addition, the transmission medium that constitutes the communication network can be any medium that can transmit program code, and is not limited to a specific structure or type. For example, it can be used for wired networks such as IEEE (Institute of Electrical and Electronic Engineers) 1394, USB, power line transmission, cable TV line, telephone line, ADSL (Asymmetric Digital Subscriber Line: asymmetric digital subscriber line) line, etc., and can also be used for infrared such as IrDA (Infrared Data Association), remote control, BlueTooth (registered trademark), IEEE802.11 wireless, HDR (High Data Rate), NFC (Near Field Communication), DLNA (Digital Living Network Alliance
[0371] (Digital Living Network Alliance): registered trademark), mobile phone network, satellite line, terrestrial digital broadcast network and other wireless. It should be noted that the embodiments of the present invention can also be implemented in the form of a computer data signal embedded in a carrier wave that embodies the program code through electronic transmission.
[0372] The embodiments of the present invention are not limited to the above-mentioned embodiments, and various modifications can be made within the scope of the claims. That is, embodiments obtained by combining technical solutions appropriately modified within the scope of the claims are also included in the technical scope of the present invention.
[0373] Industrial Applicability
[0374] The embodiments of the present invention can be preferably applied to a motion picture decoding device that decodes coded data after encoding image data and a motion picture encoding device that generates coded data after encoding image data. In addition, it can be preferably applied to the data structure of coded data generated by the motion picture encoding device and referred to by the motion picture decoding device.
[0375] Explanation of symbols
[0376] 31: Image decoding device
[0377] 301: Entropy decoding unit
[0378] 302: Parameter decoding unit
[0379] 3020: Header decoding unit
[0380] 303: Inter-frame prediction parameter decoding unit
[0381] 304: Intra-frame prediction parameter decoding unit
[0382] 308: Prediction image generation unit
[0383] 309: Inter-frame prediction image generation unit
[0384] 310: Intra-frame prediction image generation unit
[0385] 311: Inverse quantization / inverse transformation unit
[0386] 312: Addition Department
[0387] 11: Image coding device
[0388] 101: Prediction image generation unit
[0389] 102: Subtraction Department
[0390] 103: Transformation / Quantization Unit
[0391] 104: Entropy coding unit
[0392] 105: Inverse quantization / inverse transformation unit
[0393] 107: Loop filter
[0394] 110: Coding parameter determination unit
[0395] 111: Parameter encoding unit
[0396] 112: Inter-frame prediction parameter encoding unit
[0397] 113: Intra-frame prediction parameter encoding unit
[0398] 1110: Header encoding unit
[0399] 1111: CT Information Coding Department
[0400] 1112: CU encoding unit (prediction mode encoding unit)
[0401] 1114: TU encoding unit
Claims
1. A moving image decoding apparatus for decoding an encoded image, the moving image decoding apparatus comprises: a weight matrix derivation unit configured to derive a weight matrix based on an intra prediction mode and a size index; a matrix prediction image derivation unit configured to derive a prediction image by performing a sum-of-products operation using sampled values derived by using the reference samples and elements in the weight matrix; and a matrix prediction image interpolation unit configured to derive a predicted image by using the prediction image, wherein if the size of the object block width and the object block height is equal to one of 4×8 pixels, 4×16 pixels, 4×32 pixels, 4×64 pixels, 8×8 pixels, 8×4 pixels, 16×4 pixels, 32×4 pixels, and 64×4 pixels, the size index is set to be equal to 1; and if the size of the object block width and the object block height is equal to one of 8×16 pixels, 8×32 pixels, 8×64 pixels, 16×8 pixels, 16×16 pixels, 16×32 pixels, 16×64 pixels, 32×8 pixels, 32×16 pixels, 32×32 pixels, 32×64 pixels, 64×8 pixels, 64×16 pixels, 64×32 pixels, and 64×64 pixels, the size index is set to be equal to 2.
2. A moving image encoding apparatus for encoding image data, the moving image encoding apparatus comprises: a weight matrix derivation unit configured to derive a weight matrix based on an intra prediction mode and a size index; a matrix prediction image derivation unit configured to derive a prediction image by performing a sum-of-products operation using sampled values derived by using the reference samples and elements in the weight matrix; and a matrix prediction image interpolation unit configured to derive a predicted image by using the prediction image, wherein if the size of the object block width and the object block height is equal to one of 4×8 pixels, 4×16 pixels, 4×32 pixels, 4×64 pixels, 8×8 pixels, 8×4 pixels, 16×4 pixels, 32×4 pixels, and 64×4 pixels, the size index is set to be equal to 1; and if the size of the object block width and the object block height is equal to one of 8×16 pixels, 8×32 pixels, 8×64 pixels, 16×8 pixels, 16×16 pixels, 16×32 pixels, 16×64 pixels, 32×8 pixels, 32×16 pixels, 32×32 pixels, 32×64 pixels, 64×8 pixels, 64×16 pixels, 64×32 pixels, and 64×64 pixels, the size index is set to be equal to 2.
3. A non-transitory computer-readable medium storing a bitstream generated by encoding moving image data, the bitstream being decoded by the following process: deriving a weight matrix from the bitstream based on an intra prediction mode and a size index; Deriving a predicted image by performing a product-sum operation of the sampled values derived using the reference sampling and the elements in the weight matrix; and Deriving a predicted image by using the predicted image, wherein, if the size of the object block width and the object block height is equal to one of 4×8 pixels, 4×16 pixels, 4×32 pixels, 4×64 pixels, 8×8 pixels, 8×4 pixels, 16×4 pixels, 32×4 pixels, and 64×4 pixels, the size index is set to be equal to 1; and if the size of the object block width and the object block height is equal to one of 8×16 pixels, 8×32 pixels, 8×64 pixels, 16×8 pixels, 16×16 pixels, 16×32 pixels, 16×64 pixels, 32×8 pixels, 32×16 pixels, 32×32 pixels, 32×64 pixels, 64×8 pixels, 64×16 pixels, 64×32 pixels, and 64×64 pixels, the size index is set to be equal to 2.