IMAGE FILTER DEVICE, IMAGE DECODING DEVICE, AND IMAGE ENCODING DEVICE

The image filter device uses a combination of neural networks to adaptively apply filters based on image characteristics, addressing inefficiencies in conventional methods by optimizing network scaling and filtering for improved encoding and decoding performance.

JP7681164B2Active Publication Date: 2025-05-21SHARP KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024105868
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-03-20
Filing Date
2024-07-01
Publication Date
2025-05-21
Estimated Expiration
2038-08-03

AI Technical Summary

Technical Problem

Conventional neural network filters for image encoding and decoding do not adapt to the characteristics of input image data, leading to inefficient network scaling and inability to apply suitable filters to different regions, resulting in suboptimal encoding and decoding performance.

Method used

The proposed image filter device employs a combination of individual and common neural networks that selectively apply filters based on image characteristics, using a first neural network for input data, a second for encoding parameters, a combiner for combined data, and an adder for output data, allowing for adaptive filtering.

Benefits of technology

This approach enables the application of filters tailored to image characteristics, reducing network scale and improving encoding and decoding efficiency by aligning filtering processes with image-specific traits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007681164000001
    Figure 0007681164000001
  • Figure 0007681164000002
    Figure 0007681164000002
  • Figure 0007681164000003
    Figure 0007681164000003
Patent Text Reader

Abstract

To provide an image filter device, an image decoding device, and an image coding device which apply a filter corresponding to image characteristics, to input image data.SOLUTION: In the image coding device, a convolution neutral network (CNN) filter device 107 as the image filter device comprises: a first CNN filter 107b1 which inputs a pre-filter image and extracts features such as directivity, activity, etc.; and a second CNN filter 107b2 which inputs a coding parameter and outputs a post-filter image. The second CNN filter 107b2 includes Concatenate layer which concatenates a result of image processing with the coding parameter.SELECTED DRAWING: Figure 20
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] One aspect of the present invention relates to an image filtering device, an image decoding device, and an image encoding device. [Background technology]

[0002] In order to efficiently transmit or record moving images, a moving image encoding device is used that generates encoded data by encoding moving images, and a moving image decoding device is used that generates a decoded image by decoding the encoded data.

[0003] Specific examples of video coding methods include the methods proposed in H.264 / AVC and High-Efficiency Video Coding (HEVC).

[0004] In such a video coding method, images (pictures) constituting a video are divided into slices obtained by dividing the images, coding tree units (CTUs) obtained by dividing the slices, coding tree units (CTUs) obtained by dividing the coding tree units, and so on. The coding unit (sometimes called a coding unit (CU)) that is to be encoded, and The coding unit is divided into blocks, called prediction units (PUs) and transform units (TUs), and the blocks are managed in a hierarchical structure, and the blocks are coded / decoded for each CU.

[0005] In such a video coding method, a predicted image is usually generated based on a locally decoded image obtained by coding / decoding an input image, and a prediction residual (sometimes called a "difference image" or "residual image") obtained by subtracting the predicted image from the input image (original image) is coded. Methods for generating a predicted image include inter-prediction and intra-prediction.

[0006] Moreover, Non-Patent Document 1 can be cited as a recent example of a video encoding and decoding technique.

[0007] In addition, we have developed a neural network called Variable-filter-size Residue-learning CNN (VRCNN). Non-patent document 2 is an example of a technology that uses a network. [Prior art documents] [Non-patent literature]

[0008] [Non-Patent Document 1] "Algorithm Description of Joint Exploration Test Model 6", JVET-F1001, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 31 March - 7 April 2017 [Non-Patent Document 2] "A Convolutional Neural Network Approach for Post-Processing in HEVC Intra Coding" Summary of the Invention [Problem to be solved by the invention]

[0009] However, the above-mentioned neural network filter technology only switches the entire network depending on the quantization parameter, and there is a problem that the network scale becomes large when applying a filter according to the characteristics of the input image data. Also, there is a problem that a filter suitable for encoding each region cannot be applied.

[0010] The present invention has been made in view of the above problems, and has an object to provide a method for applying a filter to input image data in accordance with image characteristics while suppressing the network scale, as compared with the conventional configuration. The purpose of this project is to realize the following: [Means for solving the problem]

[0011] In order to solve the above problems, the neural network filter of the image filter device of the present invention is characterized in including: (i) a first neural network that inputs a first image and outputs first output data; (ii) a second neural network that inputs encoding parameters and outputs second output data; (iii) a combiner that derives combined data from the first output data and the second output data; (iv) a third neural network that uses the combined data to output third output data; and (v) an adder that inputs the third output data and the first image and outputs fourth output data. In addition, the image filter device of the present invention is equipped with a neural network that receives as input one or more first type input image data whose pixel values ​​are luminance or chrominance, and one or more second type input image data whose pixel values ​​are values ​​corresponding to reference parameters for generating a predicted image and a differential image, and outputs one or more first type output image data whose pixel values ​​are luminance or chrominance.

[0012] In order to solve the above problems, the image filter device of the present invention comprises a plurality of individual neural networks and a common neural network, wherein the individual neural networks selectively act on input image data to the image filter device depending on the values ​​of filter parameters in the input image data, and the common neural network commonly acts on output image data of the individual neural networks regardless of the values ​​of the filter parameters.

[0013] In order to solve the above problems, the image filter device of the present invention comprises a plurality of individual neural networks and a common neural network, wherein the common neural network acts on input image data to the image filter device, and the individual neural networks selectively act on output image data of the common neural network depending on values ​​of filter parameters in the input image data. Effect of the Invention

[0014] Compared to the conventional configuration, it is possible to apply a filter to input image data according to image characteristics. [Brief description of the drawings]

[0015] [Figure 1] FIG. 2 is a diagram showing a hierarchical structure of data of an encoded stream according to the embodiment. [Diagram 2] 11 is a diagram illustrating patterns of PU division modes, where (a) to (h) respectively show partition shapes when the PU division modes are 2Nx2N, 2NxN, 2NxnU, 2NxnD, Nx2N, nLx2N, nRx2N, and NxN. [Diagram 3] FIG. 2 is a conceptual diagram showing an example of a reference picture and a reference picture list. [Figure 4] 1 is a block diagram showing a configuration of an image encoding device according to a first embodiment. [Diagram 5] FIG. 2 is a schematic diagram showing a configuration of an image decoding device according to the first embodiment. [Figure 6] 1 is a schematic diagram showing a configuration of an inter-prediction image generating unit of an image encoding device according to the present embodiment. [Figure 7] 1 is a schematic diagram showing a configuration of an inter-prediction image generating unit of an image decoding device according to this embodiment. [Figure 8] 2 is a conceptual diagram showing an example of input / output of the image filtering device according to the first embodiment. FIG. [Figure 9] 1 is a schematic diagram showing a configuration of an image filtering device according to a first embodiment. [Figure 10] FIG. 4 is a schematic diagram showing a modified example of the configuration of the image filtering device according to the first embodiment. [Figure 11] FIG. 11 is a diagram illustrating an example of a quantization parameter. [Figure 12] FIG. 11 is a diagram for explaining an example of a prediction parameter. [Figure 13] FIG. 13 is a diagram illustrating an example of intra prediction. [Figure 14] FIG. 13 is a diagram illustrating an example of intra-prediction parameters. [Figure 15] FIG. 11 is a diagram for explaining an example of divided depth information. [Figure 16] FIG. 13 is a diagram illustrating another example of divided depth information. [Figure 17] FIG. 11 is a diagram for explaining another example of a prediction parameter. [Figure 18] FIG. 11 is a schematic diagram showing the configuration of an image filtering device according to a second embodiment. [Figure 19] FIG. 13 is a conceptual diagram illustrating an example of an image filtering device according to a third embodiment. [Figure 20] FIG. 11 is a schematic diagram showing the configuration of an image filtering device according to a third embodiment. [Figure 21] FIG. 13 is a schematic diagram showing a modified example of the configuration of an image filtering device according to the third embodiment. [Figure 22] FIG. 13 is a conceptual diagram illustrating an example of an image filtering device according to a fourth embodiment. [Diagram 23] FIG. 13 is a conceptual diagram showing a modified example of an image filtering device according to the fourth embodiment. [Figure 24] FIG. 13 is a conceptual diagram illustrating an example of an image filtering device according to a fifth embodiment. [Diagram 25] FIG. 13 is a conceptual diagram showing a modified example of an image filtering device according to the sixth embodiment. [Figure 26] FIG. 13 is a block diagram showing a configuration of an image encoding device according to a seventh embodiment. [Figure 27] 1 is a schematic diagram showing a configuration of an image filtering device according to an embodiment of the present invention. [Figure 28]10 is a conceptual diagram showing an example of updating parameters of the image filtering device according to the embodiment. FIG. [Figure 29] FIG. 13 is a diagram showing a data structure for transmitting parameters. [Diagram 30] FIG. 13 is a block diagram showing a configuration of an image decoding device according to a seventh embodiment. [Diagram 31] 1 is a diagram showing the configuration of a transmitting device equipped with an image encoding device according to the present embodiment, and a receiving device equipped with an image decoding device, where (a) shows the transmitting device equipped with the image encoding device, and (b) shows the receiving device equipped with the image decoding device. [Diagram 32] 1 is a diagram showing the configuration of a recording device equipped with an image encoding device according to the present embodiment, and a playback device equipped with an image decoding device, where (a) shows the configuration of a recording device equipped with an image encoding device, and (b) shows the configuration of a playback device equipped with an image decoding device. [Diagram 33] 1 is a schematic diagram showing a configuration of an image transmission system according to an embodiment of the present invention. [Diagram 34] 5A to 5C are conceptual diagrams showing another example of input / output of the image filtering device according to the first embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0016] (First embodiment) Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0017] FIG. 33 is a schematic diagram showing the configuration of an image transmission system 1 according to this embodiment.

[0018] The image transmission system 1 is a system that transmits a code obtained by encoding an image to be encoded, decodes the transmitted code, and displays the image. The image transmission system 1 includes an image encoding device 11, a network 21, an image decoding device 31, and an image display device 41.

[0019] An image T representing an image of a single layer or multiple layers is input to the image coding device 11. A layer is a concept used to distinguish between multiple pictures when there is one or more pictures that make up a certain time period. For example, coding the same picture using multiple layers with different image quality or resolution is called scalable coding, and coding pictures from different viewpoints using multiple layers is called view scalable coding. In the case of performing prediction (inter-layer prediction, inter-view prediction), the coding efficiency is greatly improved. Also, even when no prediction is performed (simulcast), the coded data can be consolidated.

[0020] The network 21 transmits the coded stream Te generated by the image coding device 11 to the image decoding device 31. The network 21 may be the Internet, a wide area network (WAN), a local area network (LAN), or a is a combination of these. The network 21 is not necessarily limited to a two-way communication network, and may be a one-way communication network that transmits broadcast waves such as terrestrial digital broadcasting and satellite broadcasting. In addition, the network 21 may be replaced by a storage medium on which the encoded stream Te is recorded, such as a DVD (Digital Versatile Disc) or a BD (Blue-ray Disc).

[0021] The image decoding device 31 decodes each of the coded streams Te transmitted by the network 21, and generates one or more decoded images Td.

[0022] The image display device 41 displays all or part of one or more decoded images Td generated by the image decoding device 31. The image display device 41 includes, for example, a display device such as a liquid crystal display or an organic EL (Electro-luminescence) display. Also, in spatial scalable encoding and SNR scalable encoding, when the image decoding device 31 and the image display device 41 have high processing capabilities, an extended layer image with high image quality is displayed, and when they have only lower processing capabilities, a base layer image that does not require as high a processing capability and display capability as the extended layer is displayed.

[0023] <Operator> The operators used in this specification are described below.

[0024] >> is a right bit shift, << is a left bit shift, & is a bitwise AND, and | is a bitwise OR , and |= is an OR assignment operator).

[0025] x? y : z is a ternary operator that takes y when x is true (non-zero) and z when x is false (0). is.

[0026] Clip3(a, b, c) is a function that clips c to a value between a and b (inclusive). It returns a if c < a, b if c > b, and c otherwise (where a <= b).

[0027] <Structure of the encoded stream Te> Prior to a detailed description of the image encoding device 11 and the image decoding device 31 according to the present embodiment, the data structure of the encoded stream Te generated by the image encoding device 11 and decoded by the image decoding device 31 will be described.

[0028] Fig. 1 is a diagram showing a hierarchical structure of data in an encoded stream Te. The encoded stream Te illustratively includes a sequence and a number of pictures constituting the sequence. (a) to (f) in Fig. 1 each represent an encoded video sequence that defines a sequence SEQ. , a coded picture specifying a picture PICT, a coded slice specifying a slice S, The coding tree unit includes coding slice data that specifies the slice data, a coding tree unit included in the coding slice data, and a coding unit (CU) included in the coding tree unit. FIG.

[0029] (Coded Video Sequence) In a coded video sequence, an image decoding device is used to decode the sequence SEQ to be processed. The sequence SEQ is shown in (a) of Figure 1. As shown in the figure, the Video Parameter Set and the Sequence Parameter Set are The image data includes a sequence parameter set (SPS), a picture parameter set (PPS), a picture parameter set (PICT), and supplemental enhancement information (SEI). The value after the # indicates the layer ID. In FIG. 1, an example is shown in which coded data for layers #0 and #1, i.e., layers 0 and 1, exists, but the type and number of layers are not limited to this.

[0030] The video parameter set VPS is used to determine the number of layers of a video image. A set of coding parameters common to multiple video images and a set of coding parameters related to multiple layers and individual layers included in the video images are defined.

[0031] The sequence parameter set SPS is used to decode the target sequence. A set of coding parameters that 31 refers to is specified. For example, the width and height of a picture are specified. Note that there may be multiple SPSs. In that case, one of the multiple SPSs can be selected from the PPS. Select .

[0032] The picture parameter set PPS specifies the number of pictures to be decoded for each picture in the target sequence. A set of coding parameters to be referred to by the image decoding device 31 is defined. For example, the reference value of the quantization width (pic_init_qp_minus26) used for decoding a picture and the application of weighted prediction are defined. A flag (weighted_pred_flag) indicating the weighted pred is included. Note that there may be multiple PPSs. In this case, one of multiple PPSs is selected from each picture in the target sequence.

[0033] (Encoded Picture) In the coded picture, a set of data to be referred to by the image decoding device 31 in order to decode the picture PICT to be processed is defined. As shown in (b) of FIG. 1, the picture PICT is divided into slices S0 to S NS-1 where NS is the total number of slices contained in the picture PICT.

[0034] In the following, slices S0 to S NS-1 If there is no need to distinguish between The same applies to other data to which subscripts are added that are included in the coded stream Te described below.

[0035] (Coded Slice) In the case of the coded slice, the image decoding device 31 refers to the slice S to be processed in order to decode the slice S. A slice S is a set of data that is to be processed, as shown in Figure 1(c). The slice header SH and slice data SDATA.

[0036] The slice header SH includes a group of coding parameters to be referred to by the image decoding device 31 in order to determine a decoding method for the current slice. Slice type designation information (slice_type) that designates a slice type is an example of a coding parameter included in the slice header SH.

[0037] Slice types that can be specified by the slice type specification information include (1) an I slice that uses only intra prediction during encoding, (2) a P slice that uses unidirectional prediction or intra prediction during encoding, and (3) a B slice that uses unidirectional prediction, bidirectional prediction, or intra prediction during encoding.

[0038] The slice header SH may include a reference (pic_parameter_set_id) to a picture parameter set PPS included in the coded video sequence.

[0039] (Encoded slice data) In the case of coded slice data, image decoding is performed to decode the slice data SDATA to be processed. The slice data SDATA is a set of data that the encoding device 31 refers to. As shown in (d), it includes a coding tree unit (CTU), which is a fixed-size (e.g., 64x64) block that constitutes a slice and is also called a largest coding unit (LCU).

[0040] (coding tree unit) As shown in FIG. 1(e), a set of data to be referenced by the image decoding device 31 in order to decode the coding tree unit to be processed is specified. The coding tree unit is divided by recursive quadtree division. A node of the tree structure obtained by the recursive quadtree division is called a coding node (CN: Coding Node). An intermediate node of the quadtree is a coding node, and the coding tree unit itself is also specified as the top coding node. The CTU is a division If cu_split_flag is 1, the coding node CN is split into four coding nodes CN. If cu_split_flag is 0, the coding node CN is not split and is split into one coding node CN. It has a coding unit (CU) as a node. The coding unit CU is the terminal node of the coding node and is not divided any further. The coding unit CU is the basic unit of the coding process.

[0041] Furthermore, when the size of the coding tree unit CTU is 64x64 pixels, the size of the coding unit can be any of 64x64 pixels, 32x32 pixels, 16x16 pixels, and 8x8 pixels.

[0042] (Encoding Unit) As shown in (f) of Fig. 1, a set of data to be referenced by the image decoding device 31 in order to decode the coding unit to be processed is defined. Specifically, the coding unit is composed of a prediction tree, a transform tree, and a CU header CUH. The CU header defines a prediction mode, a partitioning method (PU partitioning mode), etc.

[0043] In the prediction tree, prediction information (reference picture index, motion vector, etc.) of each prediction unit (PU) obtained by dividing the coding unit into one or more is specified. In other words, the prediction unit is one or more non-overlapping areas constituting the coding unit. The prediction tree also includes one or more prediction units obtained by the above division. In the following, a prediction unit further divided into prediction units is called a "subblock". A subblock is composed of multiple pixels. When the size of the prediction unit and the subblock are equal, there is one subblock in the prediction unit. When the size of the prediction unit is larger than the size of the subblock, the prediction unit is divided into subblocks. For example, when the prediction unit is 8x8 and the subblock is 4x4, the prediction unit is divided into four subblocks, divided into two horizontally and two vertically.

[0044] The prediction process may be performed for each prediction unit (sub-block).

[0045] Roughly speaking, there are two types of division in a prediction tree: intra prediction and inter prediction. Intra prediction is a prediction within the same picture, and inter prediction refers to a prediction process performed between different pictures (for example, between display times or between layer images).

[0046] In the case of intra prediction, there are two division methods: 2Nx2N (the same size as the coding unit) and NxN.

[0047] In the case of inter prediction, the division method is the PU division mode (part_mode) of the encoded data. It is coded by 2Nx2N (same size as coding unit), 2NxN, 2NxnU, 2NxnD, Nx2N , nLx2N, nRx2N, and NxN. Note that 2NxN and Nx2N indicate 1:1 symmetrical division, while 2NxnU, 2NxnD, nLx2N, and nRx2N indicate 1:3 and 3:1 asymmetrical division. The PUs included in a CU are represented as PU0, PU1, PU2, and PU3, respectively.

[0048] Partition shapes (positions of PU division boundaries) in each PU division mode are specifically illustrated in (a) to (h) of FIG. 2. Partition shapes (positions of PU division boundaries) in each PU division mode are specifically illustrated in (a) of FIG. 2. (b), (c), and (d) show 2NxN, 2NxnU, and 2NxnD partitions (horizontal partitions), respectively. (e), (f), and (g) show Nx2N, nLx2N, and nRx2N partitions (vertical partitions), respectively. (h) shows NxN Horizontal and vertical partitions are collectively called rectangular partitions, and 2Nx2N and NxN are collectively called square partitions.

[0049] In addition, in the transform tree, the coding unit is divided into one or more transform units, and the position and size of each transform unit are specified. In other words, a transform unit is one or more non-overlapping regions that constitute the coding unit. The transform tree includes one or more transform units obtained by the above division.

[0050] The division in the transform tree includes a method of allocating an area of ​​the same size as the coding unit as the transform unit, and a method of recursive quad-tree division similar to the above-mentioned division of the CU.

[0051] The conversion process is carried out for each conversion unit.

[0052] (Prediction parameters) The predicted image of the prediction unit (PU) is calculated based on the prediction parameters associated with the PU. The prediction parameters include intra-prediction prediction parameters and inter-prediction prediction parameters. The following describes inter-prediction prediction parameters (inter-prediction parameters). The inter-prediction parameters are composed of prediction list use flags predFlagL0 and predFlagL1, reference picture indexes refIdxL0 and refIdxL1, and motion vectors mvL0 and mvL1. The prediction list use flags predFlagL0 and predFlagL1 are flags indicating whether or not reference picture lists called L0 list and L1 list are used, respectively, and the corresponding reference picture list is used when the value is 1. Note that in this specification, when the term "flag indicating whether XX is present" is used, a flag other than 0 (for example, 1) is XX, 0 is not XX, and 1 is treated as true and 0 is treated as false in logical negation, logical product, etc. (similarly below). However, in an actual device or method, other values ​​can be used as true and false values.

[0053] Syntax elements for deriving inter prediction parameters included in the encoded data include, for example, a PU partition mode part_mode, a merge flag merge_flag, a merge index merge_idx, an inter prediction identifier inter_pred_idc, a reference picture index refIdxLX, a prediction vector index mvp_LX_idx, and a difference vector mvdLX.

[0054] (See picture list) The reference picture list is a list of reference pictures stored in the reference picture memory 306. Figure 3 is a conceptual diagram showing an example of a reference picture and a reference picture list. In Figure 3(a), the rectangles are pictures, the arrows indicate the reference relationships of pictures, the horizontal axis indicates time, I, P, and B in the rectangles indicate intra-pictures, uni-predictive pictures, and bi-predictive pictures, respectively, and the number in the rectangles indicates the number of pictures. The letters indicate the decoding order. As shown in the figure, the decoding order of the pictures is I0, P1, B2, B3, B4, and the display order is I0, B3, B2, B4, P1. FIG. 3(b) shows an example of a reference picture list. A reference picture list is a list indicating candidates for reference pictures, and one picture (slice) may have one or more reference picture lists. In the example shown in the figure, the target picture B3 is , and has two reference picture lists: L0 list RefPicList0 and L1 list RefPicList1. When the target picture is B3, the reference pictures are I0, P1, and B2, and the reference picture has these pictures as elements. The reference picture index refIdxLX specifies which picture in the picture list is actually referenced. The figure shows an example in which reference pictures P1 and B2 are referenced by refIdxL0 and refIdxL1.

[0055] (Merge prediction and AMVP prediction) There are two methods of decoding (encoding) prediction parameters: merge prediction mode and AMVP (Adaptive Motion Vector Prediction) mode. The merge flag merge_flag is a flag for identifying between these modes. The merge prediction mode is a mode in which the prediction list usage flag predFlagLX (or inter prediction identifier inter_pred_idc), reference picture index refIdxLX, and motion vector mvLX are not included in the encoded data, but are derived from prediction parameters of nearby PUs that have already been processed. The AMVP mode is a mode in which the inter prediction identifier inter_pred_idc, reference picture index refIdxLX, and motion vector mvLX are included in the encoded data. Note that the motion vector mvLX is a prediction vector that identifies the prediction vector mvpLX. The vector is encoded as a vector index mvp_LX_idx and a difference vector mvdLX.

[0056] The inter prediction identifier inter_pred_idc is a value indicating the type and number of reference pictures, and takes one of the values ​​PRED_L0, PRED_L1, and PRED_BI. PRED_L0 and PRED_L1 indicate that reference pictures managed in the reference picture lists of the L0 list and L1 list, respectively, are used. PRED_BI indicates that two reference pictures are used (uni-prediction). (Bi-predictive BiPred), and uses reference pictures managed in an L0 list and an L1 list. A predicted vector index mvp_LX_idx is an index indicating a predicted vector, and a reference picture index refIdxLX is an index indicating a reference picture managed in a reference picture list. Note that LX is a description method used when there is no distinction between L0 prediction and L1 prediction, and parameters for the L0 list and parameters for the L1 list are distinguished by replacing LX with L0 and L1.

[0057] The merge index merge_idx is the prediction parameter candidate derived from the PU for which processing has been completed. This is an index indicating which prediction parameter among the complements (merging candidates) is to be used as the prediction parameter of the decoding target PU.

[0058] (Motion Vector) The motion vector mvLX indicates the amount of displacement between blocks on two different pictures. The prediction vector and the difference vector regarding the motion vector mvLX are respectively denoted as the prediction vector mvpLX and the difference vector mvpLX. It's called KutormvdLX.

[0059] (Inter prediction identifier inter_pred_idc and prediction list usage flag predFlagLX) The relationship between the inter prediction identifier inter_pred_idc and the prediction list use flags predFlagL0 and predFlagL1 is as follows, and they can be converted into each other.

[0060] inter_pred_idc = (predFlagL1<<1) + predFlagL0 predFlagL0 = inter_pred_idc & 1 predFlagL1 = inter_pred_idc >> 1 In addition, the inter prediction parameter may use a prediction list usage flag or an inter prediction identifier. Furthermore, the determination using the prediction list usage flag may be replaced with a determination using the inter prediction identifier. Conversely, the determination using the inter prediction identifier may be replaced with a determination using the prediction list usage flag.

[0061] (Bi-predictive biPred decision) The flag biPred indicating whether or not the prediction is biPred can be derived based on whether or not two prediction list usage flags are both 1. For example, the flag can be derived by the following formula.

[0062] biPred = (predFlagL0 == 1 && predFlagL1 == 1) The flag biPred can also be derived based on whether the inter prediction identifier is a value indicating the use of two prediction lists (reference pictures). For example, the flag biPred can be derived by the following formula:

[0063] biPred = (inter_pred_idc == PRED_BI) ? 1 : 0 The above formula can also be expressed as the following formula:

[0064] biPred = (inter_pred_idc == PRED_BI) It should be noted that PRED_BI can use a value of 3, for example.

[0065] (Configuration of an image decoding device) Next, the configuration of the image decoding device 31 according to this embodiment will be described. Fig. 5 is a schematic diagram showing the configuration of the image decoding device 31 according to this embodiment. The image decoding device 31 includes an entropy decoding unit 301, a prediction parameter decoding unit (prediction image decoding device) 302, a CNN (Convolutional Neural Network) filter 305, a reference picture The image processing unit 302 includes a memory 306 , a prediction parameter memory 307 , a prediction image generating unit (prediction image generating device) 308 , an inverse quantization and inverse transform unit 311 , and an adder 312 .

[0066] The prediction parameter decoding unit 302 includes an inter prediction parameter decoding unit 303 and an intra prediction parameter decoding unit 304. The prediction image generating unit 308 includes an inter prediction image generating unit 309 and an intra prediction image generating unit 310.

[0067] The entropy decoding unit 301 performs entropy decoding on the externally input coded stream Te, and separates and decodes individual codes (syntax elements). The separated codes include prediction information for generating a predicted image and residual information for generating a difference image.

[0068] The entropy decoding unit 301 outputs a part of the separated codes to the prediction parameter decoding unit 302. The part of the separated codes includes, for example, a quantization parameter (QP), a prediction mode predMode, a PU partition mode part_mode, a merge flag merge_flag, a merge index merge_idx, an inter prediction identifier inter_pred_idc, a reference picture index refIdxLX, a prediction vector index mvp_LX_idx, and a difference vector mvdLX. Which codes are to be decoded is controlled by the following: The entropy decoding is performed based on an instruction from a prediction parameter decoding unit 302. The entropy decoding unit 301 outputs the quantized coefficients to an inverse quantization and inverse transform unit 311. The quantized coefficients are used to perform a discrete cosine transform (DCT), a discrete sine transform (DST), a Karyhnen Loeve Transform (KLT), or a dequantization transform (DLT) on the residual signal in the encoding process. These coefficients are obtained by performing a frequency transform such as the Nenloeve transform and then quantizing it.

[0069] In addition, the entropy decoding unit 301 performs a process of decoding a part of the separated code using a CNN filter 30 described later. The separated code is output to the decoder 5. The separated code part is, for example, a quantization parameter (QP), a prediction parameter, and depth information (division information).

[0070] The inter prediction parameter decoding unit 303 decodes inter prediction parameters based on the code input from the entropy decoding unit 301, with reference to the prediction parameters stored in the prediction parameter memory 307.

[0071] The inter-prediction parameter decoding unit 303 converts the decoded inter-prediction parameters into a predicted image The result is output to the generator 308 and stored in the predicted parameter memory 307 .

[0072] The intra prediction parameter decoding unit 304 decodes intra prediction parameters by referring to the prediction parameters stored in the prediction parameter memory 307, based on the code input from the entropy decoding unit 301. The intra prediction parameters are parameters used in a process of predicting a CU within one picture, such as an intra prediction mode IntraPredMode. The intra-prediction parameter decoding unit 304 outputs the decoded intra-prediction parameters to a predicted image generating unit 308 , and also stores them in a prediction parameter memory 307 .

[0073] The intra prediction parameter decoding unit 304 may derive different intra prediction modes for luminance and chrominance. In this case, the intra prediction parameter decoding unit 304 decodes a luminance prediction mode IntraPredModeY as a luminance prediction parameter, and a chrominance prediction mode IntraPredModeC as a chrominance prediction parameter. The luminance prediction mode IntraPredModeY has 35 modes, which correspond to planar prediction (0), DC prediction (1), and directional prediction (2 to 34). The chrominance prediction mode IntraPredModeC uses any of planar prediction (0), DC prediction (1), directional prediction (2 to 34), and LM mode (35). The intra-prediction parameter decoding unit 304 decodes a flag indicating whether IntraPredModeC is the same mode as the luma mode, and if the flag indicates that it is the same mode as the luma mode, assigns IntraPredModeY to IntraPredModeC, and if the flag indicates that it is a mode different from the luma mode, may decode planar prediction (0), DC prediction (1), directional prediction (2 to 34), or LM mode (35) as IntraPredModeC.

[0074] The CNN filter 305 receives the quantization parameters and the prediction The parameters are acquired, and the decoded image of the CU generated by the adder 312 is used as an input image (unfiltered image), and the unfiltered image is processed to output an output image (filtered image). The filter 305 has the same function as the CNN filter 107 provided in the image encoding device 11 described later. has.

[0075] The reference picture memory 306 stores the decoded image of the CU generated by the adder 312 in a location that is determined in advance for each picture and CU to be decoded.

[0076] The prediction parameter memory 307 stores prediction parameters in a predetermined position for each picture and prediction unit (or sub-block, fixed size block, pixel) to be decoded. Specifically, the prediction parameter memory 307 stores the inter prediction parameters decoded by the inter prediction parameter decoding unit 303, the intra prediction parameters decoded by the intra prediction parameter decoding unit 304, and the prediction mode predMode separated by the entropy decoding unit 301. The stored inter prediction parameters include, for example, a prediction list usage flag predFlagLX (inter prediction identifier inter_pred_idc), a reference picture index refIdxLX, and a motion vector mvLX.

[0077] The prediction mode predMode input from the entropy decoding unit 301 and prediction parameters input from the prediction parameter decoding unit 302 are input to the prediction image generation unit 308. The prediction image generation unit 308 also reads a reference picture from the reference picture memory 306. The prediction image generation unit 308 generates a prediction image of a PU or a sub-block using the input prediction parameters and the read reference picture (reference picture block) in the prediction mode indicated by the prediction mode predMode.

[0078] Here, when the prediction mode predMode indicates an inter prediction mode, the inter prediction image generation unit 309 generates a prediction image of a PU or sub-block by inter prediction using the inter prediction parameters input from the inter prediction parameter decoding unit 303 and the read reference picture (reference picture block).

[0079] The inter predicted image generating unit 309 reads out from the reference picture memory 306 a reference picture block located at a position indicated by a motion vector mvLX with respect to the decoding target PU, from a reference picture indicated by a reference picture index refIdxLX, for a reference picture list (L0 list or L1 list) in which a prediction list usage flag predFlagLX is 1. The inter predicted image generating unit 309 performs prediction based on the read reference picture block to generate a predicted image of the PU. The inter predicted image generating unit 309 outputs the generated predicted image of the PU to the addition unit 312. Here, the reference picture block is a set of pixels on the reference picture (usually rectangular, so called a block), and is an area referenced to generate a predicted image of the PU or sub-block.

[0080] When the prediction mode predMode indicates an intra prediction mode, the intra prediction image generating unit 310 performs intra prediction using the intra prediction parameters input from the intra prediction parameter decoding unit 304 and the read reference picture. Specifically, the intra prediction image generating unit 310 reads out from the reference picture memory 306 adjacent PUs that are in a predetermined range from the decoding target PU, which is a picture to be decoded and is already decoded. The predetermined range is, for example, one of the adjacent PUs on the left, upper left, upper, and upper right when the decoding target PU moves sequentially in a so-called raster scan order, and differs depending on the intra prediction mode. The raster scan order is an order in which each picture is moved sequentially from the left end to the right end for each row from the top end to the bottom end.

[0081] The intra-prediction image generation unit 310 performs prediction in a prediction mode indicated by the intra-prediction mode IntraPredMode based on the read neighboring PU, and generates a prediction image of the PU. The image generation unit 310 outputs the generated predicted image of the PU to the addition unit 312.

[0082] When the intra prediction parameter decoding unit 304 derives different intra prediction modes for luma and chroma, the intra prediction image generation unit 310 generates a prediction image for the luma PU by any one of planar prediction (0), DC prediction (1), or directional prediction (2 to 34) according to the luma prediction mode IntraPredModeY, and generates a prediction image for the chroma PU by any one of planar prediction (0), DC prediction (1), directional prediction (2 to 34), or LM mode (35) according to the chroma prediction mode IntraPredModeC.

[0083] The inverse quantization and inverse transform unit 311 inverse quantizes the quantized coefficients input from the entropy decoding unit 301 to obtain transform coefficients. The inverse quantization and inverse transform unit 311 performs inverse frequency transform such as inverse DCT, inverse DST, and inverse KLT on the obtained transform coefficients to calculate a residual signal. The unit 311 outputs the calculated residual signal to the addition unit 312 .

[0084] The adder 312 generates a decoded image of the PU by adding, for each pixel, the predicted image of the PU input from the inter predicted image generation unit 309 or the intra predicted image generation unit 310 and the residual signal input from the inverse quantization and inverse transform unit 311. The adder 312 stores the generated decoded image of the PU in the reference picture memory 306, and outputs to the outside a decoded image Td obtained by integrating the generated decoded images of the PU for each picture.

[0085] (Inter-prediction image generation unit 309) 7 is a schematic diagram showing the configuration of an inter-prediction image generation unit 309 included in the prediction image generation unit 308 according to this embodiment. The inter-prediction image generation unit 309 includes a motion compensation unit (prediction image generation device) 3091 and a weighted prediction unit 3094.

[0086] (Motion Compensation) The motion compensation unit 3091 receives the inter prediction parameter from the inter prediction parameter decoding unit 303. Based on prediction parameters (prediction list usage flag predFlagLX, reference picture index refIdxLX, motion vector mvLX), an interpolated image (motion compensation image predSamplesLX) is generated by reading a block located at a position shifted by the motion vector mvLX from the position of the decoding target PU in the reference picture RefX specified by the reference picture index refIdxLX from the reference picture memory 306. Here, the accuracy of the motion vector mvLX is integer accuracy. If not, a filter called a motion compensation filter is applied to generate pixels at decimal positions to generate a motion compensated image.

[0087] (Weighted prediction) The weighted prediction unit 3094 multiplies the input motion-compensated image predSamplesLX by a weighting factor. This generates a predicted image for the PU.

[0088] (Configuration of Image Encoding Device) Next, the configuration of the image encoding device 11 according to this embodiment will be described. FIG. 4 is a block diagram showing the configuration of the image encoding device 11 according to this embodiment. The image encoding device 11 includes a predicted image generating unit 101, a subtraction unit 102, a transform / quantization unit 103, an entropy encoding unit 104, an inverse quantization / inverse transform unit 105, an addition unit 106, and a CNN (Convolutional Neural Network) , convolutional neural network) filter 107, a prediction parameter memory (prediction parameter storage unit, frame memory) 108, a reference picture memory (reference image storage unit, frame memory) 109, an encoding parameter determination unit 110, and a prediction parameter encoding unit 111. The prediction parameter encoding unit 111 is configured to include an inter-prediction parameter encoding unit 112 and an intra-prediction parameter encoding unit 113.

[0089] For each picture of the image T, the prediction image generation unit 101 generates a prediction image P of a prediction unit PU for each coding unit CU, which is an area obtained by dividing the picture. Here, the prediction image generation unit 101 reads out a decoded block from the reference picture memory 109 based on a prediction parameter input from the prediction parameter coding unit 111. In the case of inter prediction, the prediction parameter input from the prediction parameter coding unit 111 is, for example, a motion vector. The prediction image generation unit 101 reads out a block located at a position on the reference image indicated by the motion vector starting from the target PU. In the case of intra prediction, the prediction parameter is, for example, an intra prediction mode. The pixel values ​​of adjacent PUs used in the intra prediction mode are read out from the reference picture memory 109, and a prediction image P of the PU is generated. The prediction image generation unit 101 reads out The predicted image generation unit 101 generates a predicted image P of the PU using one of the plurality of prediction methods for the reference picture block thus generated. The predicted image generation unit 101 outputs the generated predicted image P of the PU to the subtraction unit 102.

[0090] The predicted image generating unit 101 operates in the same manner as the predicted image generating unit 308 already described. For example, FIG. 6 is a schematic diagram showing a configuration of an inter-predicted image generating unit 1011 included in the predicted image generating unit 101. The inter-predicted image generating unit 1011 includes a motion compensation unit 10111 and a weighted prediction unit 10112. The motion compensation unit 10111 and the weighted prediction unit 10112 have the same configurations as the above-mentioned motion compensation unit 3091 and weighted prediction unit 3094, respectively, and therefore will not be described here.

[0091] The prediction image generating unit 101 generates a prediction image P of the PU based on pixel values ​​of a reference block read from a reference picture memory, using parameters input from the prediction parameter encoding unit. The predicted image generated by the predicted image generation unit 101 is output to a subtraction unit 102 and an addition unit 106.

[0092] The subtraction unit 102 subtracts the signal value of the predicted image P of the PU input from the predicted image generation unit 101 from the pixel value of the corresponding PU of the image T to generate a residual signal. The residual signal thus obtained is output to the transform / quantization unit 103 .

[0093] The transform / quantization unit 103 performs frequency transform on the residual signal input from the subtraction unit 102 to calculate transform coefficients. The transform / quantization unit 103 quantizes the calculated transform coefficients to obtain quantization coefficients. The transform / quantization unit 103 outputs the obtained quantization coefficients to the entropy coding unit 104 and the inverse quantization / inverse transform unit 105.

[0094] The entropy coding unit 104 receives the quantization coefficients from the transform / quantization unit 103 and the coding parameters from the prediction parameter coding unit 111. The input coding parameters include, for example, a quantization parameter, depth information (division information), a reference picture index refIdxLX, a prediction vector index mvp_LX_idx, a difference vector mvdLX, a prediction model The code includes the predMode, the merge index merge_idx, and other codes.

[0095] The entropy coding unit 104 entropy codes the input quantized coefficients and coding parameters to generate a coded stream Te, and outputs the generated coded stream Te to the outside.

[0096] The inverse quantization and inverse transform unit 105 inverse quantizes the quantized coefficients input from the transform and quantization unit 103 to obtain transform coefficients. The inverse quantization and inverse transform unit 105 performs inverse frequency transform on the obtained transform coefficients to calculate a residual signal. The inverse quantization and inverse transform unit 105 outputs the calculated residual signal to the addition unit 106.

[0097] The adder 106 generates a decoded image by adding, for each pixel, the signal value of the predicted image P of the PU input from the predicted image generation unit 101 and the signal value of the residual signal input from the inverse quantization and inverse transform unit 105. The adder 106 stores the generated decoded image in the reference picture memory 109.

[0098] (Configuration of Image Filter Device) The CNN filter 107 is an example of an image filter device according to the present embodiment. The image filter device according to the present embodiment functions as a filter that acts on a locally decoded image. The image filter device according to the present embodiment includes a neural network that receives one or more first-type input image data whose pixel values ​​are luminance or chrominance, and one or more second-type input image data whose pixel values ​​are values ​​according to reference parameters for generating a predicted image and a difference image, and outputs one or more first-type output image data whose pixel values ​​are luminance or chrominance.

[0099] Here, the reference parameters in this specification refer to parameters that are referenced to generate a predicted image and a differential image, and may include the above-mentioned coding parameters as an example. A specific example of the reference parameters is as follows.

[0100] Quantization parameter for the image on which the image filter device operates (hereinafter also called the input image) - Parameter indicating the type of intra prediction or inter prediction for the input image - Parameter indicating the intra prediction direction in the input image (intra prediction mode) -Parameter indicating reference picture for inter prediction in input image A parameter that indicates the partition depth in the input image. A parameter that indicates the size of the partition in the input image. Note that the reference parameters may also be simply called parameters unless there is some confusion. Also, the reference parameters may be explicitly transmitted in the encoded data.

[0101] The CNN filter 107 converts the decoded image data generated by the adder 106 into a first type input image The image filter device according to this embodiment is input as first-type (unfiltered image) data, processes the unfiltered image, and outputs first-type output image (filtered image) data. The image filter device according to this embodiment can also obtain quantization parameters and prediction parameters from the prediction parameter encoding unit 111 or the entropy decoding unit 301 as second-type input image data, and process the unfiltered image. Here, the output image after filtering by the image filter device is expected to match the original image as closely as possible.

[0102] The image filter device has the effect of reducing coding distortion, that is, blocking distortion and ringing distortion.

[0103] Here, CNN is a general term for a neural network that has at least a convolution layer (a layer in which weighting coefficients and bias / offsets in product-sum operations do not depend on the position in a picture). The weighting coefficients are also called kernels. The CNN filter 107 is a convolutional In addition to the connection layer, there is a fully connected layer (FCN) in which weight calculation depends on the position in the picture. The CNN filter 107 may include a layer in which neurons in the layer are LCN (Locally Connected Networks) is a structure in which neurons are connected to only a portion of the inputs in the layer (in other words, neurons have a spatial location and are connected only to inputs close to that spatial location). In the CNN filter 107, the input samples to the convolution layer are The input size and the output size may be different. By increasing the amount of movement (step size) when moving the position where the convolution filter is applied to more than 1, it is possible to include a layer in which the output size is smaller than the input size. Also, a deconvolution layer in which the output size is larger than the input size can be included. The deconvolution layer may also be called a transposed convolution. The CNN filter 107 may also include a pooled convolution layer. It can include a Pooling layer, a DropOut layer, etc. The Pooling layer is a layer that divides a large image into small windows and obtains representative values ​​such as the maximum and average values ​​for each divided window, and the Dropout layer is a layer that adds randomness by setting the output to a fixed value (e.g. 0) according to the probability.

[0104] FIG. 8 is a conceptual diagram showing an example of input and output of the CNN filter 107. In the example shown in FIG. For this, the unfiltered image consists of a luminance (Y) channel, a first chrominance (Cb) channel, and The filtered image includes three image channels, including a second chrominance (Cr) channel, and one coding parameter (reference parameter) channel, including a quantization parameter (QP) channel. The filtered image also includes three image channels, including a processed luma (Y') channel, a processed chrominance (Cb') channel, and a processed chrominance (Cr') channel.

[0105] 8 is an example of the input and output of the CNN filter 107. For example, the configuration in which Y (luminance), the first chrominance (Cb), and the second chrominance (Cr) of the unfiltered image are inputted separately to the respective channels is also included in the configuration of this embodiment. The present embodiment naturally includes a configuration in which the first color difference (Cb) and the second color difference (Cr) are input separately to each channel. In addition, the input pre-filter image is not limited to the Y, Cb, and Cr channels, and may be, for example, R, G, and B channels or X, Y, and Z channels. The channels may be Y, CMYK, or luminance and color difference channels. The channels are not limited to, Cb, and Cr, but may be, for example, channels expressed as Y, U, V, Y, Pb, Pr, Y, Dz, Dx, or I, Ct, and Cp. FIG. 34 shows another example of the input and output of the CNN filter 107. This is a conceptual diagram.

[0106] In FIG. 34(a), the unfiltered image is divided into the channels of luminance (Y) and quantization parameter (QP). The first chrominance (Cb) and quantization parameter (QP) channels, the second chrominance (Cr) and quantization parameter (QP) channels, The input signal is divided into a luminance (Y) and quantization parameter (QP) channel and input to a CNN filter 107. The CNN filter 107 includes a CNN filter 107-1 that processes the luminance (Y) and quantization parameter (QP) channels and outputs (Y'), a CNN filter 107-2 that processes the first chrominance (Cb) and quantization parameter (QP) channel and outputs (U'), and a CNN filter 107-3 that processes the second chrominance (Cr) and quantization parameter (QP) channel and outputs (V'). Note that the reference parameter (encoding parameter) is not limited to the quantization parameter (QP), and one or more encoding parameters can be used. Also, the CNN filters 107 are each configured to process a different The CNN filter 107 may be configured to operate in different modes, but is not limited to the configuration including the CNN filter 107-1, the CNN filter 107-2, and the CNN filter 107-3 configured using different means (circuits or software). For example, the CNN filter 107 may be configured by one of a plurality of different means (circuits or software) and may operate in different modes.

[0107] In FIG. 34(b), the unfiltered image has luminance (Y) and quantization parameter (QP) The CNN filter 107 is divided into a luminance (Y) channel, a first chrominance (Cb) channel, a second chrominance (Cr) channel, and a quantization parameter (QP) channel. A CNN filter 107-4 processes the channel of the quantization parameter (QP) and outputs (Y'); The chrominance (Cb) of the first pixel, the chrominance (Cr) of the second pixel, and the quantization parameter (QP) are processed (U', V') and the CNN filter 107-5 which outputs the reference parameter ( The coding parameter (QP) is not limited to the quantization parameter (QP), and one or more coding parameters can be used. Also, the CNN filter 107 can be implemented by different means (circuits, software, etc.). For example, the CNN filter 107 may be configured using a plurality of different means (circuits, software, etc.). The apparatus may be configured to operate in different modes, with one of the means A) and B). This configuration is used when processing an image (input) in which luminance and reference parameters are interleaved, and when processing an image (input) in which the first chrominance, the second chrominance, and reference parameters are interleaved.

[0108] In the configuration shown in FIG. 34(b), the processing of luminance (Y) and the processing of color difference (first color difference (Cb) and The second chrominance (Cr) channel is processed by interleaving the two channels, which are processed by different CNN filters. In this configuration, even if the resolution of the luminance (Y) is different from that of the first chrominance (Cb) and the second chrominance (Cr), the interlace of the first chrominance (Cb) and the second chrominance (Cr) can be performed. The amount of calculation is not large in the leave. Also, since the CNN filter 107 can process the luminance (Y) and the chrominance (the first chrominance (Cb) and the second chrominance (Cr)) separately, the luminance and the chrominance can be processed in parallel. Also, since the amount of information increases, that is, when processing the chrominance, the first chrominance, the second chrominance, and the encoding parameters can be used simultaneously, the accuracy of the CNN filter 107 can be improved in the configuration shown in FIG. 34(b).

[0109] 9 is a schematic diagram showing an example of the configuration of the CNN filter 107 according to this embodiment. The CNN filter 107 includes a plurality of convX layers.

[0110] In this embodiment, the convX layer includes at least one of the following configurations: can be done.

[0111] (1) conv(x): A configuration that performs the filtering process (convolution) (2) act(conv(x)): A configuration that performs activation (nonlinear function, e.g., sigmoid, tanh, relu, elu, selu, etc.) after convolution. (3) batch_norm(act(conv(x))): Batch normalization after convolution and activation. Configuration for implementing input range normalization (4) act(batch_norm(conv(x))): Batch normalization between convolution and activation. Configuration for implementing input range normalization (5) Pooling: A configuration that compresses and downsizes information between conv layers Furthermore, the CNN filter 107 may be configured to include at least one of the following layers in addition to the convX layer.

[0112] (5) Pooling: A configuration that compresses and downsizes information between conv layers (6) add / sub: Element-by-element addition (including subtraction) (7) concatenate / stack: stacking multiple inputs to form a new larger input (8) fcn: Configuration that implements fully connected filters (9) lcn: Configuration for implementing partially connected filters 9, the CNN filter 107 includes three convX layers (conv1, conv2, conv3) and an add layer. The input unfiltered image has a size of (N1+N2)xH1xW1. Here, N1 is the number of channels in the image. For example, the unfiltered image has the luminance (Y) channel. If it contains only Y, Cb, and Cr channels, N1 is "1". If it contains Y, Cb, and Cr channels, N1 is "3". If it contains R, G, and B channels, N1 is "3". W1 is the width patch of the picture. where H1 is the height patch size of the picture. N2 indicates the number of channels of the coding parameters. For example, if the coding parameters include only the quantization parameter (QP) channel, N2 is "1". The structure with the add layer is This is a configuration in which residuals are predicted using a CNN filter, and is known to be particularly effective when the CNN layer is deep. Note that the number of add layers is not limited to one, and multiple add layers may be used, as in the case of a known configuration called ResNet, which stacks multiple layers to derive residuals.

[0113] As described later, the network may include branches, and may have a concatenate layer that bundles the branched inputs and outputs. For example, N1xH1xW1 data and N2xH1xW1 When the data above is concatenated, the result is (N1+N2)xH1xW1 data.

[0114] The first conv layer of the CNN filter 107, conv1, receives (N1+N2)xH1xW1 data as input. The second conv layer of the CNN filter 107, conv2, receives the data of Nconv1xH1xW1 and outputs the data of Nconv2xH1xW1. The third conv layer, conv3, of 107 receives Nconv2xH1xW1 data and computes N1xH1xW1 The add layer adds the N1xH1xW1 data output from the conv layer and the N1xH1xW1 pre-filtered image for each pixel, and outputs N1xH1xW1 data.

[0115] As shown in FIG. 9, the number of channels of a picture is determined by the number of channels processed by the CNN filter 107. By doing so, the number of inputs is reduced from N1+N2 to N1. In this embodiment, the CNN filter 107 processes in a channel-first (channel×height×width) data format, but may process in a channel-last (height×width×channel) data format.

[0116] In addition, the CNN filter 107 reduces the output size by the convolution layer. A deconvolution layer can be used to increase the output size and an autoencoder layer can be used to restore it. A deep network consisting of multiple convolution layers is sometimes called a DNN (Deep Neural Network). An image filter device can also have a Recurrent Neural Network (RNN) that re-inputs part of the network output into the network. In an RNN, the re-input information can be considered as the internal state of the network.

[0117] The image filter device can further combine multiple LSTMs (Long Short-Term Memories) and GRUs (Gated Recurrent Units) that utilize sub-networks of neural networks to control the updating and transmission of re-input information (internal state).

[0118] The quantization parameter (QP) channel is used as the coding parameter channel of the unfiltered image. In addition to the channel, the channel for partition information (PartDepth) and the channel for prediction mode information (PredMode) are also available. You can add a ner.

[0119] (Quantization Parameter (QP)) The quantization parameter (QP) is a parameter that controls the compression rate and image quality of an image. In this embodiment, the quantization parameter (QP) has a characteristic that the larger the value, the lower the image quality and the smaller the code amount, and the smaller the value, the higher the image quality and the larger the code amount. As the quantization parameter (QP), for example, a parameter that derives the quantization width of a prediction residual can be used.

[0120] As the quantization parameter (QP) for each picture, one representative quantization parameter (QP) of the frame to be processed can be input. For example, the quantization parameter (QP) can be specified by a parameter set applied to the target picture. Also, the quantization parameter (QP) can be calculated based on the quantization parameters (QP) applied to the components of the picture. Specifically, the quantization parameter (QP) can be calculated based on the average value of the quantization parameters (QP) applied to the slices.

[0121] Also, as the quantization parameter (QP) for each unit into which a picture is divided, the quantization parameter (QP) for each unit into which a picture is divided according to a predetermined criterion can be input. For example, the quantization parameter (QP) can be applied to each slice. Also, the quantization parameter (QP) can be applied to blocks within a slice. Also, the quantization parameter (QP) can be specified in area units (for example, each area obtained by dividing a picture into 16x9 pieces) that are independent of existing coding units. In this case, since the quantization parameter (QP) depends on the number of slices and the number of transform units, the value of the quantization parameter (QP) corresponding to the area becomes indefinite, and a CNN filter cannot be constructed. Therefore, the average value of the quantization parameter (QP) within the area is set to 1. One possible method is to use the quantization parameter (QP) at one position in the region as the representative value. Another possible method is to use the median or mode of the quantization parameters (QP) at multiple positions in the region as the representative value.

[0122] In addition, when inputting a specific number of quantization parameters (QPs), a list of quantization parameters (QPs) is generated so that the number of quantization parameters (QPs) is constant, and the CNN function is generated. For example, a method of creating a list of quantization parameters (QP) for each slice and creating a list of three quantization parameters (QP), the maximum value, the minimum value, and the median value, and inputting the list may be considered.

[0123] In addition, the quantization parameter (QP) to be applied to the component to be processed can be input as the quantization parameter (QP) for each component. Examples of this quantization parameter (QP) include the luma quantization parameter (QP) and the chroma quantization parameter (QP).

[0124] In addition, when applying the CNN filter on a block-by-block basis, the surrounding quantization parameter (QP) is In addition, the quantization parameter (QP) of the target block and the quantization parameters (QP) of the surrounding blocks may be input.

[0125] The CNN filter 107 can be designed according to the picture and the coding parameters. That is, the CNN filter 107 can be derived from image data such as directionality and activity. Since the CNN filter 107 can be designed according to not only the picture characteristics that can be output but also the encoding parameters, it is possible to realize a filter with a different strength for each encoding parameter. Therefore, since the present embodiment includes the CNN filter 107, the encoding parameter To process images according to coding parameters without introducing a different network for each image. can be done.

[0126] 10 is a schematic diagram showing a modified example of the configuration of the image filter device according to the present embodiment. As shown in FIG. 10, the CNN filter, which is the image filter device, does not include an add layer, but a convX layer. Even in this modification that does not include an add layer, the CNN filter outputs N1*H1*W1 data.

[0127] With reference to FIG. 11, an example of the reference parameter being a quantization parameter (QP) will be described. The quantization parameter (QP) shown in FIG. 11(a) is arranged in the unit area of ​​the transform unit (or the unit area has the same quantization parameter (QP)). FIG. 11(b) shows a case where the quantization parameter (QP) shown in FIG. 11(a) is input in the unit area such as a pixel. When input in pixel units, the quantization parameter (QP) directly corresponding to each pixel can be used for processing, and processing according to each pixel can be performed. In addition, the boundary of the transform unit can be found from the change position of the quantization parameter (QP), and information on whether the pixel is in the same transform unit or in a different adjacent transform unit can be used for filtering. In addition, not only the change in pixel value but also the magnitude of the change in the quantization parameter (QP) can be used. For example, information on whether the quantization parameter (QP) is flat, changes slowly, changes sharply, or changes continuously can be used. In addition, before inputting to the CNN filter 107, the average Normalization or standardization may be performed so that the value approaches 0 and the variance approaches 1. This also applies to coding parameters and pixel values ​​other than the quantization parameter.

[0128] 12 illustrates an example of the reference parameters being prediction parameters. The prediction parameters include information indicating intra prediction or inter prediction, and if it is inter prediction, a prediction mode indicating the number of reference pictures to be used for prediction.

[0129] The prediction parameters shown in (a) of FIG. 12 are arranged in units of coding units (prediction units). The prediction parameters shown in (a) of FIG. 12 are input in units of pixels, etc., as shown in (b) of FIG. 12. When input in units of pixels, as in the example shown in (b) of FIG. 11, prediction parameters that directly correspond spatially to each pixel can be used for processing, and processing according to each pixel can be performed. That is, prediction parameters of (x, y) coordinates can be used simultaneously with (R, G, B) and (Y, Cb, Cr) that are pixel values ​​of (x, y) coordinates. In addition, the boundary of coding units can be found from the change position of the prediction parameters, and information on whether a pixel is in the same coding unit or in a different adjacent coding unit can be used for filtering. In addition, not only the change in pixel value but also the magnitude of change in the prediction parameters can be used. For example, information on whether the prediction parameters are flat, change slowly, change sharply, or change continuously can be used. Note that the values ​​assigned to the prediction parameters are not limited to the example of numbers shown in (b) of FIG. 12, as long as values ​​close to prediction modes of similar nature are assigned. For example, "-2" may be assigned to intra-prediction, "2" may be assigned to uni-prediction, and "4" may be assigned to bi-prediction.

[0130] Definitions of prediction modes for intra prediction will be described with reference to FIG. 13. FIG. 13 shows the definitions of prediction modes. As shown in the figure, 67 types of prediction modes are defined for luminance pixels, and each prediction mode is identified by a number from "0" to "66" (intra prediction mode index). In addition, the following names are assigned to each prediction mode. That is, "0" is "Planar (planar prediction)" and "1" is "DC ( For chrominance pixels, "Planar (planar prediction)", "VER (vertical prediction)", "HOR (horizontal prediction)", "DC (DC prediction)", "VDIR (45 degree prediction)", LM prediction (chrominance prediction mode), and DM prediction (using the intra prediction mode for luma) can be used. LM prediction is a method of predicting chrominance based on luma prediction. In other words, LM prediction uses the correlation between luminance pixel values ​​and chrominance pixel values. This is a prediction.

[0131] 14 shows an example of intra prediction parameters in which the reference parameters are luminance pixels. The intra prediction parameters include values ​​of prediction parameters determined for each partition. The intra prediction parameters may include, for example, an intra prediction mode.

[0132] The intra prediction parameters shown in (a) of FIG. 14 are arranged in units of coding units (prediction units). The prediction parameters shown in (a) of FIG. 14 are input in units of pixels, etc., as shown in (b) of FIG. 14. When input in units of pixels, as in the example shown in (b) of FIG. 11, the intra prediction parameters directly corresponding to each pixel can be used for processing, and processing according to each pixel can be performed. In addition, the boundary of the coding units can be known from the change position of the intra prediction parameters, and information on whether the pixels are the same coding unit or adjacent different coding units can be used. In addition, not only the change in pixel value but also the magnitude of the change in the intra prediction parameters can be used. For example, information on whether the intra prediction parameters change slowly, abruptly, or continuously can be used.

[0133] 15 shows an example of the reference parameter being depth information (partition information). The depth information is determined for each partition according to the transform unit. The depth information is determined according to the number of divisions of the coding unit, for example, and corresponds to the size of the coding unit.

[0134] The depth information shown in (a) of FIG. 15 is arranged in coding units (prediction unit units). FIG. 15(b) shows a case where the depth information shown in (a) of FIG. 15 is input in unit areas such as pixels. As in the example shown in (b) of FIG. 11, when input in pixel units, the depth information directly corresponding to each pixel can be used for processing, and processing according to each pixel can be performed. In addition, the boundary of the coding unit can be known from the change position of the depth information, and information on whether the pixel is the same coding unit or an adjacent different coding unit can be used. In addition, not only the change in pixel value but also the magnitude of the change in the depth information can be used. For example, information on whether the depth information changes slowly, abruptly, or continuously can be used.

[0135] Also, size information indicating the horizontal and vertical sizes of a partition may be used instead of the depth information.

[0136] FIG. 16 shows an example in which the reference parameters are size information including the horizontal and vertical sizes of the partition. In the example shown in FIG. 16, two pieces of information, that is, information on the horizontal size and the vertical size, are input for each unit area. In the example of FIG. 16, the horizontal size W (width ) and log2(W), which is the logarithm of the vertical size H (height) plus a specified offset (-2). -2,log2(H)-2 is used as the reference parameter. For example, the partition sizes (W, H) of (1,1), (2,1), (0,0), (2,0), (1,2), and (0,1) are (8,8), (16,8), (4,4), (16,4), (8,16), and (4,8), respectively.

[0137] The partition size in the example shown in FIG. 16 can also be considered as a value (3-log2(D)) obtained by subtracting the logarithm of a value D indicating the number of times the transform unit is divided horizontally and vertically from a predetermined value.

[0138] The size information shown in (a) of Fig. 16 is arranged in units of transform blocks. (b) of Fig. 16 shows a case where the size information shown in (a) of Fig. 16 is input in units of pixels or other units. The example shown in (b) of Fig. 16 also has the same effect as the example shown in (b) of Fig. 15.

[0139] Fig. 17 shows another example in which the coding parameters are made up of a plurality of prediction parameters. In the example shown in Fig. 17, in addition to the prediction mode, reference picture information is included in the prediction parameters.

[0140] The prediction parameters shown in (a) of Fig. 17 are arranged in units of transform blocks. Fig. 17(b) shows a case where the prediction parameters shown in (a) of Fig. 17 are input in units of pixels or other units. The example shown in (b) of Fig. 17 also has the same effect as the example shown in (b) of Fig. 12.

[0141] (CNN filter training method) The CNN filter 107 learns using the training data and the error function.

[0142] As training data for the CNN filter 107, the above-mentioned unfiltered image, reference parameters, and A set of raw images can be input. Also, the filter output from the CNN filter 107 The filtered image is expected to minimize the error with the original image under certain reference parameters.

[0143] Furthermore, as the error function of the CNN filter 107, a function (for example, mean absolute error or mean square error) that evaluates the error of the filtered image filtered by the CNN filter 107 with respect to the original image can be used. In addition to the image error, the magnitude of the parameter can also be added to the error function as a regularization term. In regularization, the absolute value, squared value, or both of the parameters can be used (these are called lasso, ridge, and elasticnet, respectively).

[0144] Furthermore, as will be described later, in a method for transmitting CNN parameters, the code amount of the CNN parameters may be further added to the error of the error function.

[0145] It can also be trained using a separate CNN network that evaluates the quality of images. In this case, the output of the CNN filter 107 (Generator) to be evaluated is The inputs are input in series to the work (Discriminator) to minimize (or maximize) the evaluation value of the evaluation CNN network. It is also appropriate to train the evaluation CNN network at the same time as training the CNN filter 107. The method of training two networks, one for generation and one for evaluation, at the same time is called Generative Adversarial Networks (GAN).

[0146] The CNN filter 305 of the image decoding device 31 is trained in the same manner as the CNN filter 107 of the image encoding device 11. In a configuration in which the same CNN filter is used in both the image encoding device 11 and the image decoding device 31, the CNN parameters of the two CNN filters are are considered to be the same.

[0147] The prediction parameter memory 108 stores the prediction parameters generated by the coding parameter determination unit 110 in a location that is determined in advance for each picture and CU to be coded.

[0148] The reference picture memory 109 stores the decoded image generated by the CNN filter 107 as the encoding target. The picture and CU are stored in a predetermined location.

[0149] The coding parameter determination unit 110 selects one set from among a plurality of sets of coding parameters. The coding parameters are the above-mentioned prediction parameters and parameters to be coded that are generated in relation to the prediction parameters. The predicted image generation unit 101 generates a predicted image P of the PU using each of the sets of coding parameters.

[0150] The coding parameter determination unit 110 determines the size of the amount of information and the code for each of the multiple sets. The coding parameter determination unit 110 calculates a cost value indicating a quantization error. The cost value is, for example, the sum of the code amount and the value obtained by multiplying the squared error by a coefficient λ. The code amount is the information amount of the coding stream Te obtained by entropy coding the quantization error and the coding parameters. The squared error is the sum between pixels for the squared value of the residual value of the residual signal calculated in the subtraction unit 102. The coefficient λ is a real number greater than zero that is set in advance. The coding parameter determination unit 110 selects a set of coding parameters that minimizes the calculated cost value. As a result, the entropy coding unit 104 outputs the selected set of coding parameters to the outside as the coding stream Te, and does not output the set of coding parameters that was not selected. The coding parameter determination unit 110 stores the determined coding parameters in the prediction parameter memory 108.

[0151] The prediction parameter coding unit 111 derives a format for coding from the parameters input from the coding parameter determination unit 110, and outputs the format to the entropy coding unit 104. Deriving a format for coding means, for example, deriving a difference vector from a motion vector and a prediction vector. The prediction parameter coding unit 111 also derives parameters necessary for generating a prediction image from the parameters input from the coding parameter determination unit 110, and outputs the parameters to the prediction image generation unit 101. The parameters necessary for generating a prediction image are, for example, motion vectors in units of subblocks.

[0152] The inter prediction parameter coding unit 112 derives inter prediction parameters such as a difference vector based on the prediction parameters input from the coding parameter determination unit 110. The inter prediction parameter coding unit 112 includes, as a configuration for deriving parameters necessary for generating a prediction image to be output to the prediction image generation unit 101, a configuration that is partially the same as the configuration for the inter prediction parameter decoding unit 303 (see FIG. 5, etc.) to derive inter prediction parameters.

[0153] The intra-prediction parameter encoding unit 113 derives a format for encoding (for example, MPM_idx, rem_intra_luma_pred_mode, etc.) from the intra-prediction mode IntraPredMode input from the encoding parameter determination unit 110.

[0154] Second embodiment Another embodiment of the present invention will be described below with reference to FIG. 18. For the sake of convenience, the description of members having the same functions as those described in the above embodiment will be omitted. Various types of network configurations of the CNN filter are possible. The second embodiment shown in FIG. 18 shows an example of a CNN filter having a different network configuration from the network configuration (FIGS. 9 and 10) described in the first embodiment. has the same effect.

[0155] In this embodiment, as shown in FIG. 18, the CNN filter 107a includes two convX layers (convolution layers) conv1 and conv2, a pooling layer pooling, and a Deconv layer (deconvolution layer). The convX layers conv1 and conv2 are convolution layers, and the Deconv layer conv3 is a deconvolution layer. The pooling layer pooling is disposed between the convX layers conv2 and conv3.

[0156] The input pre-filter image has a size of (N1+N2)*H1*W1. Here, N1 indicates the number of channels of the image, W1 is the width patch size of the picture, and H1 is the height patch size of the picture. N2 indicates the number of channels of the coding parameters.

[0157] The first convX layer conv1 of the CNN filter 107a receives (N1+N2)*H1*W1 data as input and outputs Nconv1*H1*W1 data. The second convX layer conv2 of the CNN filter 107a receives Nconv1*H1*W1 data as input and outputs Nconv2*H1*W1 data. The pooling layer pooling after the convX layer conv2 receives Nconv2*H1*W1 data as input and outputs Nconv2*H2*W2 That is, the pooling layer pooling outputs the data output from the convX layer conv2. The data with height*width of H1*W1 is converted to data with size H2*W2. The Deconv layer conv3 after the pooling layer receives Nconv2*H2*W2 data and converts it to N1*H1*W1 data. That is, the Deconv layer conv3 outputs the height*width output from the pooling layer pooling. The data of size H2*W2 is converted back to data of size H1*W1. Transposed Convolution is used here.

[0158] In this embodiment, the CNN filter in the image decoding device is the same as that in the image encoding device. The CNN filter 107a has a similar function to that of the CNN filter 107a in the

[0159] In this embodiment, as in (a) of FIG. 34 in the first embodiment, the unfiltered image is divided into a luminance (Y) and quantization parameter (QP) channel, a first chrominance (Cb) and a quantization parameter (QP) channel, and a second chrominance (Cb) and a quantization parameter (QP) channel. Alternatively, the input may be divided into a channel for the parameter (QP) and a channel for the second chrominance (Cr) and quantization parameter (QP) and input to the CNN filter 107a. As shown in FIG. 34(b), the unfiltered image has luminance (Y) and quantization parameter (QP). The CNN filter 107a may be configured to perform a filter process on an image (input) in which luminance and reference parameters are interleaved, and a filter process on an image (input) in which the first chrominance, the second chrominance, and the reference parameters are interleaved. Note that the reference parameters (coding parameters) are not limited to the quantization parameter (QP), and the CNN filter 107a may perform a filter process on a single image (input) in which luminance and reference parameters are interleaved. The above coding parameters can be used.

[0160] According to the configuration of the second embodiment, a kind of autoencoder type network configuration is used in which data reduced in the convolution layer and the pooling layer is expanded in the transpooling layer, thereby making it possible to perform filtering processing taking into account higher-level conceptual features. In other words, when performing filtering processing according to encoding parameters, it is possible to change the filter strength taking into account higher-level conceptual features that integrate edges and colors.

[0161] (Third embodiment) Another embodiment of the present invention will be described below with reference to Figures 19 to 20. For the sake of convenience, the description of members having the same functions as those described in the above embodiment will be omitted.

[0162] In this embodiment, as shown in FIG. 19, the CNN filter 107b includes a first CNN filter 107b1 and a second CNN filter 107b2. An unfiltered image is input to the first CNN filter 107b1. The first CNN filter 107b1 is a filter that detects directionality and activity. The second CNN filter 107b2 receives the data processed by the first CNN filter 107b1 and a quantization parameter (QP) as an encoding parameter.

[0163] That is, the first CNN filter 107b1 is a second neural network The first type of input image data input to the second CNN filter 107b2 is set as an output image.

[0164] The second CNN filter 107b2 performs a filtering process to weight the extracted features. The second CNN filter 107b2 uses the coding parameters to determine how to weight the The second CNN filter 107b2 controls whether to perform filtering. will be output.

[0165] The first CNN filter 107b1 is different from the above-mentioned CNN filter 107 in that it is a luminance (Y) An unfiltered image consisting of three channels including the first channel, the first color difference (Cb) channel, and the second color difference (Cr) channel is input, and a filtered image consisting of three channels including luminance and two color differences is output. Note that the channels of the unfiltered image and the filtered image are not limited to Y, Cb, Cr, but may be R, G, B, and further alpha and depth may be added. You can also add this.

[0166] In addition to the quantization parameter (QP), the second CNN filter 107b2 also includes a prediction parameter Alternatively, other coding parameters such as a prediction parameter may be input in addition to the quantization parameter (QP). It is to be noted that the configuration of inputting other coding parameters such as a prediction parameter in addition to the quantization parameter (QP) is not limited to this embodiment, and is similar to other embodiments.

[0167] Also in this embodiment, as in (a) of FIG. 34 in the first embodiment, the pre-filter image has a luminance (Y) channel, a first chrominance (Cb) channel, and a second chrominance (Cr) channel. The CNN filter 107b1 may be configured to input the CNN filter 107b1. As shown in FIG. 34(b) in the first embodiment, the pre-filter image has a luminance (Y) channel and The first color difference (Cb) channel and the second color difference (Cr) channel are separated and input to the CNN filter 107b1. That is, the CNN filter b1 may be configured to receive the luminance and the reference parameter and a filter process may be performed on an image (input) in which the first chrominance, the second chrominance, and a reference parameter are interleaved. Note that the reference parameter (encoding parameter) is not limited to the quantization parameter (QP), and the CNN filter b1 may use one or more encoding parameters.

[0168] FIG. 20 is a schematic diagram showing the configuration of the CNN filter 107b according to the present embodiment. As shown in FIG. 1, the first CNN filter 107b1 includes two convX layers (conv1, conv2). The second CNN filter 107b2 includes two convX layers (conv3, conv4) and a Concatenate layer. Includes.

[0169] The conv1 of the first CNN filter 107b1 receives N1*H1*W1 data and outputs Nconv1*H1*W1 data. The conv2 of the first CNN filter 107b1 receives Nconv1*H1*W1 data and outputs Nconv2*H1*W1 data as the image processing result.

[0170] The second CNN filter 107b2's conv4 receives N2*H1*W1 data and outputs Nconv4*H1*W1 data. The second CNN filter 107b2's Concatenate layer receives the image processing result, the coding parameters processed by conv4, and (Nconv2+Nconv4)*H1*W1, concatenates them, and outputs Nconv3*H1*W1 data. The second CNN filter 107b2's conv3 receives Nconv3*H1*W1 data and outputs N1*H1*W1 data.

[0171] 21 is a schematic diagram showing a modified example of the configuration of the image filter device according to the present embodiment. As shown in FIG. 21, the second CNN filter 107c2 of the CNN filter 107c includes an add layer. The add layer may include the N1*H1*W1 conv3 output of the second CNN filter 107b2. It inputs N1*H1*W1 image data and N1*H1*W1 data, and outputs N1*H1*W1 data.

[0172] In this embodiment, the CNN filter in the image decoding device is the same as that in the image encoding device. The CNN filters 107b and 107c in the above embodiment have the same functions as those in the above embodiment.

[0173] According to the configuration of the third embodiment, image data and encoded data are inputted in separate networks. With such a configuration, the input size of image data and the input size of encoded data can be made different. Furthermore, by using a network CNN1 specialized only for image data, not only is learning easy, but the overall network configuration can be made small. Furthermore, the network CNN2, which inputs encoding parameters, The filter generates the filtered image and feature extraction data derived by the first CNN filter CNN1. The output can be weighted and further feature extracted using coding parameters, allowing for more sophisticated filtering.

[0174] (Fourth embodiment) Another embodiment of the present invention will be described below with reference to Fig. 22. For the sake of convenience, the description of members having the same functions as those described in the above embodiment will be omitted.

[0175] In this embodiment, as shown in FIG. 22, the CNN filter 107d includes a first-stage CNN filter 107d1, which is a plurality of individual neural networks including n+1 CNN filters CNN0, CNN1, . . . , CNNn, a selector 107d2, and a common neural network. and a second-stage CNN filter 107d3.

[0176] In the first stage CNN filter 107d1, the CNN filter CNN0 is the above-mentioned CNN filter 107b1, but with a filter parameter FP having a value smaller than FP1. The CNN filter CNN1 is a filter optimized for FP1 and above, and is smaller than FP2. The filter is optimized for the filter parameter FP with small values. Filter CNNn is a filter optimized for filter parameters FP with values ​​equal to or greater than FPn. It's Ruta.

[0177] Each of the CNN filters CNN0, CNN1, ..., CNNn included in the first-stage CNN filter 107d1 outputs a filtered image to the selector 107d2. The selector 107d2 receives a filter parameter FP and selects a filtered image to be output to the second-stage CNN filter 107d3 according to the input filter parameter FP. The second stage CNN filter 107d3 receives the filter parameters input to the selector 107d2. An image that has been filtered by an optimal filter is input to the meter FP. In other words, the individual neural network in this embodiment selectively acts on the input image data depending on the values ​​of the filter parameters in the input image data to the image filter device.

[0178] Note that the filter parameter FP for selecting the CNN filter is explicitly set in the encoded data. For example, the filter parameter FP may be derived from a representative value (such as an average value) of a quantization parameter, which is one of the encoding parameters.

[0179] The second-stage CNN filter 107d3 filters the input image and outputs the filtered image In other words, the common neural network in this embodiment acts commonly on the output image data of the individual neural networks, regardless of the values ​​of the filter parameters.

[0180] The filter parameter FP used for selection by the selector 107d2 is not limited to a representative value of the quantization parameter (QP) in the input image. The filter parameter FP may be explicitly transmitted in the encoded data. In addition to the quantization parameter in the input image, the filter parameter FP may include a parameter indicating the type of intra prediction and inter prediction in the input image, a parameter indicating the intra prediction direction in the input image (intra prediction mode), a parameter indicating the partition division depth (depth information, division information) of the input image, and a parameter indicating the size of the partition in the input image. In addition, representative values ​​such as a value at a specific position (upper left or center), an average value, a minimum value, a maximum value, a median value, and a mode value may also be used for these parameters.

[0181] 23 is a schematic diagram showing a modified example of the configuration of the image filter device according to the present embodiment. As shown in FIG. 23, the CNN filter 107e includes a CNN filter 107e2, which is a plurality of individual neural networks including n+1 CNN filters CNN0, CNN1, ..., CNNn, and a CNN filter 107e3, which is a CNN filter 107e4, which is a CNN filter 107e5, which is a CNN filter 107e6, which is a CNN filter 107e7, which is a CNN filter 107e8, which is a CNN filter 107e9, which is a CNN filter 107e10, which is a CNN filter 107e11, which is a CNN filter 107e12, which is a CNN filter 107e13, which is a CNN filter 107e14, which is a CNN filter 107e The rectifier 107e3 is the latter stage of the CNN filter 107e1, which is a common neural network. In this case, the CNN filter 107e1 acts on the input image data to the CNN filter 107e, and the CNN filter 107e2 acts on the filter in the input image data. Depending on the value of the data parameter, the CNN filter 107e1 selectively applies the following to the output image data: In addition, the selector 107e3 outputs the filtered image.

[0182] In this embodiment, the CNN filter in the image decoding device is the same as that in the image encoding device. It has the same function as the CNN filter in

[0183] According to the configuration of the fourth embodiment, by using a part (107e2) that switches the network depending on the magnitude of the filter parameter FP and a part (107e1) that uses the same network regardless of the magnitude of the filter parameter, the network configuration can be made smaller than a configuration in which the entire filter is switched by an encoding parameter such as a quantization parameter. In addition to the fact that the smaller the network configuration, the smaller the amount of calculation is and the faster the process is, the more robust the learning parameters are, and the more suitable the filter processing can be performed for many input images.

[0184] In this embodiment, as in (a) of FIG. 34 in the first embodiment, the pre-filter image has a luminance (Y) channel, a first chrominance (Cb) channel, and a second chrominance (Cr) channel. The CNN filter 107d1 may be configured to input the CNN filter 107d1. As shown in FIG. 34(b) in the first embodiment, the pre-filter image has a luminance (Y) channel and The first color difference (Cb) channel and the second color difference (Cr) channel are separated and fed to the CNN filter 107d1. That is, the CNN filter d1 may be configured to receive the luminance and the reference parameter Alternatively, the CNN filter d1 may perform filtering on an image (input) in which the first chrominance, the second chrominance, and the reference parameters are interleaved. Note that the reference parameters (coding parameters) are not limited to the quantization parameters (QP), and the CNN filter d1 may use one or more coding parameters.

[0185] Fifth embodiment Another embodiment of the present invention will be described below with reference to Fig. 24. For the sake of convenience, the description of members having the same functions as those described in the above embodiment will be omitted.

[0186] In this embodiment, as shown in FIG. 24, the CNN filter 107f includes a first-stage CNN filter 107f1 including n+1 CNN filters CNN0, CNN1, . . . , CNNn, and a selector 10 7f2 and a second stage CNN filter 107f3.

[0187] In the first stage CNN filter 107f1, the CNN filter CNN1 is a filter optimized for a quantization parameter (QP) having a value larger than QP1L and smaller than QP1H. The CNN filter CNN2 is a filter optimized for a quantization parameter (QP) having a value larger than QP2L and smaller than QP2H. The CNN filter CNN3 is a filter optimized for the QP3L. The CNN filter CNN4 is a filter optimized for a quantization parameter (QP) that is greater than QP4L and smaller than QP4H. This filter is optimized for the quantization parameter (QP). The same goes for other CNN filters. is a filter.

[0188] Specific examples of thresholds QP1L, QP1H, ... QP4L, QP4H are QP1L=0, QP1H=18, QP2L=12, QP2H=30. , QP3L=24, QP3H=42, QP4L=36, QP4H=51 can be assigned values.

[0189] In this case, for example, if the quantization parameter (QP) is 10, the selector 107f2 selects the CNN If the quantization parameter (QP) is 15, the selector 107f2 selects the CNN filter CNN1 and the CNN filter CNN2. If the quantization parameter (QP) is 20, the selector 107f2 selects the CNN filter CNN2. If the quantization parameter (QP) is 25, the selector 107f2 selects the CNN filter CNN2 and the CNN filter CNN3. If the quantization parameter (QP) is 30, the selector 107f2 selects the CNN filter CNN3.

[0190] The second-stage CNN filter 107f3 outputs the input image as a filtered image when the selector 107f2 selects one type of CNN filter, and outputs the average value of the two input images as a filtered image when the selector 107f2 selects two types of CNN filters. .

[0191] In this embodiment, the CNN filter in the image decoding device is the same as that in the image encoding device. The CNN filter 107f has a similar function to that of the CNN filter 107f in the

[0192] According to the configuration of the fifth embodiment, by using a part (107f1) that switches the network depending on the size of the quantization parameter QP and a part (107f2) that uses the same network regardless of the size of the quantization parameter QP, the network configuration can be made smaller than a configuration in which the entire filter is switched by an encoding parameter such as a quantization parameter. A smaller network configuration has the effect of reducing the amount of calculation and speeding up the process, as well as making the learning parameters robust and enabling appropriate filter processing for many input images. Also, each CNN filter By overlapping the optimization ranges of the filters, we can avoid visual distortion at patch boundaries when switching filters.

[0193] Sixth embodiment Another embodiment of the present invention will be described below with reference to Fig. 25. For the sake of convenience, the description of members having the same functions as those described in the above embodiment will be omitted.

[0194] As described above, the image filter device may use a function for reducing block distortion and a filter for reducing ringing distortion. Also, a deblocking filter (DF) for reducing block distortion and a sample adaptation filter for reducing ringing distortion may be used. Even if you use other filters and processing such as Sample Adaptive Offset (SAO), good.

[0195] In this embodiment, a configuration will be described in which a deblocking filter (DF) process, a sample adaptive offset (SAO) process, and a CNN filter are used in combination.

[0196] (First example) A first example of this embodiment is shown in (a) of Fig. 25. In the first example, the image filter device 107g includes a CNN filter 107g1 and a sample adaptive offset (SAO) 107g2. The CNN filter 107g1 functions as a filter that reduces block noise.

[0197] (Second example) FIG. 25B shows a second example of this embodiment. In the second example, the image filter device 107h includes a deblocking filter (DF) 107h1 and a CNN filter 107h2. The CNN filter 107h2 further reduces ringing noise after the deblocking filter. It acts as a filter to reduce noise.

[0198] (Third example) A third example of this embodiment is shown in (c) of Fig. 25. In the third example, the image filter device 107i includes a first CNN filter 107i1 and a second CNN filter 107i2. The first CNN filter 107i1 functions as a filter that reduces block noise, and the second CNN filter 107i2 functions as a filter that further reduces ringing noise after the filter that reduces block noise.

[0199] In any of the examples, the CNN filter in the image decoding device is the same as the CNN filter in the image coding device. It has the same function as the CNN filter.

[0200] In addition, the unfiltered images input to the image filter devices 107g to 107i in this embodiment are, as in the other embodiments, a luminance (Y) channel, a first color difference (1 / 2) channel, and a second color difference (1 / 2) channel. Alternatively, the pre-filtered image may include three image channels including a luminance (Y) and quantization parameter (QP) channel, and a first chrominance (Cb) and quantization parameter (QP) channel. Also, as shown in (a) of FIG. 34, the pre-filtered image may include three image channels including a luminance (Y) and quantization parameter (QP) channel, and a first chrominance (Cb) and quantization parameter (QP) channel. Alternatively, the pre-filter image may be divided into a channel for luminance (Y) and a quantization parameter (QP) and a channel for a second chrominance (Cr) and a quantization parameter (QP) and input to the image filter devices 107g to 107i. Also, as shown in FIG. 34(b), the pre-filter image may be divided into a channel for luminance (Y) and a quantization parameter (QP) and a channel for a second chrominance (Cr) and a quantization parameter (QP). The input signal may be divided into a channel for a luminance (QP) and channels for a first chrominance (Cb), a second chrominance (Cr), and a quantization parameter (QP), and input to the image filter devices 107g to 107i. That is, the image filter devices 107g to 107i may perform a filter process on an image (input) in which the luminance and the reference parameters are interleaved, and may perform a filter process on an image (input) in which the first chrominance, the second chrominance, and the reference parameters are interleaved. Note that the reference parameter (encoding parameter) is not limited to the quantization parameter (QP), and the image filter devices 107g to 107i may use one or more encoding parameters.

[0201] Seventh embodiment Another embodiment of the present invention will be described below with reference to Figures 26 to 30. For the sake of convenience, the description of members having the same functions as those described in the above embodiment will be omitted.

[0202] FIG. 26 is a block diagram showing the configuration of an image encoding device according to this embodiment. The image encoding device 11j according to this embodiment differs from the above embodiment in that a CNN filter 107j acquires CNN parameters and performs filtering using the acquired CNN parameters. In addition, the CNN parameters used by the CNN filter 107j are different from those in the above embodiment in that they are dynamically updated on a sequence-by-sequence, picture-by-picture basis, etc. The meter is a fixed, pre-determined value and is not updated.

[0203] As shown in FIG. 26, the image encoding device 11j according to this embodiment includes a CNN parameter determination unit 114, a CNN parameter encoding unit 115, and a multiplexing unit 116 in addition to the components included in the image encoding device 11 shown in FIG.

[0204] The CNN parameter determination unit 114 receives the image T (input image) and the output of the adder 106 (f The neural network parameters (CNN parameters) are updated so that the difference between the input image and the unfiltered image is reduced.

[0205] FIG. 27 is a schematic diagram showing an example of the configuration of a CNN filter 107j. As described above, a CNN filter includes multiple layers such as a convX layer, and the CNN filter 107j shown in FIG. 27 includes three layers. Each layer can be identified by a layer ID. In the CNN filter 107j shown in FIG. 27, The layer ID of the input layer is L-2, the layer ID of the middle layer is L-1, and the layer ID of the output layer is L.

[0206] Each layer contains multiple units, and each unit can be identified by a unit ID. The unit ID of the top unit of the middle layer L-1 is (L-1, 0), the unit ID of the unit above the output layer L is (L, 0), and the unit ID of the unit below the output layer L is (L, 1). As shown in Figure 27, each unit in each layer is connected to a unit in the next layer. In Figure 27, the connections between units are indicated by arrows. The weight of each connection is different and is controlled by a weight coefficient.

[0207] The CNN parameter determination unit 114 determines the filter including both the weighting coefficients and the bias (offset). The CNN parameter determination unit 114 outputs a filter coefficient. In addition, the CNN parameter determination unit 114 outputs an identifier as a CNN parameter. When the CNN filter is composed of multiple CNN layers, the identifier identifies the CNN layer. Also, if the CNN layer is identified by the layer ID and unit ID, the identifier is Layer ID, unit ID.

[0208] In addition, the CNN parameter determination unit 114 outputs data indicating the unit structure as the CNN parameters. The data indicating the unit structure may be, for example, a filter size such as 3*3. The data indicating the filter size is output as a CNN parameter when the filter size is variable. When the filter size is fixed, the filter size There is no need to output data indicating this.

[0209] The CNN parameter determination unit 114 performs a full update to update all parameters or a partial update to update only a part of the parameters. Partial update is performed to update the parameters of the units of the layer. CNN parameter determination unit 114 The data indicating whether to output the update contents as differences is added to the CNN parameters. To exert effort.

[0210] The CNN parameter determination unit 114 can output CNN parameter values ​​such as filter coefficients as they are. The CNN parameter determination unit 114 can also output differential parameter values ​​such as differences from CNN parameter values ​​before updating, differences from default values, etc. The CNN parameter determination unit 114 can also compress the CNN parameter values ​​using a predetermined method and output them.

[0211] Referring to Figure 28, the layer and unit configuration and the CNN parameters (filter coefficients, weight coefficients ) update will be explained.

[0212] At each layer, the input value Z (L-1) ijk and the L layer parameters (filter coefficients) h pqr , h 0The sum of products of and is passed to the activation function (equation (1) shown in Figure 28), and the value Z L ijk to the next layer, where N is the number of channels in the input of the layer, and W is the width of the input of the layer. where H is the input height of the layer, and kN is the number of input channels of the kernel (filter). is essentially equal to N, and kW is the kernel width and kH is the kernel height. .

[0213] In this embodiment, the CNN parameter determination unit 114 determines the CNN parameters (filter coefficients) h pqr , h 0 At least a portion of the can be dynamically updated.

[0214] In this embodiment, the CNN parameters are data of the Network Abstraction Layer (NAL) structure. FIG. 29(a) shows a sequence of data having the NAL structure according to the present embodiment. In this embodiment, the sequence parameter set SPS (Sequence Parameter Set) included in the sequence SEQ is updated according to the update type ( (indicating whether it is partial / full / differential, etc.), CNN layer ID (L), CNN unit ID (m), L layer, Filter size (kW*kH) and filter coefficient (h pqr , h 0 ) and other image sequences It also transmits update parameters that are applied to the entire sequence SEQ. The picture parameter set PPS (Picture Parameter Set) includes the update type (indicating whether it is partial / full / differential, etc.), layer ID (L), unit ID (m), filter size (kW*kH), filter coefficient (Kp), and the like. Number (h pqr , h 0 ) to be applied to a certain picture.

[0215] Also, as shown in (b) of FIG. 29, a sequence includes multiple pictures. The data determination unit 114 can output the CNN parameters on a sequence-by-sequence basis. In this case, the CNN parameters for the entire sequence can be updated. Also, the CNN parameter determination unit 114 can output the CNN parameters for each picture. In this case, The data can be updated.

[0216] The matters described with reference to FIGS. 27 to 29 are common to the encoding side and the decoding side, and the same applies to the CNN filter 305j described later. The above-mentioned matters are applied to the CNN parameters output to the CNN filter 107j of the image encoding device 11j, and also to the CNN parameters output to the CNN filter 305j of the image decoding device 31j.

[0217] In addition, the unfiltered image input to the CNN filter 107j in this embodiment is As in the embodiment, a luminance (Y) channel, a first chrominance (Cb) channel, and a second The pre-filter image may be an image including three image channels including a chrominance (Cr) channel and one coding parameter channel including a quantization parameter (QP) channel. Also, as shown in FIG. 34(a), the pre-filter image may be an image including three image channels including a chrominance (Cr) channel and one coding parameter channel including a quantization parameter (QP) channel. The signal is divided into a first channel for color difference (Cb) and a quantization parameter (QP), and a second channel for color difference (Cr) and a quantization parameter (QP), and the two channels are input to the CNN filter 107j. Also, as shown in FIG. 34(b), the unfiltered image may be a luminance (Y) and quantization parameter. Alternatively, the signal may be divided into a channel for the first chrominance (Cb), a channel for the second chrominance (Cr), and a channel for the quantization parameter (QP), and the channels for the first chrominance (Cb), the second chrominance (Cr), and the quantization parameter (QP) are input to the CNN filter 107j. That is, the CNN filter 107j uses an image with interleaved brightness and reference parameters (input Alternatively, the CNN filter 107j may perform a filter process on the first chrominance input (input) and perform a filter process on an image (input) in which the first chrominance input, the second chrominance input, and the reference parameter are interleaved. Note that the reference parameter (coding parameter) is not limited to the quantization parameter (QP), and the CNN filter 107j may perform a filter process on a single CNN filter 107j. The above coding parameters can be used.

[0218] The CNN parameter encoding unit 115 encodes the CNN parameters output by the CNN parameter determination unit 114. The CNN parameter is then encoded and output to the multiplexing unit 116.

[0219] The multiplexing unit 116 multiplexes the encoded data output by the entropy encoding unit 104 and the CNN parameters. The meter encoding unit 115 multiplexes the encoded CNN parameters to generate a stream, Output the stream to the outside.

[0220] 30 is a block diagram showing the configuration of an image decoding device according to this embodiment. In the image decoding device 31j according to this embodiment, a CNN filter 305j acquires CNN parameters and performs filtering using the acquired CNN parameters. The parameters are dynamically updated on a sequence-by-sequence, picture-by-picture basis, and so on.

[0221] As shown in FIG. 30, the image decoding device 31j according to this embodiment includes a demultiplexing unit 313 and a CNN parameter decoding unit 314 in addition to the components included in the image decoding device 31 shown in FIG. It is possible.

[0222] The demultiplexer 313 receives the stream and demultiplexes the encoded data and the encoded CNN parameters. The data is demultiplexed into

[0223] The CNN parameter decoding unit 314 decodes the encoded CNN parameters and outputs the CNN filter 3 Output to 05j.

[0224] In addition, in the above-described embodiment, the image encoding device 11 and a part of the image decoding device 31, for example, the entropy decoding unit 301, the prediction parameter decoding unit 302, the CNN filter 305, the prediction a predicted image generating unit 308, an inverse quantization and inverse transformation unit 311, an addition unit 312, a predicted image generating unit 101, a subtraction unit 102, a transformation and quantization unit 103, an entropy coding unit 104, an inverse quantization and inverse transformation unit 105, a CNN filter 107, a coding parameter determination unit 110, a predicted parameter code Alternatively, the encoding unit 111 may be realized by a computer. In that case, a program for realizing this control function may be recorded on a computer-readable recording medium, and the program recorded on the recording medium may be read and executed by a computer system to realize the control function. Note that the "computer system" referred to here is a computer system built into either the image encoding device 11 or the image decoding device 31, and includes hardware such as an OS and peripheral devices. Also, the "computer-readable recording medium" refers to a portable medium such as a flexible disk, a magneto-optical disk, a ROM, a CD-ROM, or a computer It refers to a storage device such as a hard disk built into the system. Furthermore, "computer-readable recording medium" may also include a device that dynamically stores a program for a short period of time, such as a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, or a device that stores a program for a certain period of time, such as a volatile memory inside a computer system that serves as a server or client in such a case. Furthermore, the above program may be one that realizes part of the above-mentioned functions, or may be one that can realize the above-mentioned functions in combination with a program already recorded in the computer system.

[0225] In addition, a part or the whole of the image encoding device 11 and the image decoding device 31 in the above-mentioned embodiment may be realized as an integrated circuit such as an LSI (Large Scale Integration). Each functional block of the image encoding device 11 and the image decoding device 31 may be individually made into a processor, or a part or the whole may be integrated into a processor. The integrated circuit method is not limited to LSI, and may be realized by a dedicated circuit or a general-purpose processor. In addition, when an integrated circuit technology that replaces LSI appears due to the progress of semiconductor technology, an integrated circuit based on that technology may be used.

[0226] Although one embodiment of the present invention has been described in detail above with reference to the drawings, the specific configuration is not limited to the above, and various design changes, etc. are possible within the scope that does not deviate from the gist of the present invention.

[0227] [Application example] The above-mentioned image encoding device 11 and image decoding device 31 can be mounted on various devices that transmit, receive, record, and play back moving images. The moving images may be natural moving images captured by a camera or the like, or artificial moving images (including CG and GUI) generated by a computer or the like.

[0228] First, the image encoding device 11 and the image decoding device 31 are used for transmitting and receiving moving images. The following describes the use of this function with reference to FIG.

[0229] Fig. 31(a) is a block diagram showing the configuration of a transmission device PROD_A equipped with an image coding device 11. As shown in Fig. 31(a), the transmission device PROD_A includes an encoding unit PROD_A1 that obtains encoded data by encoding a moving image, a modulation unit PROD_A2 that obtains a modulated signal by modulating a carrier wave with the encoded data obtained by the encoding unit PROD_A1, and a transmission unit PROD_A3 that transmits the modulated signal obtained by the modulation unit PROD_A2. The image coding device 11 described above includes: This encoding unit is used as PROD_A1.

[0230] The transmission device PROD_A captures moving images as a supply source of moving images to be input to the encoding unit PROD_A1. The apparatus further includes a camera PROD_A4 for recording moving images, a recording medium PROD_A5 for recording moving images, an input terminal PROD_A6 for inputting moving images from the outside, and an image processing unit A7 for generating or processing images. Although Fig. 31(a) illustrates an example in which the transmitting device PROD_A includes all of these components, some of them may be omitted.

[0231] The recording medium PROD_A5 may be one in which unencoded moving images are recorded. Alternatively, the recording medium PROD_A5 may be a recording medium that has recorded thereon a moving image encoded by a recording encoding method different from the transmission encoding method. In the latter case, a decoding unit ( It is advisable to use a conductor (not shown) between the two.

[0232] Fig. 31(b) is a block diagram showing the configuration of a receiving device PROD_B equipped with an image decoding device 31. As shown in Fig. 31(b), the receiving device PROD_B includes a receiving unit PROD_B1 that receives a modulated signal, a demodulation unit PROD_B2 that obtains coded data by demodulating the modulated signal received by the receiving unit PROD_B1, and a decoding unit PROD_B3 that obtains a moving image by decoding the coded data obtained by the demodulation unit PROD_B2. The above-mentioned image decoding device 31 is used as this decoding unit PROD_B3.

[0233] The receiving device PROD_B is a supply destination of the video output by the decoding unit PROD_B3, and displays the video. In (b) of FIG. 31, the display PROD_B4 may further include a recording medium PROD_B5 for recording the moving image, and an output terminal PROD_B6 for outputting the moving image to the outside. Although the receiving device PROD_B is illustrated as having all of these components, some of them may be omitted.

[0234] The recording medium PROD_B5 is for recording unencoded moving images. Alternatively, the signal may be encoded by a coding method for recording that is different from the coding method for transmission. In the latter case, a signal from the decoding unit PROD_B3 to the recording medium PROD_B5 may be provided between the decoding unit PROD_B3 and the recording medium PROD_B5. It is preferable to interpose an encoding unit (not shown) that encodes the acquired video image according to an encoding method for recording.

[0235] The transmission medium for transmitting the modulated signal may be wireless or wired. The transmission mode for transmitting the modulated signal may be broadcast (here, this refers to a transmission mode in which the destination is not specified in advance) or communication (here, this refers to a transmission mode in which the destination is specified in advance). In other words, the transmission of the modulated signal may be realized by any of wireless broadcast, wired broadcast, wireless communication, and wired communication.

[0236] For example, a broadcasting station (such as a broadcasting facility) / receiving station (such as a television receiver) for terrestrial digital broadcasting is an example of a transmitting device PROD_A / receiving device PROD_B that transmits and receives modulated signals by wireless broadcasting. Also, a broadcasting station (such as a broadcasting facility) / receiving station (such as a television receiver) for cable television broadcasting is an example of a transmitting device PROD_A / receiving device PROD_B that transmits and receives modulated signals by cable broadcasting. .

[0237] Also, a server (such as a workstation) / client (such as a television receiver, a personal computer, a smartphone, etc.) of a VOD (Video On Demand) service or a video sharing service using the Internet is an example of a transmitting device PROD_A / receiving device PROD_B that transmits and receives modulated signals by communication (usually, in a LAN, either wireless or wired is used as a transmission medium, and in a WAN, wired is used as a transmission medium). Here, personal computers include desktop PCs, laptop PCs, and tablet PCs. Also, smartphones include multi-function mobile phone terminals.

[0238] The client of the video hosting service has a function to decode the encoded data downloaded from the server and display it on a display, as well as a function to encode the video captured by the camera and upload it to the server. That is, the client of the video hosting service functions as both the transmitting device PROD_A and the receiving device PROD_B.

[0239] Next, it will be described with reference to FIG. 32 that the above-mentioned image encoding device 11 and image decoding device 31 can be used for recording and reproducing moving images.

[0240] Fig. 32(a) is a block diagram showing the configuration of a recording device PROD_C equipped with the above-mentioned image coding device 11. As shown in Fig. 32(a), the recording device PROD_C includes an encoding unit PROD_C1 that obtains encoded data by encoding a moving image, and a writing unit PROD_C2 that writes the encoded data obtained by the encoding unit PROD_C1 onto a recording medium PROD_M. The encoding device 11 is used as this encoding unit PROD_C1.

[0241] The recording medium PROD_M may be (1) a type built into the recording device PROD_C, such as an HDD (Hard Disk Drive) or SSD (Solid State Drive), (2) a type connected to the recording device PROD_C, such as an SD memory card or a USB (Universal Serial Bus) flash memory, or (3) a drive device built into the recording device PROD_C, such as a DVD (Digital Versatile Disc) or a BD (Blu-ray Disc: registered trademark). (not shown).

[0242] The recording device PROD_C supplies a video image to the encoding unit PROD_C1. The recording device PROD_C may further include a camera PROD_C3 for capturing an image, an input terminal PROD_C4 for inputting a moving image from the outside, a receiving unit PROD_C5 for receiving the moving image, and an image processing unit PROD_C6 for generating or processing an image. In (a) of Fig. 32, a configuration in which the recording device PROD_C includes all of these components is illustrated, but some of them may be omitted.

[0243] The receiving unit PROD_C5 may receive uncoded video. Alternatively, the receiving unit PROD_C5 may receive encoded data that has been encoded by a transmission encoding method different from the recording encoding method. In the latter case, a transmission decoding unit (not shown) that decodes the encoded data that has been encoded by the transmission encoding method may be provided between the receiving unit PROD_C5 and the encoding unit PROD_C1.

[0244] Examples of such a recording device PROD_C include a DVD recorder, a BD recorder, and an HDD (Hard Disk Drive) recorder (in this case, the input terminal PROD_C4 or the receiving unit PROD_C5 is the main source of the moving image). Also, a camcorder (in this case, the camera PROD_C3 is the main source of the moving image), a personal computer (in this case, the receiving unit PROD_C5 or the image The processing unit C6 is the main source of video images), a smartphone (in this case, the camera PROD_C3 Or the receiving unit PROD_C5 is the main source of video images) This is an example.

[0245] FIG. 32(b) is a block diagram showing the configuration of a playback device PROD_D equipped with the above-mentioned image decoding device 31. As shown in FIG. 32(b), the playback device PROD_D includes a reading unit PROD_D1 that reads out coded data written to a recording medium PROD_M, and a decoding unit PROD_D2 that obtains a moving image by decoding the coded data read by the reading unit PROD_D1. The image decoding device 31 is used as this decoding unit PROD_D2.

[0246] The recording medium PROD_M may be (1) a type built into the playback device PROD_D, such as an HDD or SSD, or (2) a type of storage device, such as an SD memory card or a USB flash memory. (3) DVD, BD, etc. The DVD-ROM 10 may be loaded into a drive device (not shown) built into the playback device PROD_D.

[0247] In addition, the playback device PROD_D receives the video output from the decoding unit PROD_D2 and outputs the video to the playback device PROD_D. The device may further include a display PROD_D3 for displaying the moving image, an output terminal PROD_D4 for outputting the moving image to the outside, and a transmission unit PROD_D5 for transmitting the moving image. Although the playback device PROD_D is illustrated as having all of these components, some of them may be omitted.

[0248] The transmission unit PROD_D5 may transmit uncoded video. Alternatively, the decoder unit PROD_D2 may transmit encoded data encoded by a transmission encoding method different from the recording encoding method. In the latter case, it is preferable to interpose an encoding unit (not shown) between the decoder unit PROD_D2 and the transmitter unit PROD_D5, which encodes the video image by the transmission encoding method.

[0249] Examples of such a playback device PROD_D include a DVD player, a BD player, and a HDD player (in this case, the output terminal PROD_D4 to which a television receiver or the like is connected operates). In addition, television receivers (in this case, the display PROD_D3 is the main destination of moving images), digital signage (also called electronic billboards or electronic bulletin boards, etc.) a display PROD_D3 or a transmission unit PROD_D5 being the main destination of the moving image), a desktop PC (in this case, the output terminal PROD_D4 or the transmission unit PROD_D5 being the main destination of the moving image), a laptop or tablet PC (in this case, the display PROD_D3 or the transmission unit PROD_D5 being the main destination of the moving image), Examples of such a playback device PROD_D include a playback device PROD_D for a smartphone (in which case the display PROD_D3 or the transmission unit PROD_D5 is the main destination of the video images), and a smartphone (in which case the display PROD_D3 or the transmission unit PROD_D5 is the main destination of the video images).

[0250] (Hardware and Software Realizations) Furthermore, each block of the image decoding device 31 and the image encoding device 11 described above may be realized in hardware by a logic circuit formed on an integrated circuit (IC chip), or may be realized by a CPU This may be realized in software using a Central Processing Unit (Central Processing Unit).

[0251] In the latter case, each of the above devices includes a CPU that executes the instructions of a program that realizes each function, Storage devices such as ROM (Read Only Memory) that stores programs, RAM (Random Access Memory) that expands the programs, and memory that stores the programs and various data. The object of the embodiment of the present invention can also be achieved by supplying each of the above devices with a recording medium on which program code (executable program, intermediate code program, source program) of a control program for each of the above devices, which is software for realizing the above functions, is recorded in a computer-readable manner, and the computer (or CPU or MPU) reads out and executes the program code recorded on the recording medium.

[0252] Examples of the recording medium include tapes such as magnetic tapes and cassette tapes, magnetic disks such as floppy disks (registered trademark) and hard disks, disks including optical disks such as CD-ROMs (Compact Disc Read-Only Memory), MO disks (Magneto-Optical discs), MDs (Mini Discs), DVDs (Digital Versatile Discs), CD-Rs (CD Recordable), and Blu-ray Discs (registered trademark), cards such as IC cards (including memory cards) and optical cards, mask ROMs, EPROMs (Erasable Programmable Read-Only Memory), EEPROMs (Electrically Erasable and Programmable Read-Only Memory), flash memory, and the like. For example, semiconductor memories such as a ROM or logic circuits such as a programmable logic device (PLD) or a field programmable gate array (FPGA) can be used.

[0253] Moreover, each of the above devices may be configured to be connectable to a communication network, and the above program code may be supplied via the communication network. This communication network is not particularly limited as long as it is capable of transmitting the program code. For example, the Internet, an intranet, an extranet, a LAN (Local Area Network), an ISDN (Integrated Services Digital Network), a VAN (Value-Added Network), a CATV (Community Antenna television / Cable Television) communication network, a virtual private network, a telephone line network, a mobile communication network, a satellite communication network, etc. may be used. Moreover, the transmission medium constituting this communication network is not limited to a specific configuration or type as long as it is a medium capable of transmitting the program code. For example, it can be used in wired communication such as IEEE (Institute of Electrical and Electronic Engineers) 1394, USB, power line carrier, cable TV line, telephone line, ADSL (Asymmetric Digital Subscriber Line) line, etc., or in wireless communication such as infrared such as IrDA (Infrared Data Association) or remote control, BlueTooth (registered trademark), IEEE802.11 wireless, HDR (High Data Rate), NFC (Near Field Communication), DLNA (Digital Living Network Alliance: registered trademark), mobile phone network, satellite line, terrestrial digital broadcasting network, etc. The embodiment of the present invention can also be realized in the form of a computer data signal embedded in a carrier wave in which the program code is embodied by electronic transmission.

[0254] The present invention is not limited to the above-described embodiment, and various modifications are possible within the scope of the claims. In other words, the technical scope of the present invention also includes embodiments obtained by combining technical means that are appropriately modified within the scope of the claims.

[0255] 〔summary〕 The image filter device (CNN filter 107, 305) according to the first aspect of the present invention is a luminance or The device is equipped with a neural network that receives input of one or more first type input image data having pixel values ​​represented by chrominance, and one or more second type input image data having pixel values ​​represented by values ​​corresponding to reference parameters for generating a predicted image and a differential image, and outputs one or more first type output image data having pixel values ​​represented by luminance or chrominance.

[0256] According to the above arrangement, a filter according to image characteristics can be applied to input image data.

[0257] The image filter device (CNN filter 107, 305) according to the second aspect of the present invention is In the above embodiment, the present invention may further include a parameter determination unit (CNN parameter determination unit 114) that updates neural network parameters used by the neural network.

[0258] According to the above configuration, the parameters used by the neural network can be updated.

[0259] The image filter device (CNN filter 107, 305) according to the third aspect of the present invention is In 1 or 2, the reference parameters may include a quantization parameter in the image on which the image filter device operates.

[0260] The image filter device (CNN filter 107, 305) according to the fourth aspect of the present invention is In 1 to 3, the reference parameters may include a parameter indicating the type of intra prediction or inter prediction for the image on which the image filter device operates.

[0261] The image filter device (CNN filter 107, 305) according to the fifth aspect of the present invention is In 1 to 4, the reference parameters may include a parameter (intra prediction mode) indicating an intra prediction direction in an image on which the image filter device acts.

[0262] The image filter device (CNN filter 107, 305) according to the sixth aspect of the present invention is In 1 to 4, the reference parameters may include a parameter indicating a division depth of a partition in an image on which the image filter device operates.

[0263] The image filter device (CNN filter 107, 305) according to the seventh aspect of the present invention is In 1 to 6, the reference parameters may include a parameter indicating the size of a partition in the image on which the image filter device operates.

[0264] The image filter device (CNN filter 107, 305) according to the eighth aspect of the present invention is In any one of 1 to 7, a second neural network may be provided that outputs an image based on the first type of input image data input to the neural network.

[0265] In the image filter device (CNN filter 107, 305) according to the ninth aspect of the present invention, The neural network in the above-described embodiments 1 to 8 may receive as input the first type of input image data having a first color difference (Cb) and a second color difference (Cr) as pixel values ​​and the second type of input image data, and output a first type of output image data having a first color difference (Cb) and a second color difference (Cr) as pixel values.

[0266] In the image filter device (CNN filter 107, 305) according to the tenth aspect of the present invention, The neural network in the above aspects 1 to 8 may include a means for receiving the first type of input image data having luminance as pixel values ​​and the second type of input image data, and outputting first type of output image data having luminance as pixel values, and a means for receiving the first type of input image data having the first color difference and the second color difference as pixel values ​​and outputting first type of output image data having the first color difference and the second color difference as pixel values.

[0267] The image filter device (CNN filters 107d and 107f) according to the eleventh aspect of the present invention The image filter device (107) comprises a number of individual neural networks (107d1, 107f1) and a common neural network (107d3, 107f3), the individual neural networks (107d1, 107f1) acting selectively on input image data according to values ​​of filter parameters in the input image data to the image filter device (107), and the common neural network (107d3, 107f3) acting commonly on output image data of the individual neural networks regardless of the values ​​of the filter parameters.

[0268] According to the above arrangement, both filtering processes based on the values ​​of the filter parameters and filtering processes independent of the values ​​of the filter parameters can be applied to the image data.

[0269] The image filter device (CNN filter 107e) according to the twelfth aspect of the present invention is a CNN filter that The image filter device (107) includes a common neural network (107e1) that operates on input image data to the image filter device (107) and a separate neural network (107e2) that operates on the input image data to the image filter device (107). , selectively operates on the output image data of the common neural network depending on values ​​of filter parameters in the input image data.

[0270] According to the above configuration, the same effects as those of the third aspect are achieved.

[0271] An image filter device according to the thirteenth aspect of the present invention (CNN filters 107d, 107e, 107 f) In the above aspect 11 or 12, the filter parameter may be a quantization parameter in the image on which the image filter device operates.

[0272] According to the above configuration, it is possible to use filter parameters according to the image.

[0273] An image filter device according to the fourteenth aspect of the present invention (CNN filters 107d, 107e, 107 In f) of the above aspect 11 or 12, the filter parameter may be a parameter indicating the type of intra prediction or inter prediction for the image on which the image filter device acts.

[0274] An image filter device according to a fifteenth aspect of the present invention (CNN filters 107d, 107e, 107 In f) of the above aspect 11 or 12, the filter parameter may be a parameter indicating an intra prediction direction in an image on which the image filter device acts (intra prediction mode).

[0275] An image filter device according to the sixteenth aspect of the present invention (CNN filters 107d, 107e, 107 In aspect f), in the eleventh or twelfth aspect, the filter parameter may be a parameter indicating a division depth of a partition in an image on which the image filter device operates.

[0276] An image filter device according to the seventeenth aspect of the present invention (CNN filters 107d, 107e, 107 In aspect f), in the eleventh or twelfth aspect, the filter parameter may be a parameter indicating a size of a partition in an image on which the image filter device operates.

[0277] An image decoding device (31, 31j) according to aspect 18 of the present invention is an image decoding device that decodes an image, and includes the image filter device of aspects 1 to 15 as a filter that acts on the decoded image.

[0278] An image coding device (11, 11j) according to a nineteenth aspect of the present invention is an image coding device that codes an image, and includes the image filter device of aspects 1 to 15 as a filter that acts on a locally decoded image.

[0279] Moreover, aspects 13 to 17 of the present invention may have the following configurations.

[0280] An image filter device according to the thirteenth aspect of the present invention (CNN filters 107d, 107e, 107 f) In the above aspect 9 or 10, the filter parameter may be an average value of the quantization parameters in the image on which the image filter device acts.

[0281] According to the above configuration, it is possible to use filter parameters that correspond to the entire image.

[0282] An image filter device according to the fourteenth aspect of the present invention (CNN filters 107d, 107e, 107 In f), in aspect 9 or 10, the filter parameter may be an average value of parameters indicating the type of intra prediction or inter prediction in an image on which the image filter device acts.

[0283] An image filter device according to a fifteenth aspect of the present invention (CNN filters 107d, 107e, 107 In f) of aspect 9 or 10, the filter parameter may be an average value of parameters indicating an intra prediction direction (intra prediction mode) in an image on which the image filter device acts.

[0284] An image filter device according to the sixteenth aspect of the present invention (CNN filters 107d, 107e, 107 f) In the above aspect 9 or 10, the filter parameter may be an average value of a parameter indicating a division depth of partitions in an image on which the image filter device operates.

[0285] An image filter device according to the seventeenth aspect of the present invention (CNN filters 107d, 107e, 107 f) In the above aspect 9 or 10, the filter parameter may be an average value of a parameter indicating a size of a partition in an image on which the image filter device operates.

[0286] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. The technical scope of the present invention also includes embodiments obtained by appropriately combining the technical means disclosed in the different embodiments. Furthermore, new technical features can be formed by combining the technical means disclosed in the respective embodiments. [Industrial Applicability]

[0287] The embodiments of the present invention can be suitably applied to an image decoding device that decodes coded data obtained by coding image data, and an image coding device that generates coded data obtained by coding image data, and can also be suitably applied to the data structure of coded data generated by an image coding device and referenced by the image decoding device.

[0288] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority from application Ser. No. 2017-155903, filed Aug. 10, 2017, and application Ser. No. 2018-053226, filed Mar. 20, 2018, the contents of which are incorporated herein by reference. [Explanation of symbols]

[0289] 11 Image encoding device 31 Image Decoding Device 107 CNN Filter (Image Filter Device) 114 CNN Parameter Determination Unit (Parameter Determination Unit)

Claims

1. 1. An image filter device having a neural network filter, The above neural network filter is (i) a first neural network that receives a first image and outputs first output data; (ii) a second neural network that receives the encoding parameters and outputs second output data; and (iii) a combiner that derives combined data from the first output data and the second output data; and (iv) a third neural network that uses the combined data to output third output data; and (v) an adder which receives the third output data and the first image and outputs fourth output data.

2. An image decoding device for decoding an image, comprising: Equipped with a neural network filter, The above neural network filter is (i) a first neural network that receives a first image and outputs first output data; (ii) a second neural network that receives the encoding parameters and outputs second output data; and (iii) a combiner that derives combined data from the first output data and the second output data; and (iv) a third neural network that uses the combined data to output third output data; and (v) an adder which receives the third output data and the first image and outputs fourth output data.

3. An image encoding device that encodes an image, comprising: Equipped with a neural network filter, The above neural network filter is (i) a first neural network that receives a first image and outputs first output data; (ii) a second neural network that receives the encoding parameters and outputs second output data; and (iii) a combiner that derives combined data from the first output data and the second output data; and (iv) a third neural network that uses the combined data to output third output data; and (v) an adder which receives the third output data and the first image and outputs fourth output data.

Citation Information

Patent Citations

  • Video intra-prediction method and device therefor

    JP2017055434A

  • Image coding method, image decoding method, image coding device and image decoding device

    WO2016199330A1