method

By encoding and decoding specific partial image regions within a picture, the challenges of encoding and decoding specific parts of the same screen are addressed, enhancing efficiency and reducing resource usage in video encoding and decoding processes.

JP7851992B2Active Publication Date: 2026-04-27SHARP KK
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
SHARP KK
Filing Date
2024-07-01
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies face challenges in efficiently encoding and decoding specific parts of the same screen, with existing technologies struggling to efficiently encode and decode specific parts of the same screen.

Method used

Implementing a mechanism for encoding and decoding specific partial image regions within a picture, where intra-prediction, inter-prediction, and loop filtering are restricted to these regions, treating areas outside the partial image region as outside the picture, and allowing independent decoding of these regions within the same screen.

Benefits of technology

Enables efficient encoding and decoding of specific partial image regions within a picture, allowing independent decoding and reducing encoding complexity and resource usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007851992000001
    Figure 0007851992000001
  • Figure 0007851992000002
    Figure 0007851992000002
  • Figure 0007851992000003
    Figure 0007851992000003
Patent Text Reader

Abstract

To provide a video coding apparatus and a video decoding apparatus, capable of independently decoding only a specific partial portion on the same screen.SOLUTION: A video decoding apparatus identifies a first picture in a set of pictures as being associated with a gradual refreshed picture; identifies as a sequentially decoder refresh (SDR) network abstraction layer (NAL) unit type and stores the nal_unit_type in a coding stream; stores an enable flag for identifying use of a gradual refresh for decoding a subset of the picture in the set of pictures in the coding stream; and identifies the number of pictures from the first picture until the entire picture is properly decoded, transmits the number of pictures when the enable flag is true, and bypasses the number of pictures from the first picture until the entire picture is properly decoded when the enable flag is false.SELECTED DRAWING: Figure 14
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to a moving image decoding apparatus and a moving image encoding apparatus.

Background Art

[0002] In order to efficiently transmit or record a moving image, a moving image encoding apparatus that generates encoded data by encoding the moving image and a moving image decoding apparatus that generates a decoded image by decoding the encoded data are used.

[0003] Specific moving image encoding methods include, for example, methods proposed in H.264 / AVC (Advanced Video Coding) and H.265 / HEVC (High-Efficiency Video Coding).

[0004] In HEVC, a method of dividing a picture called a tile into a rectangle is introduced. Tiles are mainly intended to divide the screen and perform encoding and decoding in parallel, and intra prediction, motion vector prediction, and entropy encoding operate independently for each tile.

[0005] In addition, Non-Patent Document 1 can be cited as a recent moving image encoding and decoding technology.

Prior Art Documents

Non-Patent Documents

[0006]

Non-Patent Document 1

Non-Patent Document 2

[0007] In tile layouts, intra-prediction and motion vectors within the same screen are restricted, but inter-prediction is not.

[0008] To independently decode only a specific sub-image region within the same screen, interpretation that references areas outside the sub-image region will prevent correct decoding. Therefore, conventional methods have involved restricting the direction of motion vectors on the encoding side. However, recent methods such as HEVC use merged modes or other methods that utilize previously encoded motion vectors, making it difficult to explicitly restrict motion vectors and resulting in a significant decrease in encoding efficiency.

[0009] Therefore, the present invention has been made in view of the above problems, and its objective is to provide a mechanism for realizing video encoding and decoding that can independently decode only a specific part of the same screen. [Means for solving the problem]

[0010] A motion image decoding device according to one aspect of the present invention is characterized in that, for intra prediction, inter prediction, loop filtering, etc., a partial image region is set within the picture, the area outside the partial image region is treated the same as the area outside the picture, and no such restrictions are applied to areas within the picture other than the partial image region. [Effects of the Invention]

[0011] According to one embodiment of the present invention, partial decoding within a picture can be achieved by setting a partial image region within the picture in which prediction processing and loop filtering processing are restricted. [Brief explanation of the drawing]

[0012] [Figure 1] This diagram shows the hierarchical structure of the encoded stream data. [Figure 2] This figure shows an example of CTU partitioning. [Figure 3] This is a conceptual diagram showing an example of a reference picture and a reference picture list. [Figure 4] This is a schematic diagram showing the types (mode numbers) of intra-prediction modes. [Figure 5] This figure illustrates the partial and non-partial image regions of the present invention. [Figure 6] This figure illustrates the scope of the target block of the present invention that can be referenced. [Figure 7] This flowchart shows the decoding process flow of the parameter decoding unit. [Figure 8] This figure shows an example of syntax used to define a partial image region. [Figure 9] This figure shows an example of syntax used to define a partial image region. [Figure 10] This flowchart shows the procedure for setting a partial image area. [Figure 11] This diagram illustrates the settings for a partial image region map. [Figure 12] This is a diagram illustrating a gradual refresh process. [Figure 13] This diagram illustrates the syntax required for a gradual refresh. [Figure 14] This is a schematic diagram showing the configuration of a video decoding device. [Figure 15] This is a block diagram showing the configuration of a video encoding device. [Figure 16]This is a diagram showing the configurations of a transmission device equipped with a moving image encoding device according to this embodiment and a reception device equipped with a moving image decoding device. (a) shows the transmission device equipped with the moving image encoding device, and (b) shows the reception device equipped with the moving image decoding device. [Figure 17] This is a diagram showing the configurations of a recording device equipped with a moving image encoding device according to this embodiment and a playback device equipped with a moving image decoding device. (a) shows the recording device equipped with the moving image encoding device, and (b) shows the playback device equipped with the moving image decoding device. [Figure 18] This is a schematic diagram showing the configuration of an image transmission system according to this embodiment.

Mode for Carrying Out the Invention

[0013] (First Embodiment) Hereinafter, embodiments of the present invention will be described with reference to the drawings.

[0014] FIG. 18 is a schematic diagram showing the configuration of an image transmission system 1 according to this embodiment.

[0015] The image transmission system 1 is a system that transmits an encoded stream obtained by encoding an image to be encoded and decodes the transmitted encoded stream to display an image. The image transmission system 1 includes a moving image encoding device (image encoding device) 11, a network 21, a moving image decoding device (image decoding device) 31, and a moving image display device (image display device) 41.

[0016] An image T is input to the moving image encoding device 11.

[0017] Network 21 transmits the encoded stream Te generated by the video encoding device 11 to the video decoding device 31. Network 21 is the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. Network 21 is not necessarily limited to a bidirectional communication network; it may also be a unidirectional communication network that transmits broadcast waves such as terrestrial digital broadcasting or satellite broadcasting. Furthermore, network 21 may be replaced by a storage medium that records the encoded stream Te, such as a DVD (Digital Versatile Disc) or a BD (Blu-ray Disc).

[0018] The video decoding device 31 decodes each of the encoded streams Te transmitted by the network 21 and generates one or more decoded images Td.

[0019] The video display device 41 displays all or part of one or more decoded images Td generated by the video decoding device 31. The video display device 41 includes a display device such as a liquid crystal display or an organic EL (electro-luminescence) display. Examples of display forms include stationary, mobile, and HMD (head-mounted display). Furthermore, if the video decoding device 31 has high processing power, it displays high-quality images, and if it has lower processing power, it displays images that do not require high processing power or display power.

[0020] <operators> The operators used in this specification are listed below.

[0021] >> is a right bit shift, << is a left bit shift, & is a bitwise AND, | is a bitwise OR, |= is the OR assignment operator, and || represents logical OR.

[0022] x?y:z is a ternary operator that takes the value y if x is true (non-zero) and z if x is false (0).

[0023] Clip3(a, b, c) is a function that clips c to a value between a and b (inclusive). If c < a, it returns a; if c > b, it returns b; otherwise (when a <= b), it returns c.

[0024] abs(a) is a function that returns the absolute value of a.

[0025] Int(a) is a function that returns the integer value of a.

[0026] floor(a) is a function that returns the largest integer less than or equal to a.

[0027] ceil(a) is a function that returns the smallest integer greater than or equal to a.

[0028] a / d represents the division of a by d, rounded down to the nearest integer.

[0029] <Structure of the Encoded Stream Te> Prior to the detailed description of the moving image encoding device 11 and the moving image decoding device 31 according to this embodiment, the data structure of the encoded stream Te generated by the moving image encoding device 11 and decoded by the moving image decoding device 31 will be described.

[0030] Figure 2 is a diagram showing the hierarchical structure of the data in the encoded stream Te. The encoded stream Te illustratively includes a sequence and a plurality of pictures that make up the sequence. Figures 2(a) to 2(f) respectively show an encoded video sequence that defines the sequence SEQ, an encoded picture that defines the picture PICT, an encoded slice that defines the slice S, encoded slice data that defines the slice data, an encoded tree unit included in the encoded slice data, and an encoded unit included in the encoded tree unit.

[0031] (Encoded Video Sequence) In an encoded video sequence, a set of data that the video decoding device 31 references to decode the sequence SEQ to be processed is defined. As shown in Figure 2(b), the sequence SEQ includes a Video Parameter Set, a Sequence Parameter Set SPS, a Picture Parameter Set PPS, a Picture PICT, and Supplemental Enhancement Information SEI.

[0032] The Video Parameter Set (VPS) defines a set of encoding parameters common to multiple video layers in a video composed of multiple layers, as well as a set of encoding parameters associated with the multiple layers included in the video and with each individual layer.

[0033] The sequence parameter set (SPS) defines a set of encoding parameters that the video decoding device 31 references to decode the target sequence. For example, the width and height of the picture are defined. Multiple SPSs may exist. In that case, one of the multiple SPSs is selected from the PPS.

[0034] The Picture Parameter Set (PPS) defines a set of encoding parameters that the video decoding device 31 references to decode each picture in the target sequence. For example, it includes a reference value for the quantization width used for decoding the picture (pic_init_qp_minus26) and a flag indicating the application of weighted prediction (weighted_pred_flag). Multiple PPSs may exist. In that case, one of the multiple PPSs is selected for each picture in the target sequence.

[0035] (Encoded picture) The encoded picture specifies the set of data that the video decoding device 31 references to decode the picture PICT to be processed. The picture PICT includes slices 0 to NS-1, as shown in Figure 2(b) (NS is the total number of slices included in the picture PICT).

[0036] In the following, if it is not necessary to distinguish between slices 0 through NS-1, the code subscripts may be omitted. The same applies to other data included in the coded stream Te described below that have subscripts.

[0037] (Encoded slice) In an encoded slice, a set of data is defined that the video decoding device 31 references to decode the slice S to be processed. As shown in Figure 2(b), the slice includes a slice header and slice data.

[0038] The slice header contains a set of encoding parameters that the video decoding device 31 references to determine the decoding method for the target slice. The slice type specification information (slice_type), which specifies the slice type, is an example of the encoding parameters included in the slice header.

[0039] The slice types that can be specified by the slice type specification information include (1) I slices that use only intra prediction during encoding, (2) P slices that use unidirectional prediction or intra prediction during encoding, and (3) B slices that use unidirectional prediction, bidirectional prediction or intra prediction during encoding. Note that interpretation is not limited to single or bidirectional prediction, and prediction images may be generated using more reference pictures. Hereinafter, when referring to P slices and B slices, we mean slices that contain blocks on which interpretation can be used.

[0040] The slice header may also include a reference to the picture parameter set PPS (pic_parameter_set_id).

[0041] (Encoded slice data) The encoded slice data defines the set of data that the video decoding device 31 references in order to decode the slice data to be processed. As shown in Figure 1(d), the slice data includes CTUs. A CTU is a fixed-size (e.g., 64x64) block that makes up a slice, and is sometimes called a Largest Coding Unit (LCU).

[0042] (Code tree unit) Figure 2(e) defines the set of data that the video decoding device 31 references to decode the CTU to be processed. The CTU is divided into coding units CU, which are the basic units of encoding processing, by recursive quad tree partitioning (QT), binary tree partitioning (BT), or ternary tree partitioning (TT). BT and TT partitioning together are called multi-tree partitioning (MT). The nodes of the tree structure obtained by recursive quad tree partitioning are called coding nodes. The intermediate nodes of quad trees, binary trees, and ternary trees are coding nodes, and the CTU itself is defined as the top-level coding node.

[0043] The CT (Computer Transcription) includes the following information: a QT splitting flag (cu_split_flag) indicating whether or not to perform QT ​​splitting; an MT splitting mode (split_mt_mode) indicating the splitting method for MT splitting; an MT splitting direction (split_mt_dir) indicating the splitting direction for MT splitting; and an MT splitting type (split_mt_type) indicating the splitting type for MT splitting. cu_split_flag, split_mt_flag, split_mt_dir, and split_mt_type are transmitted for each encoding node.

[0044] If cu_split_flag is 1, the coding node is split into four coding nodes (Figure 2(b)). If cu_split_flag is 0, or if split_mt_flag is 0, the coding node is not split and has one CU as a node (Figure 2(a)). The CU is the terminal node of the coding node and cannot be split further. The CU is the basic unit of the coding process.

[0045] When split_mt_flag is 1, the encoding node is split into MT as follows: When split_mt_type is 0 and split_mt_dir is 1, the encoding node is horizontally split into 2 encoding nodes (Figure 2(d)), and when split_mt_dir is 0, the encoding node is vertically split into 2 encoding nodes (Figure 2(c)). Also, when split_mt_type is 1 and split_mt_dir is 1, the encoding node is horizontally split into 3 encoding nodes (Figure 2(f)), and when split_mt_dir is 0, the encoding node is vertically split into 3 encoding nodes (Figure 2(e)).

[0046] Furthermore, if the size of the CTU is 64x64 pixels, the size of the CU can be any of the following: 64x64 pixels, 64x32 pixels, 32x64 pixels, 32x32 pixels, 64x16 pixels, 16x64 pixels, 32x16 pixels, 16x32 pixels, 16x16 pixels, 64x8 pixels, 8x64 pixels, 32x8 pixels, 8x32 pixels, 16x8 pixels, 8x16 pixels, 8x8 pixels, 64x4 pixels, 4x64 pixels, 32x4 pixels, 4x32 pixels, 16x4 pixels, 4x16 pixels, 8x4 pixels, 4x8 pixels, and 4x4 pixels.

[0047] (Encoding Unit) As shown in Figure 1(f), the set of data that the video decoding device 31 references to decode the encoding unit to be processed is defined. Specifically, the CU consists of a CU header CUH, prediction parameters, transformation parameters, quantization transformation coefficients, etc. The CU header defines the prediction mode, etc.

[0048] Prediction processing can be performed at the CU (Unit) level or at the subCU level, which is a further division of the CU. If the size of the CU and the subCU are equal, there is one subCU within the CU. If the CU is larger than the size of the subCU, the CU is divided into subCUs. For example, if the CU is 8x8 and the subCU is 4x4, the CU is divided into four subCUs, each consisting of two horizontal and two vertical divisions.

[0049] There are two types of predictions (prediction modes): intra-prediction and inter-prediction. Intra-prediction is prediction within the same picture, while inter-prediction refers to prediction processing performed between different pictures (for example, between display times or between layer images).

[0050] The transformation and quantization processes are performed in units of CUs, but the quantization transformation coefficients may be entropy-encoded in subblock units such as 4x4.

[0051] (Prediction parameters) The predicted image is derived from the prediction parameters associated with the block. These prediction parameters include intra-prediction and inter-prediction parameters.

[0052] The following describes the prediction parameters for interpretation. The interpretation parameters consist of the prediction list usage flags predFlagL0 and predFlagL1, the reference picture indices refIdxL0 and refIdxL1, and the motion vectors mvL0 and mvL1. The prediction list usage flags predFlagL0 and predFlagL1 indicate whether or not the reference picture lists called L0 list and L1 list are used, respectively. If the value is 1, the corresponding reference picture list is used. In this specification, when referring to "flags indicating whether or not XX is true," a flag value other than 0 (e.g., 1) is considered true, and 0 is considered false. In logical negation, logical AND, etc., 1 is treated as true and 0 as false (the same applies below). However, in actual devices and methods, other values ​​may be used as true and false values.

[0053] Syntax elements for deriving interprediction parameters include, for example, the merge flag merge_flag, merge index merge_idx, interprediction identifier inter_pred_idc, reference picture index refIdxLX, prediction vector index mvp_LX_idx, and difference vector mvdLX.

[0054] (Reference picture list) The reference picture list is a list of reference pictures stored in the reference picture memory 306. Figure 4 is a conceptual diagram showing an example of a reference picture and reference picture list in a picture structure for low latency. In Figure (a), rectangles represent pictures, arrows represent the reference relationships between pictures, the horizontal axis represents time, I, P, and B in the rectangles represent intra-picture, single-prediction picture, and double-prediction picture, respectively, and the numbers in the rectangles represent the decoding order. As shown in the figure, the decoding order of the pictures is I0, P1 / B1, P2 / B2, P3 / B3, P4 / B4, and the display order is the same. Figure (b) shows an example of the reference picture list for picture B3 (target picture). The reference picture list is a list that represents candidates for reference pictures, and one picture (slice) may have one or more reference picture lists. In the example in the figure, the target picture B3 has two reference picture lists: L0 list RefPicList0 and L1 list RefPicList1. Each CU specifies which picture in the reference picture list RefPicListX (X=0 or 1) to actually reference using the reference picture index refIdxLX. The figure shows an example where refIdxL0=2 and refIdxL1=0. If the target picture is P3, the reference picture list is only the L0 list. Note that LX is a notation used when there is no distinction between L0 prediction and L1 prediction, and from now on, parameters for the L0 list and parameters for the L1 list will be distinguished by replacing LX with L0 and L1.

[0055] (Merge prediction and AMVP prediction) There are two methods for decoding (encoding) prediction parameters: merge prediction mode and AMVP (Adaptive Motion Vector Prediction) mode. The merge flag, merge_flag, is used to distinguish between these two modes.

[0056] The merge prediction mode is a mode in which the prediction list usage flag predFlagLX (or inter-prediction identifier inter_pred_idc), reference picture index refIdxLX, and motion vector mvLX are not included in the encoded data, but are derived from the prediction parameters of already processed neighboring blocks. The merge index merge_idx is an index that indicates which of the prediction parameter candidates (merge candidates) derived from the processed block will be used as the prediction parameter for the target block.

[0057] AMVP mode is a mode in which the inter-prediction identifier inter_pred_idc, reference picture index refIdxLX, and motion vector mvLX are included in the encoded data. The motion vector mvLX is encoded as a prediction vector index mvp_LX_idx that identifies the prediction vector mvpLX and a difference vector mvdLX. The inter-prediction identifier inter_pred_idc is a value that indicates the type and number of reference pictures, and can take one of the values ​​PRED_L0, PRED_L1, or PRED_BI. PRED_L0 and PRED_L1 indicate single prediction using one reference picture managed in the L0 list and L1 list, respectively. PRED_BI indicates bi-prediction BiPred using two reference pictures managed in the L0 list and L1 list.

[0058] (Motion vector) The motion vector mvLX represents the amount of shift between blocks in two different pictures. The prediction vector and difference vector for the motion vector mvLX are called the prediction vector mvpLX and the difference vector mvdLX, respectively.

[0059] The following describes the prediction parameters for intra-prediction. The intra-prediction parameters consist of the luminance prediction mode (IntraPredModeY) and the color difference prediction mode (IntraPredModeC). Figure 5 is a schematic diagram showing the types (mode numbers) of intra-prediction modes. As shown in the figure, there are, for example, 67 types (0 to 66) of intra-prediction modes. For example, Planar prediction (0), DC prediction (1), and Angular prediction (2 to 66). Furthermore, LM modes (67 to 72) may be added for color difference.

[0060] Syntax elements for deriving intra prediction parameters include, for example, prev_intra_luma_pred_flag, mpm_idx, rem_selected_mode_flag, rem_selected_mode, and rem_non_selected_mode.

[0061] (MPM) `prev_intra_luma_pred_flag` is a flag indicating whether the target block's luminance prediction mode `IntraPredModeY` matches its MPM (Most Probable Mode). `MPM` is a prediction mode included in the MPM candidate list `mpmCandList[]`. The MPM candidate list stores candidates that are estimated to have a high probability of being applied to the target block, based on the intra-prediction modes of adjacent blocks and a given intra-prediction mode. If `prev_intra_luma_pred_flag` is 1, the target block's luminance prediction mode `IntraPredModeY` is derived using the MPM candidate list and the index `mpm_idx`.

[0062] IntraPredModeY = mpmCandList[mpm_idx] (REM) If prev_intra_luma_pred_flag is 0, an intra-prediction mode is selected from the remaining modes, RemIntraPredMode, which excludes the intra-prediction modes included in the MPM candidate list from the entire intra-prediction mode set. Intra-prediction modes that can be selected as RemIntraPredMode are called "non-MPM" or "REM". The flag rem_selected_mode_flag specifies whether to select an intra-prediction mode by referring to rem_selected_mode or by referring to rem_non_selected_mode. RemIntraPredMode is derived using rem_selected_mode or rem_non_selected_mode.

[0063] (Partial image region encoding / decoding region) This invention describes a video encoding and decoding method characterized by setting a partial image region within the same picture, performing encoding and decoding on the partial image region without using pixels from the other regions, and performing encoding and decoding on the other regions using the entire picture.

[0064] Figure 5 illustrates regions A and B of the present invention. In the video encoding and decoding device of the present invention, regions A and B are defined within a picture. For example, regions A and B are defined by a partial image region control unit, which will be described later. Region A can only be predicted from within region A, and the area outside the region is subjected to processing such as padding, similar to the area outside a picture or tile. On the other hand, region B can be predicted from the entire picture, including region A. Prediction processing here refers to intra-prediction, inter-prediction, loop filtering, etc. Because the encoding and decoding processes are closed within region A, only region A can be decoded.

[0065] Hereafter, region A will be referred to as the partial image region (first region, control region, clean region, refreshed region, region A). Conversely, regions other than the partial image region will also be referred to as the non-partial image region (second region, uncontrolled region, dirty region, unrefreshed region, region B, outside the restricted region).

[0066] For example, a region that is encoded and decoded solely from intra-predictions, and which has already been encoded in the intra-prediction (the new refresh region IRA consisting only of intra-predictions, as described later), is a partial image region. A region that is encoded and decoded by further referencing this partial image region composed of intra-predictions is also a partial image region. Furthermore, a region that is encoded and decoded by referencing a partial image region in a reference picture, such as in inter-prediction, is also a partial image region. In short, a partial image region is a region that is encoded and decoded by referencing only the pixels of the partial image region, without referencing pixels of the non-partial image region.

[0067] In the following, the top-left position of a partial image region is denoted by (xRA_st, yRA_st), the bottom-right position by (xRA_en, yRA_en), and the size by (wRA, hRA). Furthermore, since the position and size have the following relationship, one can be derived from the other.

[0068] xRA_en = xRA_st + wRA - 1 yRA_en = yRA_st + hRA - 1 It can also be derived as follows.

[0069] wRA = xRA_en - xRA_st + 1 hRA = yRA_en - yRA_st + 1 Furthermore, the upper-left position of the time-limiting reference region j is indicated by (xRA_st[j], yRA_st[j]), the lower-right position by (xRA_en[j], yRA_en[j]), and the size by (wRA[j], hRA[j]). Alternatively, the position of the limiting reference region of the referenced picture Ref may be indicated by (xRA_st[Ref], yRA_st[Ref]), the lower-right position by (xRA_en[Ref], yRA_en[Ref]), and the size by (wRA[Ref], hRA[Ref]).

[0070] (Determination of partial image regions) For example, if a picture is at time i and a block is at position (x, y), the following formula can be used to determine whether the pixel at position is within the sub-image region.

[0071] IsRA(x, y) = (xRA_st[i] <= x && x <= xRA_en[i] && yRA_st[i] <= y && y <= yRA_en[i]) Alternatively, the following judgment formula may also be used.

[0072] IsRA(x, y) = (xRA_st[i] <= x && x < xRA_st[i]+wRA[i] && yRA_st[i] <= y && y < yRA_st[i]+hRA[i]) IsRA(xRef, yRef) = (xRA_st[Ref] <= xRef && xRef <= xRA_en[Ref] && yRA_st[Ref] <= yRef && yRef <= yRA_en[Ref]) For example, if the target picture is at time i, the top-left coordinates of the target block Pb are (xPb, yPb), and the width and height are bW and bH, the intra-prediction unit, motion compensation unit, and loop filter of the video decoding and video encoding devices derive IsRA(Pb) using the following determination formula if the target block Pb is within a partial image region.

[0073] IsRA(Pb) = (xRA_st[i] <= xPb && xPb <= xRA_en[i] && yRA_st[i] <= yPb && yPb <= yRA_en[i]) Alternatively, the following judgment formula may also be used.

[0074] IsRA(Pb) = (xRA_st[i] <= xPb && xPb < xRA_st[i]+wRA[i] && yRA_st[i] <= yPb && yPb < yRA_st[i]+hRA[i]) (Basic operation of the reference area of ​​a partial image region) The video encoding device and video decoding device described herein perform the following operations.

[0075] Figure 6 shows the range that a partial image region can reference in the intra-prediction, inter-prediction, and loop filter of the present invention. Figure 6(a) shows the range that a target block included in a partial image region can reference. In the picture in Figure 6(a), the area enclosed by a thick line is an already encoded / decoded area included in the partial image region. The already encoded / decoded area included in the partial image region of the same picture (target image i) as the target block is the range that the target block can reference in intra-prediction, inter-prediction, and loop filter. Similarly, the partial image region in the reference picture (reference image j) is the range that the target block can reference in inter-prediction and loop filter. Figure 6(b) shows the range that a target block included in a non-partial image region can reference. In the picture in Figure 6(b), the area enclosed by a thick line is an already encoded / decoded area in the target picture. The already encoded or decoded area in the target picture (target image i) is the range that the target block can reference in intra-prediction and inter-prediction. Similarly, all areas within the reference picture (reference image j) are within the range that can be referenced by interpretation. Note that when using parallel processing or reference restrictions such as tiling, slicing, or wavefronts, additional restrictions may be added on top of the above. • Target blocks included in a partial image region will undergo either intra-prediction, which refers only to pixels in the partial image region within the target picture, or inter-prediction, which refers to the limited reference region of the reference picture. The encoding parameters of the target block are derived by referring to the encoding parameters of the target block within the partial image region (e.g., intra-prediction direction, motion vector, reference picture index) in the target picture, or by referring to the encoding parameters of the limiting reference region of the reference picture. • For target blocks included in a partial image region, the loop filtering process is performed by referencing only the pixels of that partial image region in the target picture.

[0076] (Determination and availability of partial image regions) In intra-prediction MPM derivation and inter-prediction merge candidate derivation, prediction parameters (intra-prediction mode, motion vector) of the target block are sometimes derived using prediction parameters of adjacent regions. In such cases, the following processing may be performed: In intra-prediction and inter-prediction, if the target block is a partial image region (IsRA(xPb, yPb) is true) and the reference position (xNbX, yNbX) of the target block's adjacent block is not a partial image region (IsRA(xNbX, yNbX) is false), the values ​​of the adjacent block are not used in the prediction parameter derivation. That is, if the target block is a partial image region (IsRA(xPb, yPb) is true) and the reference position (xNbX, yNbX) of the target block's adjacent block is also a partial image region (IsRA(xNbX, yNbX) is true), then that position (xNbX, yNbX) is used in the prediction parameter derivation.

[0077] As explained above in the derivation of prediction candidates, the determination of a partial image region may also be used for determining areas outside the screen, similar to the determination of areas outside the screen or parallel processing units (slice boundaries, tile boundaries). In this case, if the target block is a partial image region (IsRA(xPb, yPb) is true) and the reference position of the target block (xNbX, yNbX) is also a partial image region (IsRA(xNbX, yNbX) is true), then the reference position (xNbX, yNbX) is determined to be unreachable (availableNbX=0). In other words, the reference position (xNbX, yNbX) is determined to be accessible (availableNbX=1) if the target block is within the screen, and the reference position is not in a different parallel processing unit with the same reference position as the target block, and the target block is in a non-partial image region, or if the reference position (xNbX, yNbX) of the target block is in a partial image region (IsRA(xNbX, yNbX) is true). In intra-prediction and inter-prediction, if the reference position (xNbX, yNbX) is accessible (availableNbX=1), the prediction parameters of that reference position are used to derive the prediction parameters of the target block.

[0078] (Determining the restricted reference area and clipping the restricted reference area) Furthermore, if the reference picture is at time j and the top-left position of the reference pixel is (xRef, yRef), the motion compensation unit derives the condition that the reference pixel is within the limited reference area using the following determination formula.

[0079] IsRA(xRef, yRef) = (xRA_st[j] <= xRef && xRef <= xRA_en[j] && yRA_st[j] <= yRef && yRef <= yRA_en[j] ) Alternatively, the following judgment formula may also be used.

[0080] IsRA(xRef, yRef) = (xRA_st[j] <= xRef && xRef < xRA_st[j]+wRA[j] && yRA_st[i] <= yRef && yRef < yRA_st[j]+hRA[j]) Furthermore, the motion compensation unit may clip the reference pixel to a position within the partial image area using the following formula.

[0081] xRef = Clip3(xRA_st[j], xRA_en[j], xRef) yRef = Clip3(yRA_st[j], yRA_en[j], yRef) Alternatively, the following derivation formula may also be used.

[0082] xRef = Clip3(xRA_st[j], xRA_st[j]+wRA[j]-1, xRef) yRef = Clip3(yRA_st[j], yRA_st[j]+hRA[j]-1, yRef) The position of the partial image region is transmitted from the video encoding device to the video decoding device using the stepwise refresh information described later. The position and size of the partial image region may not be derived according to time (e.g., POC), but rather after decoding the target picture, or at the start of decoding the target picture, by setting a reference picture Ref in the reference memory. In this case, the position and size of the partial image region can be derived by specifying the reference picture Ref.

[0083] (SDR picture) In AVC and HEVC, IDR (Instantaneous Decoder Refresh) pictures have the entire picture as an intra-CTU, enabling random access to encoded data as a picture that can be randomly accessed and decoded independently. In this embodiment, a picture in which the entire partial image region is intra-encoded is made an SDR (Sequentially Decoder Refresh) picture and can be identified by the nal_unit_type of the NAL (Network Abstraction Layer).

[0084] In SDR pictures, it is possible to decode individual areas of the picture independently, and random access is possible to these areas. Unlike conventional IDR pictures, where the entire picture is intranet, SDR pictures have only a portion of the picture as intranet, resulting in smaller fluctuations in the amount of encoding.

[0085] (Parameter decoding unit 302) The parameter decoding unit 302 sets a partial image region in the SDR picture, for example, as follows: - The partial image area is defined as a rectangle specified by the coordinates of the top-left CTU and the number of CTUs for width and height. - A partial image area is defined as a rectangle, specified by the pixel position in the upper left corner and the number of pixels in its width and height. • Set multiple sub-image regions within a single picture. • Set the partial image regions so that multiple partial image regions overlap each other.

[0086] The statement that multiple sub-image regions overlap means, for example, that multiple sub-image regions contained within a single picture may contain the same CTU (Critical Point Unit).

[0087] Furthermore, the partial image regions of multiple pictures in a GOP (Group of Picture) may overlap. Here, overlapping partial image regions means that the partial image region set in an SDR picture and the partial image region set in the next picture after the SDR picture contain the same CTU. The number of pictures in which partial image regions overlap is not particularly limited and is simply multiple pictures in the GOP that are consecutive to the SDR picture.

[0088] (Processing flow by parameter decoding unit 302 (SDR picture)) Figure 8 is a flowchart showing the processing flow performed by the parameter decoding unit 302.

[0089] (Step S1) Start decryption and proceed to step S2.

[0090] (Step S2) The parameter decoding unit 302 determines whether the target picture is an SDR picture using the nal_unit_type of the NAL. If it is an SDR picture, proceed to S3; otherwise, proceed to S4.

[0091] (Step S3) The partial image region contained within the target picture is set as the region to be decoded using intra-prediction, and the process proceeds to S4.

[0092] (Step S4) The parameter decoding unit 302 decodes the target picture.

[0093] By setting the partial image region in this way, the video decoding device 31 can decode only the partial image region of the picture that is continuous with the SDR picture.

[0094] (Example of region information 1) The syntax for setting a partial image region may be included in the picture parameter set. Figure 8 shows an example of the syntax notified for setting a partial image region. partial_region_mode is information that identifies whether or not to define a partial image region in the picture. The entropy decoding unit 301 of the video decoding device 31 determines that setting a partial image region is necessary if partial_region_mode included in the picture parameter set is 1, and decodes num_of_patial_region_minus1.

[0095] num_of_patial_region_minus1 indicates the number of partial image regions in the picture minus 1. position_ctu_adress[i] indicates the address of the top-left CTU of the i-th partial image region among multiple regions in the picture. region_ctu_width_minus1[i] indicates the number of horizontal CTUs of the i-th partial image region among multiple regions in the picture minus 1. region_ctu_height_minus1[i] indicates the number of vertical CTUs of the i-th partial image region among multiple regions in the picture minus 1.

[0096] The entropy decoding unit 301 adds 1 to i until i is equal to the value of num_of_patial_region_minus1, and decodes position_ctu_adress[i], region_ctu_width_minus1[i], and region_ctu_height_minus1[i].

[0097] Then, the partial image region control unit 320 of the video decoding device 31 performs the following for each i: position_ctu_adress[i] region_ctu_width_minus1[i] region_ctu_height_minus1[i] A partial image region having the position and size specified by [the specified method] is set within the target picture.

[0098] In addition, num_of_patial_region_minus1 position_ctu_adress[i] region_ctu_width_minus1[i] region_ctu_height_minus1[i] This is an example of region information used to identify a sub-image area.

[0099] (Example of region information 2) The syntax for setting a partial image region may be included in the slice header. Figure 9 shows an example of the syntax notified for setting a partial image region. first_slice_segment_in_pic_flag is a flag that indicates whether the slice is the first slice in the decoding order. If first_slice_segment_in_pic_flag is 1, it indicates that it is the first slice. If first_slice_segment_in_pic_flag is 0, it indicates that it is not the first slice. If first_slice_segment_in_pic_flag is 1, the entropy decoding unit 301 of the video decoding device 31 sets partial_region_mode and decodes num_of_patial_region_minus1.

[0100] num_of_patial_region_minus1 indicates the number of partial image regions in the slice minus 1. position_ctu_adress[i] indicates the address of the top-left CTU of the i-th partial image region among multiple regions in the slice. region_ctu_width_minus1[i] indicates the number of horizontal CTUs of the i-th partial image region among multiple regions in the slice minus 1. region_ctu_height_minus1[i] indicates the number of vertical CTUs of the i-th partial image region among multiple regions in the slice minus 1.

[0101] The entropy decoding unit 301 adds 1 to i until i is equal to the value of num_of_patial_region_minus1, and decodes position_ctu_adress[i], region_ctu_width_minus1[i], and region_ctu_height_minus1[i].

[0102] Then, the partial image region control unit 320 of the video decoding device 31 performs the following for each i: position_ctu_adress[i] region_ctu_width_minus1[i] region_ctu_height_minus1[i] A partial image region having the position and size identified by [the specified method] is set within the target slice.

[0103] In addition, num_of_patial_region_minus1 position_ctu_adress[i] region_ctu_width_minus1[i] region_ctu_height_minus1[i] This is an example of region information used to identify a sub-image area.

[0104] In the example above, 1 CTU is used as the smallest unit, but you may also set one or more CTU columns, or one or more CTU rows and multiple CTUs as the smallest unit.

[0105] (Example of procedure for setting a partial image area) Figure 10 is a flowchart showing the processing flow performed by the video decoding device 31 when a partial image region is defined in the picture parameter set.

[0106] (Step S1) The decryption process begins and the process proceeds to step S2.

[0107] (Step S2) The entropy decoding unit 301 proceeds to step S3 if it is in partial_region_mode (when partial_region_mode is 1), and to step S4 if it is not in partial_region_mode (when partial_region_mode is 0).

[0108] (Step S3) In partial_region_mode, the entropy decoding unit 301 decodes each syntax included in the region information, and the partial image region control unit 320 defines the partial image region specified by each syntax and terminates the process. The specific process for setting the partial image region is as described above.

[0109] (Step S4) If the mode is not partial_region_mode, the video decoding device 31 erases the partial image region and terminates the process.

[0110] (Partial image region map) The parameter decoding unit 302 may be configured to set a partial image region map (partial_region_map) as information representing the position of a partial image region for each picture.

[0111] Figure 11 shows an example of syntax notified to set a partial image region. partial_region_map is syntax that indicates whether or not each CTU in the picture is a partial screen region. The entropy decoding unit 301 of the video decoding device 31 determines that if partial_region_map is 1, it is a partial image region, and if partial_region_map is 0, it is a non-partial image region. partial_region_mode is information used to determine whether or not to define a partial image region in the picture.

[0112] PicHeightInCtbsY indicates the vertical CTU count of the picture, and PicWidthInCtbsY indicates the horizontal CTU count of the picture.

[0113] Increment i by 1 until i equals num_of_partial_region_minus1+1, and calculate y=position_ctu_adress[i] / PicWidthInCtbsY and x=position_ctu_adress[i]%PicWidthInCtbsY.

[0114] Increment j by 1 until j is equal to region_ctu_width_minus1[i], and increment k by 1 until k is equal to region_ctu_height_minus1[i], then set the corresponding partial_region_map[h+j][w+k] to 1.

[0115] The partial image region control unit 320 of the parameter decoding unit 302 may be configured to set the partial image region by referring to the partial_region_map generated in this way.

[0116] Furthermore, the partial_region_map information, which represents the location of a partial image region and is saved for each picture, is managed in the DPB (Decoder Picture Buffer) of the decoded picture memory. In addition, the partial_region_map information is stored in the reference picture list of the reference picture memory 306 for use in interpretation performed by the prediction image generation unit 308.

[0117] Here, the order of decoding and encoding of the CTUs is as follows: the video decoding device 31 decodes the CTUs in the raster scan order on a picture or tile basis without distinguishing between CTUs in partial image areas and CTUs in non-partial image areas, and the video encoding device 11 encodes the CTUs in the raster scan order without distinguishing between CTUs in partial image areas and CTUs in non-partial image areas.

[0118] Furthermore, the entropy coding unit 104 performs entropy coding without distinguishing between partial and non-partial image regions, and the entropy decoding unit 301 performs entropy decoding independently for the partial and non-partial image regions. More specifically, the entropy coding unit 104 and the entropy decoding unit 301 are configured to continuously update the context for the partial and non-partial image regions.

[0119] Since the concepts of partial image regions and non-partial image regions described in this embodiment are independent of the decoding and encoding order, the decoding and encoding order of the CTU may be independent of the partial image regions and non-partial image regions.

[0120] For example, the video decoding device 31 may be configured to decode CTUs in the partial image area and CTUs in the non-partial image area independently in raster scan order, and the video encoding device 11 may be configured to encode CTUs in the partial image area and CTUs in the non-partial image area independently in raster scan order.

[0121] Alternatively, the entropy coding unit 104 may independently perform entropy coding on the partial image region and the non-partial image region, and the entropy decoding unit 301 may independently perform entropy decoding on the partial image region and the non-partial image region. More specifically, the entropy coding unit 104 and the entropy decoding unit 301 may be configured to independently update the context on the partial image region and the non-partial image region.

[0122] (Decoding of a partial image region) With the above configuration, partial image regions are initialized using an SDR picture. The video encoding device 11 encodes the video signal with temporally consecutive partial image regions to create a bitstream. The video decoding device 31 first finds an SDR picture from the nal_unit_type of the NAL in the bitstream, and then performs intra-coding and loop filtering on the partial image region of the SDR picture without referencing the non-partial region. Therefore, the partial image region can be correctly decoded. Subsequently, in the case of inter-coding, the partial image region of the decoded picture does not refer to the non-partial image region, and intra-coding and loop filtering do not refer to the non-partial image region of the picture, thus guaranteeing that the partial image region can be correctly decoded.

[0123] (Second embodiment) (Gradual refresh) This section describes an embodiment of the present invention when the partial image region encoding and decoding method is applied to intra-refresh. Generally, intra-refresh is a method in which a region to be intra-encoded is set within a part of a picture, and that region is moved within the picture over time so that the entire picture can be intra-encoded within a certain period of time. By dividing the picture into sections within a certain period of time and performing intra-encoding, the aim is to intra-encode the entire picture without increasing the encoding amount of a particular picture, to achieve random access, and to enable error recovery in the event of an error in the bitstream. In this embodiment, partial screen region encoding and decoding are performed, an SDR picture is used, and a stepwise refresh function equivalent to intra-refresh is realized.

[0124] Figure 12(a) is a diagram illustrating the overview of the stepwise refresh in this embodiment. In the stepwise refresh of this embodiment, first, the parameter coding unit 111 sets a partial image region A in a part of the picture, and starts from the SDR picture obtained by intra-coding that partial image region. The stepwise refresh is completed when the partial image region A includes the previous partial image region in time and the region increases until the partial image region A becomes the entire picture.

[0125] In the video decoding device 31, the SDR picture is used as an access point, and decoding begins from the bitstream. By decoding until area A becomes the entire picture, the entire picture can be correctly decoded.

[0126] The method for setting the partial image region may be explicitly defined in the PPS or slice header as described in the first embodiment, or, if the seq_refresh_enable_flag described later is 1, after setting the partial image region in the SDR picture, the partial image region may be implicitly increased by 1 CTU column, 1 CTU row, or 1 CTU at a time in the encoding order of the picture.

[0127] In addition, with gradual refresh, for non-referenced pictures that reference other pictures but are not referenced by other pictures, it is not necessary to set a partial image area.

[0128] Figure 12(b) is a diagram illustrating an overview of another stepwise refresh in this embodiment. In the stepwise refresh of this embodiment, first, the parameter coding unit 111 sets a partial image region A in a part of the picture, and starts from the SDR picture obtained by intra-coding that partial image region, and the partial image region A increases in size over time, including the previous partial image region. At this time, the increased partial image region is intra-coded. The stepwise refresh is considered to end when the partial image region A becomes the entire picture. Since the inter-prediction predicted in the time direction is difficult to obtain for the increased partial image region, the partial image region may be coded by referring to coding parameters.

[0129] In the video decoding device 31, the SDR picture is used as an access point, and decoding begins from the bitstream. By decoding until area A becomes the entire picture, the entire picture can be correctly decoded.

[0130] Figure 13 shows an example of syntax notified to achieve stepwise refresh. Figure 13 shows the syntax (stepwise refresh information) notified in the sequence parameter set (SPS). seq_refresh_enable_flag is a flag indicating whether or not to use stepwise refresh for pictures after the SDR picture. The parameter decoding unit 302 decodes the stepwise refresh information, and the video decoding device 31 decodes using stepwise refresh if the seq_refresh_enable_flag flag is 1, and does not use stepwise refresh if it is 0. The parameter decoding unit 302 decodes seq_refresh_period if seq_refresh_enable_flag is 1. seq_refresh_period indicates the number of pictures from the SDR picture, which is a random access point, until the entire picture can be correctly decoded. Note that the number of non-referenced pictures does not need to be counted at this time.

[0131] (Configuration of the video decoding device) The configuration of the video decoding device 31 (Figure 14) according to this embodiment will be described below.

[0132] The video decoding device 31 includes an entropy decoding unit 301, a parameter decoding unit (predictive image decoding device) 302, a loop filter 305, a reference picture memory 306, a prediction parameter memory 307, a prediction image generation unit (predictive image generation device) 308, an inverse quantization / inverse transform unit 311, and an addition unit 312. Note that, in accordance with the video encoding device 11 described later, there is also a configuration in the video decoding device 31 that does not include the loop filter 305.

[0133] The parameter decoding unit 302 includes a partial image region control unit 320, which includes a header decoding unit 3020 (not shown), a CT information decoding unit 3021, and a CU decoding unit 3022 (prediction mode decoding unit). The CU decoding unit 3022 further includes a TU decoding unit 3024. The header decoding unit 3020 decodes parameter set information such as VPS, SPS, and PPS from the encoded data. The header decoding unit 3020 decodes the slice header (slice information) from the encoded data. The CT information decoding unit 3021 decodes the CT from the encoded data. The CU decoding unit 3022 decodes the CU from the encoded data. The TU decoding unit 3024 decodes the QP update information (quantization correction value) and quantization prediction error (residual_coding) from the encoded data if the TU contains a prediction error.

[0134] Furthermore, the parameter decoding unit 302 is configured to include an inter-prediction parameter decoding unit 303 and an intra-prediction parameter decoding unit 304 (not shown). The prediction image generation unit 308 is configured to include an inter-prediction image generation unit 309 and an intra-prediction image generation unit 310.

[0135] Furthermore, while the following examples use CTU and CU as processing units, processing may be performed in sub-CU units, not limited to these examples. Alternatively, CTU, CU, and TU may be replaced with blocks, and sub-CU with sub-blocks, and processing may be performed in block or sub-block units.

[0136] The entropy decoding unit 301 performs entropy decoding on the encoded stream Te input from an external source to separate and decode individual codes (syntax elements). Entropy coding can be performed in two ways: one method uses a context (probability model) adaptively selected according to the type of syntax element and the surrounding circumstances to encode the syntax elements to a variable length, and the other uses a predetermined table or formula to encode the syntax elements to a variable length. A representative example of the former is CABAC (Context Adaptive Binary Arithmetic Coding). The separated codes include prediction information for generating a predicted image and prediction errors for generating a difference image.

[0137] The entropy decoding unit 301 outputs a portion of the separated codes to the parameter decoding unit 302. This portion of the separated codes includes, for example, the prediction mode (predMode), merge flag (merge_flag), merge index (merge_idx), inter-prediction identifier (inter_pred_idc), reference picture index (refIdxLX), prediction vector index (mvp_LX_idx), and difference vector (mvdLX). The control of which codes to decode is performed based on the instructions of the parameter decoding unit 302. The entropy decoding unit 301 outputs the quantization conversion coefficients to the inverse quantization / inverse conversion unit 311.

[0138] The loop filter 305 is a filter installed within the encoding loop that removes block distortion and ringing distortion, thereby improving image quality. The loop filter 305 applies filters such as the deblocking filter 3051, sample-adaptive offset (SAO), and adaptive loop filter (ALF) to the decoded CU image generated by the summing unit 312.

[0139] The reference picture memory 306 stores the decoded image of the CU generated by the adder 312 at a predetermined location for each target picture and target CU.

[0140] The prediction parameter memory 307 stores prediction parameters at predetermined locations for each CTU or CU to be decoded. Specifically, the prediction parameter memory 307 stores the parameters decoded by the parameter decoding unit 302 and the prediction mode predMode separated by the entropy decoding unit 301.

[0141] The prediction image generation unit 308 receives the prediction mode (predMode), prediction parameters, etc. as input. The prediction image generation unit 308 also reads a reference picture from the reference picture memory 306. The prediction image generation unit 308 generates a block or subblock prediction image using the prediction parameters and the read reference picture (reference picture block) in the prediction mode indicated by the prediction mode (intra prediction, inter prediction). Here, a reference picture block is a collection of pixels on the reference picture (usually rectangular, hence called a block), and is the region referenced to generate the prediction image.

[0142] (Interpretation image generation unit 309) When the prediction mode predMode indicates inter-prediction mode, the inter-prediction image generation unit 309 generates a block or sub-block prediction image by inter-prediction using the inter-prediction parameters input from the inter-prediction parameter decoding unit 303 and a reference picture.

[0143] (Motion compensation) The motion compensation unit 3091 (interpolation image generation unit) generates an interpolation image (motion-compensated image) by reading a block from the reference picture memory 306 that is located at a position shifted by the motion vector mvLX from the position of the target block in the reference picture RefLX specified by the reference picture index refIdxLX, based on the inter-prediction parameters (prediction list usage flag predFlagLX, reference picture index refIdxLX, motion vector mvLX) input from the inter-prediction parameter decoding unit 303. Here, if the precision of the motion vector mvLX is not integer precision, a filter called a motion compensation filter is applied to generate pixels at decimal positions to generate the motion-compensated image.

[0144] The motion compensation unit 3091 first derives the integer position (xInt, yInt) and phase (xFrac, yFrac) corresponding to the coordinates (x, y) within the prediction block using the following formula.

[0145] xInt = xPb+(mvLX[0]>>(log2(MVBIT)))+x xFrac = mvLX[0]&(MVBIT-1) yInt = yPb+(mvLX[1]>>(log2(MVBIT)))+y yFrac = mvLX[1]&(MVBIT-1) Here, (xPb, yPb) is the top-left coordinate of a block of size wPb*hPb, x=0...wPb-1, y=0...hPb-1, and MVBIT represents the precision of the motion vector mvLX (1 / MVBIT pixel precision).

[0146] The motion compensation unit 3091 derives a temporary image temp[][] by performing horizontal interpolation on the reference picture refImg using an interpolation filter. The following Σ is the sum with respect to k of k=0..NTAP-1, shift1 is a normalization parameter that adjusts the range of values, and offset1=1<<(shift1-1).

[0147] temp[x][y] = (ΣmcFilter[xFrac][k]*refImg[xInt+k-NTAP / 2+1][yInt]+offset1)>>shift1 Next, the motion compensation unit 3091 derives the interpolated image Pred [][] from the temporary image temp[][] by vertical interpolation. The following Σ is the sum with respect to k for k=0..NTAP-1, and shift2 adjusts the range of values. The normalization parameter is offset2 = 1 << (shift2 - 1).

[0148] Pred[x][y] = (ΣmcFilter[yFrac][k]*temp[x][y+k-NTAP / 2+1]+offset2)>>shift2 In the case of biprediction, the above Pred[][] is derived for each L0 list and L1 list (referred to as interpolated images PredL0[][] and PredL1[][]), and an interpolated image Pred[][] is generated from the interpolated images PredL0[][] and PredL1[][].

[0149] (Weight prediction) The weight prediction unit 3094 generates a predicted block image by multiplying the motion-compensated image PredLX by a weight coefficient. If one of the prediction list usage flags (predFlagL0 or predFlagL1) is 1 (single prediction) and weight prediction is not used, the motion-compensated image PredLX (LX is L0 or L1) is adjusted to the number of pixels bit depth using the following process:

[0150] Pred[x][y] = Clip3(0,(1<<bitDepth)-1,(PredLX[x][y]+offset1)> >shift1) Here, shift1 = 14-bit depth and offset1 = 1 << (shift1 - 1). Furthermore, if both reference list usage flags (predFlagL0 and predFlagL1) are 1 (biprediction BiPred) and weight prediction is not used, the following process is performed to average the motion-compensated images PredL0 and PredL1 and adjust them to the number of pixels.

[0151] Pred[x][y] = Clip3(0,(1<<bitDepth)-1,(PredL0[x][y]+PredL1[x][y]+offset2)> >shift2) Here, shift2 = 15-bit depth and offset2 = 1 << (shift2 - 1).

[0152] Furthermore, when performing single prediction and weight prediction, the weight prediction unit 3094 derives the weight prediction coefficient w0 and offset o0 from the encoded data and performs the following processing.

[0153] Pred[x][y] = Clip3(0,(1<<bitDepth)-1,((PredLX[x][y]*w0+2^(log2WD-1))> >log2WD)+o0) Here, log2WD is a variable that indicates a predetermined shift amount.

[0154] Furthermore, when performing dual prediction BiPred and weight prediction, the weight prediction unit 3094 derives the weight prediction coefficients w0, w1, o0, and o1 from the encoded data and performs the following processing.

[0155] Pred[x][y] = Clip3(0,(1< <bitDepth)-1,(PredL0[x][y]*w0+PredL1[x][y]*w1+((o0+o1+1)<<log2WD))> >(log2WD+1)) The interpretation image generation unit 309 outputs the predicted image of the generated block to the addition unit 312.

[0156] (Intra predictive image generation unit 310) When the prediction mode predMode indicates intra-prediction mode, the intra-prediction image generation unit 310 performs intra-prediction using the intra-prediction parameters input from the intra-prediction parameter decoding unit 304 and the reference pixels read from the reference picture memory 306.

[0157] Specifically, the intra-predictive image generation unit 310 reads adjacent blocks on the target picture that are within a predetermined range from the target block from the reference picture memory 306. The predetermined range consists of adjacent blocks to the left, upper left, top, and upper right of the target block, and the area referenced differs depending on the intra-predictive mode.

[0158] The intra-predictive image generation unit 310 generates a predicted image of the target block by referring to the read decoded pixel value and the prediction mode indicated by the intra-predictive mode IntraPredMode. The intra-predictive image generation unit 310 outputs the generated predicted image to the summing unit 312.

[0159] The inverse quantization / inverse transformation unit 311 inversely quantizes the quantization transformation coefficients input from the entropy decoding unit 301 to obtain the transformation coefficients. These quantization transformation coefficients are obtained by quantizing the prediction error in the encoding process by applying frequency transformations such as DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), and KLT (Karyhnen Loeve Transform). The inverse quantization / inverse transformation unit 311 performs inverse frequency transformations such as inverse DCT, inverse DST, and inverse KLT on the obtained transformation coefficients to calculate the prediction error. The inverse quantization / inverse transformation unit 311 outputs the prediction error to the summing unit 312.

[0160] The summing unit 312 adds the predicted image of the block input from the prediction image generation unit 308 and the prediction error input from the inverse quantization / inverse transform unit 311 pixel by pixel to generate a decoded image of the block. The summing unit 312 stores the decoded image of the block in the reference picture memory 306 and also outputs it to the loop filter 305.

[0161] (Configuration of the video encoding device) Next, the configuration of the video encoding device 11 according to this embodiment will be described. Figure 27 is a block diagram showing the configuration of the video encoding device 11 according to this embodiment. The video encoding device 11 includes a prediction image generation unit 101, a subtraction unit 102, a transformation / quantization unit 103, an inverse quantization / inverse transformation unit 105, an addition unit 106, a loop filter 107, a prediction parameter memory (prediction parameter storage unit, frame memory) 108, a reference picture memory (reference image storage unit, frame memory) 109, an encoding parameter determination unit 110, a parameter encoding unit 111, and an entropy encoding unit 104.

[0162] The predictive image generation unit 101 generates a predictive image for each CU, which is a region obtained by dividing each picture of image T. The predictive image generation unit 101 operates in the same way as the predictive image generation unit 308 already described, so its explanation is omitted.

[0163] The subtraction unit 102 subtracts the pixel values ​​of the predicted image of the block input from the predicted image generation unit 101 from the pixel values ​​of image T to generate a prediction error. The subtraction unit 102 outputs the prediction error to the conversion / quantization unit 103.

[0164] The conversion / quantization unit 103 calculates conversion coefficients by frequency conversion for the prediction error input from the subtraction unit 102, and derives quantized conversion coefficients by quantization. The conversion / quantization unit 103 outputs the quantized conversion coefficients to the entropy coding unit 104 and the inverse quantization / inverse conversion unit 105.

[0165] The inverse quantization / inverse transformation unit 105 is the same as the inverse quantization / inverse transformation unit 311 (Figure 26) in the video decoding device 31, and therefore its explanation is omitted. The calculated prediction error is output to the summing unit 106.

[0166] The parameter coding unit 111 consists of a partial image region control unit 120, an inter-predictive parameter coding unit 112 (not shown), and an intra-predictive parameter coding unit 113.

[0167] The partial image region control unit 120 includes a header coding unit 1110, a CT information coding unit 1111, a CU coding unit 1112 (prediction mode coding unit), and an inter-prediction parameter coding unit 112 and an intra-prediction parameter coding unit 113 (not shown). The CU coding unit 1112 further includes a TU coding unit 1114.

[0168] The following describes the general operation of each module. The parameter coding unit 111 performs encoding processing on parameters such as header information, partitioning information, prediction information, and quantization conversion coefficients.

[0169] The CT information encoding unit 1111 encodes QT, MT (BT, TT) division information, etc., from the encoded data.

[0170] The CU encoding unit 1112 encodes CU information, prediction information, TU splitting flag split_transform_flag, CU residual flags cbf_cb, cbf_cr, cbf_luma, etc.

[0171] The TU encoding unit 1114 encodes the QP update information (quantization correction value) and the quantization prediction error (residual_coding) when the TU contains a prediction error.

[0172] The entropy coding unit 104 converts the syntax elements supplied from the source into binary data, generates coded data using an entropy coding scheme such as CABAC, and outputs it. The sources of the syntax elements are the CT information coding unit 1111 and the CU coding unit 1112. The syntax elements include inter-prediction parameters (prediction mode predMode, merge flag merge_flag, merge index merge_idx, inter-prediction identifier inter_pred_idc, reference picture index refIdxLX, prediction vector index mvp_LX_idx, difference vector mvdLX), intra-prediction parameters (prev_intra_luma_pred_flag, mpm_idx, rem_selected_mode_flag, rem_selected_mode, rem_non_selected_mode), quantization conversion coefficients, etc.

[0173] The entropy coding unit 104 entropy codes the division information, prediction parameters, quantization conversion coefficients, etc., to generate and output an encoded stream Te.

[0174] (Configuration of the interpredictive parameter coding unit) The interprediction parameter coding unit 112 derives interprediction parameters based on the prediction parameters input from the coding parameter determination unit 110. The interprediction parameter coding unit 112 includes a configuration that is partially identical to the configuration in which the interprediction parameter decoding unit 303 derives interprediction parameters.

[0175] (Configuration of the intra-predictive parameter coding unit 113) The intra-prediction parameter coding unit 113 derives a format for encoding (e.g., mpm_idx, rem_intra_luma_pred_mode, etc.) from the intra-prediction mode IntraPredMode input from the coding parameter determination unit 110. The intra-prediction parameter coding unit 113 includes some configurations identical to those of the intra-prediction parameter decoding unit 304, which derives the intra-prediction parameters.

[0176] The addition unit 106 generates a decoded image by adding the pixel values ​​of the predicted image of the block input from the predicted image generation unit 101 and the prediction error input from the inverse quantization / inverse transform unit 105 for each pixel. The addition unit 106 stores the generated decoded image in the reference picture memory 109.

[0177] The loop filter 107 applies a deblocking filter, SAO, and ALF to the decoded image generated by the summing unit 106. Note that the loop filter 107 does not necessarily have to include the above three types of filters; for example, it may consist of only a deblocking filter.

[0178] The prediction parameter memory 108 stores the prediction parameters generated by the coding parameter determination unit 110 in predetermined locations for each target picture and CU.

[0179] The reference picture memory 109 stores the decoded images generated by the loop filter 107 at predetermined locations for each target picture and CU.

[0180] The coding parameter determination unit 110 selects one set from among several sets of coding parameters. The coding parameters are the QT, BT, or TT segmentation information, prediction parameters, or parameters that are to be coded and generated in relation to these. The prediction image generation unit 101 generates a prediction image using these coding parameters.

[0181] The coding parameter determination unit 110 calculates an RD cost value for each of the multiple sets, which indicates the magnitude of the information and the coding error. The RD cost value is, for example, the sum of the code amount and the squared error multiplied by a coefficient λ. The code amount is the information amount of the coded stream Te obtained by entropy coding the quantization error and coding parameters. The squared error is the sum of the squares of the prediction errors calculated in the subtraction unit 102. The coefficient λ is a real number greater than zero that is set in advance. The coding parameter determination unit 110 selects the set of coding parameters that minimizes the calculated cost value. As a result, the entropy coding unit 104 outputs the selected set of coding parameters as the coded stream Te. The coding parameter determination unit 110 stores the determined coding parameters in the prediction parameter memory 108.

[0182] Furthermore, some parts of the video encoding device 11 and video decoding device 31 in the above-described embodiment, such as the entropy decoding unit 301, parameter decoding unit 302, loop filter 305, prediction image generation unit 308, inverse quantization / inverse transformation unit 311, addition unit 312, prediction image generation unit 101, subtraction unit 102, transformation / quantization unit 103, entropy encoding unit 104, inverse quantization / inverse transformation unit 105, loop filter 107, encoding parameter determination unit 110, and parameter encoding unit 111, may be implemented using a computer. In that case, a program for realizing this control function may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be loaded into a computer system and executed. Hereinafter, "computer system" refers to a computer system built into either the video encoding device 11 or the video decoding device 31, and includes hardware such as an OS and peripheral devices. Furthermore, "computer-readable recording media" refers to portable media such as flexible disks, magneto-optical disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into computer systems. In addition, "computer-readable recording media" may also include those that dynamically hold programs for a short period of time, such as communication lines used when transmitting programs over networks such as the Internet or communication lines such as telephone lines, and those that hold programs for a certain period of time, such as volatile memory inside computer systems that act as servers or clients in such cases. Moreover, the above-mentioned programs may be for the purpose of realizing some of the functions described above, and may also be programs that can realize the aforementioned functions in combination with programs already recorded in the computer system.

[0183] A motion image decoding device according to one aspect of the present invention comprises a picture division unit that divides a picture into partial image regions and non-partial image regions, with one of the CTU, CTU column, and CTU row as the smallest unit, and a predictive image generation unit that generates a predictive image. The above-mentioned predictive image generation unit uses intra-prediction and loop filtering, which references only the decoded pixels of the partial image region in the picture, or inter-prediction, which references the partial image region of the reference picture of the picture, for blocks included in the partial image region, and uses intra-prediction and loop filtering, which references the decoded pixels in the picture, or inter-prediction, which references the reference picture of the picture, for blocks included in the non-partial image region. The video decoding device, after decoding the picture, sets the above-mentioned partial image region of the picture as the partial image region of the reference picture.

[0184] A motion image decoding device according to one aspect of the present invention comprises a picture division unit that divides a picture into partial image regions and non-partial image regions, with one of the CTU, CTU column, and CTU row as the smallest unit, and a prediction image generation unit that generates a prediction image, wherein the prediction image generation unit refers to information indicating whether the picture is randomly accessible or not, and if random access is possible, it uses intra-prediction and loop filtering processing that refers only to the decoded pixels of the partial image region in the picture for the blocks included in the partial image region, and if random access is not possible, it uses intra-prediction and loop filtering processing that refers only to the decoded pixels of the partial image region in the picture, or inter-prediction that refers to the partial image region of the reference picture for the blocks included in the partial image region, and regardless of whether random access is possible or not, it uses intra-prediction and loop filtering processing that refers to the decoded pixels in the picture, or inter-prediction that refers to the reference picture of the picture, and after decoding the picture, the motion image decoding device sets the partial image region of the picture as the partial image region of the reference picture.

[0185] A motion image decoding device according to one aspect of the present invention is characterized in that the picture division unit divides the picture into a partial image region and a non-partial image region by referring to region information decoded from encoded data.

[0186] A motion image decoding device according to one aspect of the present invention is characterized in that the region information includes information indicating the position and size of the partial image region.

[0187] A motion image decoding device according to one aspect of the present invention is characterized by decoding refresh information indicating the number of pictures from a picture containing randomly accessible information until the entire picture constitutes a partial image area.

[0188] A motion image encoding device according to one aspect of the present invention comprises a picture division unit that divides a picture into partial image regions and non-partial image regions using one of CTU, CTU columns, and CTU rows as the smallest unit, and a prediction image generation unit that generates a prediction image, wherein the prediction image generation unit sets the partial image region of the picture as the restriction reference region after encoding the picture, by using intra-prediction and loop filtering processing that references only the decoded pixels of the partial image region in the picture for blocks included in the partial image region, or by using inter-prediction that references the restricted reference region of the reference picture for blocks included in the non-partial image region, and by using intra-prediction and loop filtering processing that references the decoded pixels in the picture, or by using inter-prediction that references the reference picture of the picture.

[0189] Furthermore, some or all of the video encoding device 11 and video decoding device 31 in the above-described embodiment may be implemented as an integrated circuit such as an LSI (Large Scale Integration). Each functional block of the video encoding device 11 and video decoding device 31 may be individually implemented as a processor, or some or all of them may be integrated into a single processor. In addition, the method of implementing the integrated circuit is not limited to LSIs; it may also be implemented using dedicated circuits or general-purpose processors. Furthermore, if an integrated circuit technology that can replace LSIs emerges due to advances in semiconductor technology, an integrated circuit using that technology may be used.

[0190] Although one embodiment of this invention has been described in detail above with reference to the drawings, the specific configuration is not limited to that described above, and various design changes can be made without departing from the spirit of this invention.

[0191] [Application Examples] The video encoding device 11 and video decoding device 31 described above can be installed and used in various devices that transmit, receive, record, and play back video. The video may be natural video captured by a camera or the like, or it may be artificial video (including CG and GUI) generated by a computer or the like.

[0192] First, with reference to Figure 16, we will explain how the aforementioned video encoding device 11 and video decoding device 31 can be used for transmitting and receiving video.

[0193] Figure 16(a) is a block diagram showing the configuration of the transmitter PROD_A equipped with the video encoding device 11. As shown in Figure 16(a), the transmitter PROD_A includes an encoding unit PROD_A1 that obtains encoded data by encoding video, a modulation unit PROD_A2 that obtains a modulated signal by modulating a carrier wave with the encoded data obtained by the encoding unit PROD_A1, and a transmitter PROD_A3 that transmits the modulated signal obtained by the modulation unit PROD_A2. The video encoding device 11 described above is used as this encoding unit PROD_A1.

[0194] The transmitting device PROD_A may further include a camera PROD_A4 for capturing moving images, a recording medium PROD_A5 for recording the moving images, an input terminal PROD_A6 for receiving moving images from an external source, and an image processing unit A7 for generating or processing images, as sources for supplying moving images to the encoding unit PROD_A1. Figure 16(a) illustrates a configuration in which the transmitting device PROD_A includes all of these components, but some may be omitted.

[0195] The recording medium PROD_A5 may contain unencoded video footage, or it may contain video footage encoded using a recording encoding method different from the transmission encoding method. In the latter case, a decoding unit (not shown) may be interposed between the recording medium PROD_A5 and the encoding unit PROD_A1 to decode the encoded data read from the recording medium PROD_A5 according to the recording encoding method.

[0196] Figure 16(b) is a block diagram showing the configuration of the receiver PROD_B equipped with the video decoding device 31. As shown in Figure 16(b), the receiver PROD_B includes a receiver PROD_B1 that receives a modulated signal, a demodulation unit PROD_B2 that obtains encoded data by demodulating the modulated signal received by the receiver PROD_B1, and a decoding unit PROD_B3 that obtains a video by decoding the encoded data obtained by the demodulation unit PROD_B2. The video decoding device 31 described above is used as this decoding unit PROD_B3.

[0197] The receiving device PROD_B may further include a display PROD_B4 for displaying the video, a recording medium PROD_B5 for recording the video, and an output terminal PROD_B6 for outputting the video to the outside, as recipients of the video output from the decoding unit PROD_B3. Figure 16(b) illustrates a configuration in which the receiving device PROD_B includes all of these, but some may be omitted.

[0198] The recording medium PROD_B5 may be for recording unencoded video, or it may be encoded using a recording encoding method different from the transmission encoding method. In the latter case, it is preferable to interpose an encoding unit (not shown) between the decoding unit PROD_B3 and the recording medium PROD_B5, which encodes the video acquired from the decoding unit PROD_B3 according to the recording encoding method.

[0199] The transmission medium for transmitting the modulated signal may be wireless or wired. Furthermore, the transmission method for transmitting the modulated signal may be broadcasting (referring here to a transmission method where the destination is not predetermined) or communication (referring here to a transmission method where the destination is predetermined). In other words, the transmission of the modulated signal may be achieved by wireless broadcasting, wired broadcasting, wireless communication, or wired communication.

[0200] For example, a terrestrial digital broadcasting station (broadcasting equipment, etc.) / receiving station (television receiver, etc.) is an example of a transmitting device PROD_A / receiving device PROD_B that transmits and receives modulated signals wirelessly. Similarly, a cable television broadcasting station (broadcasting equipment, etc.) / receiving station (television receiver, etc.) is an example of a transmitting device PROD_A / receiving device PROD_B that transmits and receives modulated signals via wired broadcasting.

[0201] Furthermore, servers (such as workstations) and clients (such as television sets, personal computers, and smartphones) for internet-based VOD (Video On Demand) services and video sharing services are examples of transmitting devices PROD_A and receiving devices PROD_B that transmit and receive modulated signals via communication (typically, either wireless or wired transmission is used as the transmission medium in a LAN, and wired transmission is used in a WAN). Here, personal computers include desktop PCs, laptop PCs, and tablet PCs. Smartphones also include multi-function mobile phones.

[0202] Furthermore, the video sharing service client has the function of decrypting encoded data downloaded from the server and displaying it on the screen, as well as the function of encoding video images captured by the camera and uploading them to the server. In other words, the video sharing service client functions as both a transmitting device PROD_A and a receiving device PROD_B.

[0203] Next, with reference to Figure 17, we will explain how the aforementioned video encoding device 11 and video decoding device 31 can be used for recording and playing back video.

[0204] Figure 17(a) is a block diagram showing the configuration of the recording device PROD_C equipped with the video encoding device 11 described above. As shown in the figure, the recording device PROD_C includes an encoding unit PROD_C1 that obtains encoded data by encoding video, and a writing unit PROD_C2 that writes the encoded data obtained by the encoding unit PROD_C1 to the recording medium PROD_M. The video encoding device 11 described above is used as this encoding unit PROD_C1.

[0205] The recording medium PROD_M may be (1) a type built into the recording device PROD_C, such as an HDD (Hard Disk Drive) or SSD (Solid State Drive); (2) a type connected to the recording device PROD_C, such as an SD memory card or USB (Universal Serial Bus) flash memory; or (3) a type loaded into a drive device (not shown) built into the recording device PROD_C, such as a DVD (Digital Versatile Disc: registered trademark) or BD (Blu-ray(registered trademark) Disc: registered trademark).

[0206] Furthermore, the recording device PROD_C may also include a camera PROD_C3 for capturing moving images, an input terminal PROD_C4 for receiving moving images from an external source, a receiving unit PROD_C5 for receiving moving images, and an image processing unit PROD_C6 for generating or processing images, as sources of moving images to be input to the encoding unit PROD_C1. In the figure, a configuration in which the recording device PROD_C includes all of these is shown as an example, but some may be omitted.

[0207] The receiving unit PROD_C5 may receive unencoded video footage, or it may receive encoded data encoded using a transmission encoding scheme different from the recording encoding scheme. In the latter case, it is preferable to interpose a transmission decoding unit (not shown) between the receiving unit PROD_C5 and the encoding unit PROD_C1 to decode the encoded data encoded using the transmission encoding scheme.

[0208] Examples of such recording devices PROD_C include DVD recorders, BD recorders, and HDD (Hard Disk Drive) recorders (in this case, the input terminal PROD_C4 or the receiver PROD_C5 is the main source of moving images). Camcorders (in this case, the camera PROD_C3 is the main source of moving images), personal computers (in this case, the receiver PROD_C5 or the image processing unit C6 is the main source of moving images), and smartphones (in this case, the camera PROD_C3 or the receiver PROD_C5 is the main source of moving images) are also examples of such recording devices PROD_C.

[0209] Figure 17(b) is a block diagram showing the configuration of the playback device PROD_D equipped with the video decoding device 31 described above. As shown in the figure, the playback device PROD_D includes a reading unit PROD_D1 that reads encoded data written to the recording medium PROD_M, and a decoding unit PROD_D2 that obtains a video by decoding the encoded data read by the reading unit PROD_D1. The video decoding device 31 described above is used as this decoding unit PROD_D2.

[0210] The recording medium PROD_M may be (1) a type built into the playback device PROD_D, such as an HDD or SSD; (2) a type connected to the playback device PROD_D, such as an SD memory card or USB flash memory; or (3) a type loaded into a drive device (not shown) built into the playback device PROD_D, such as a DVD or BD.

[0211] Furthermore, the playback device PROD_D may also include a display PROD_D3 for displaying the video, an output terminal PROD_D4 for outputting the video externally, and a transmission unit PROD_D5 for transmitting the video, as recipients of the video output from the decoding unit PROD_D2. In the figure, a configuration in which the playback device PROD_D includes all of these is shown as an example, but some may be omitted.

[0212] The transmitting unit PROD_D5 may transmit unencoded video footage, or it may transmit encoded data encoded using a transmission encoding scheme different from the recording encoding scheme. In the latter case, it is preferable to interpose an encoding unit (not shown) between the decoding unit PROD_D2 and the transmitting unit PROD_D5 to encode the video footage using the transmission encoding scheme.

[0213] Examples of such playback devices PROD_D include DVD players, BD players, and HDD players (in this case, the output terminal PROD_D4 to which a television receiver is connected becomes the main destination for the video). Other examples of such playback devices PROD_D include television receivers (in this case, the display PROD_D3 becomes the main destination for the video), digital signage (also called electronic billboards or electronic display boards, etc., where the display PROD_D3 or the transmitter PROD_D5 becomes the main destination for the video), desktop PCs (in this case, the output terminal PROD_D4 or the transmitter PROD_D5 becomes the main destination for the video), laptop or tablet PCs (in this case, the display PROD_D3 or the transmitter PROD_D5 becomes the main destination for the video), and smartphones (in this case, the display PROD_D3 or the transmitter PROD_D5 becomes the main destination for the video).

[0214] (Hardware implementation and software implementation) Furthermore, each block of the video decoding device 31 and video encoding device 11 described above may be implemented in hardware by logic circuits formed on an integrated circuit (IC chip), or it may be implemented in software using a CPU (Central Processing Unit).

[0215] In the latter case, each of the above devices includes a CPU that executes instructions for the program that realizes each function, a ROM (Read Only Memory) that stores the program, a RAM (Random Access Memory) that loads the program, and a storage device (recording medium) such as memory that stores the program and various data. The objective of the embodiment of the present invention can also be achieved by supplying each of the above devices with a recording medium on which the program code (executable program, intermediate code program, source program) of the control program of each of the above devices, which is software that realizes the above functions, is recorded in a way that can be read by a computer, and the computer (or CPU or MPU) reads and executes the program code recorded on the recording medium.

[0216] Examples of recording media that can be used include tapes such as magnetic tapes and cassette tapes, magnetic disks such as floppy disks and hard disks, optical disks such as CD-ROMs (Compact Disc Read-Only Memory), MO disks (Magneto-Optical discs), MDs (Mini Discs), DVDs (Digital Versatile Discs), CD-Rs (CD Recordable), and Blu-ray Discs (Blu-ray Discs). Other recording media include IC cards (including memory cards) and optical cards, semiconductor memories such as mask ROMs, EPROMs (Erasable Programmable Read-Only Memory), EEPROMs (Electrically Erasable and Programmable Read-Only Memory), and flash ROMs, and logic circuits such as PLDs (Programmable logic devices) and FPGAs (Field Programmable Gate Arrays).

[0217] Furthermore, each of the above devices may be configured to be connectable to a communication network, and the program code may be supplied via the communication network. This communication network is not particularly limited, as long as it is capable of transmitting the program code. For example, the Internet, intranet, extranet, LAN (Local Area Network), ISDN (Integrated Services Digital Network), VAN (Value-Added Network), CATV (CommunityRAntenna television / Cable Television) communication network, Virtual Private Network, telephone line network, mobile communication network, satellite communication network, etc., can be used. Also, the transmission medium constituting this communication network is not limited to a specific configuration or type, as long as it is capable of transmitting the program code. For example, it can be used with wired connections such as IEEE (Institute of Electrical and Electronic Engineers) 1394, USB, power line carriers, cable TV lines, telephone lines, and ADSL (Asymmetric Digital Subscriber Line) lines, as well as wireless connections such as IrDA (Infrared Data Association), infrared (like remote controls), Bluetooth®, IEEE 802.11 wireless, HDR (High Data Rate), NFC (Near Field Communication), DLNA® (Digital Living Network Alliance), mobile phone networks, satellite lines, and terrestrial digital broadcasting networks. Furthermore, embodiments of the present invention can also be realized in the form of computer data signals embedded in a carrier wave, where the above program code is embodied through electronic transmission.

[0218] The embodiments of the present invention are not limited to those described above, and various modifications are possible within the scope of the claims. That is, embodiments obtained by combining technical means that have been appropriately modified within the scope of the claims are also included in the technical scope of the present invention. [Industrial applicability]

[0219] Embodiments of the present invention can be suitably applied to a video decoding device that decodes encoded data from image data, and a video encoding device that generates encoded data from image data. Furthermore, they can be suitably applied to the data structure of encoded data generated by the video encoding device and referenced by the video decoding device. (Cross-reference of related applications) This application claims priority to Japanese Patent Application No. 2018-160712, filed on 29 August 2018, and all of its contents are included herein by reference. [Explanation of symbols]

[0220] 31 Image Decoder 301 Entropy Decoder 302 Parameter Decoding Unit 3020 Header Decoding Section 303 Interpretation parameter decoding unit 304 Intra Prediction Parameter Decoding Unit 308 Predictive Image Generation Unit 309 Interpretation Image Generation Unit 310 Intra Predictive Image Generation Unit 311 Inverse Quantization / Inverse Transformation Section 312 Addition section 320 Partial Image Area Control Unit 11 Image encoding device 101 Predictive Image Generation Unit 102 Subtraction Unit 103 Conversion / Quantization Section 104 Entropy coding unit 105 Inverse Quantization / Inverse Transformation Section 107 Loop Filter 110 Encoding parameter determination unit 111 Parameter coding section 112 Interpretation Parameter Coding Unit 113 Intra Prediction Parameter Coding Unit 120 Partial Image Area Control Unit 1110 Header Encoding Section 1111 CT information encoder 1112 CU coding unit (predictive mode coding unit) 1114 TU encoding section

Claims

1. A method for decoding an encoded stream that can be decoded by stepwise refresh, Decode nal_unit_type from the encoded stream, The system identifies whether the nal_unit_type is a sequential decoded refresh (SDR) network abstraction layer (NAL) unit type. When identified as an SDR NAL UNIT type, the nal_unit_type is associated with the first picture of the slice and picture set, and the first picture is an SDR picture in which the entire sub-image region is intra-encoded. In order to decode at least a subset of the pictures in the set of pictures, decode the available flags indicating the use of stepwise refresh, If the available flag is true, decode the syntax element indicating the period from the first picture until the entire picture is correctly decoded. A method characterized by the following:

2. The method according to claim 1, characterized in that the syntax element indicates the number of pictures.

3. A method for signaling a stepwise refresh to decode a set of pictures, Identify the first picture in the set of pictures as being associated with a sequential refresh picture, Encode nal_unit_type as a Sequential Decode-Refresh (SDR) Network Abstraction Layer (NAL) Unit type. The nal_unit_type is associated with the slice and the first picture, and the first picture is an SDR picture in which the entire partial image region is intra-encoded. Encode an available flag indicating the use of stepwise refresh to decode at least a subset of the pictures in the set of pictures, If the available flag is true, encode a syntax element indicating the period of time from the first picture until the entire picture is correctly decoded. A method characterized by the following:

Citation Information

Patent Citations

  • A Method for Random Access and Gradual Image Update in Image Coding

    JP2005533444A

  • Region of interest and signaling of incremental decoding refresh in video coding

    JP2015534775A

  • Incremental decoding refresh with temporal scalability support in video coding

    JP2016509404A

  • Video encoding / decoding device, method, and program

    WO2014002385A1