Image encoding device, image decoding device, image encoding method, image decoding method

By configuring an image decoding device to identify the necessary information for decoding the VVC bitstream, the redundancy in the syntax is reduced, leading to a more efficient encoding process and decreased processing time.

JP7689564B2Active Publication Date: 2025-06-06CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023208812
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-12-11
Publication Date
2025-06-06
Estimated Expiration
2039-06-21

AI Technical Summary

Technical Problem

In the Versatile Video Coding (VVC) method, the syntax element num_entry_point_offset is redundant as the number of entry_point_offset_minus1 can be derived from other syntax elements, leading to unnecessary code in the bitstream.

Method used

An image decoding device is configured to decode an image by identifying the number of pieces of information for the starting position of the coded data of the block row based on specific flags and information, allowing for the reduction of redundant syntax in the bitstream.

Benefits of technology

The proposed solution reduces the amount of code in the bitstream by eliminating redundant syntax, thereby improving encoding efficiency and reducing processing time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007689564000001
    Figure 0007689564000001
  • Figure 0007689564000002
    Figure 0007689564000002
  • Figure 0007689564000003
    Figure 0007689564000003
Patent Text Reader

Abstract

To provide a technique to reduce the code amount of a bit stream by reducing redundant syntax.SOLUTION: In a state where the value of a first flag is 1 while a second flag being a flag decoded from a picture parameter set of a bit stream and related to the mode of a slice indicates that a mode in which the slice is a rectangle is used, and a target slice included in an image includes a plurality of rectangular areas in the horizontal direction or the vertical direction, the number of pieces of information for specifying the head position of code data of block rows is specified for the target slice. The code data of the block rows is decoded based at least on the specified number of pieces of information for specifying the head position, and the information for specifying the head position.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an image encoding / decoding technique. [Background technology]

[0002] The High Efficiency Video Coding (HEVC) coding method (hereafter referred to as HEVC) is known as a coding method for compressing and recording moving images. In order to improve coding efficiency, HEVC employs basic blocks larger than conventional macroblocks (16x16 pixels). These large basic blocks are called coding tree units (CTUs) and have a maximum size of 64x64 pixels. CTUs are further divided into subblocks, which are units for prediction and transformation.

[0003] In addition, HEVC allows a picture to be divided into multiple tiles or slices for encoding. There is little data dependency between tiles or slices, so encoding and decoding can be performed in parallel. One of the major advantages of dividing into tiles and slices is that it allows parallel processing on a multi-core CPU, etc., reducing processing time.

[0004] Moreover, each slice is coded by the conventional binary arithmetic coding method adopted in HEVC. That is, each syntax element is binarized to generate a binary signal. For each syntax element, the occurrence probability is given in advance as a table (hereinafter, occurrence probability table), and the binary signal is arithmetically coded based on the occurrence probability table. During decoding, this occurrence probability table is used as decoding information for decoding the subsequent code. During coding, it is used as coding information for the subsequent coding. Then, every time coding is performed, the occurrence probability table is updated based on statistical information on whether the coded binary signal was a symbol with a high occurrence probability or not.

[0005] HEVC also has a method for parallel processing of entropy coding and decoding called Wavefront Parallel Processing (hereinafter, WPP). In WPP, a table of occurrence probabilities at the time of coding processing of a block at a pre-specified position is applied to the leftmost block of the next row, thereby enabling parallel coding processing of blocks on a row-by-row basis while suppressing a decrease in coding efficiency. To enable parallel processing on a block-row basis, entry_point_offset_minus1 indicating the start position of each block row in the bitstream and num_entry_point_offsets indicating the number of block rows are coded in the slice header. Patent Document 1 discloses a technology related to WPP.

[0006] In recent years, activities have been started to internationally standardize a more efficient coding method as a successor to HEVC. JVET (Joint Video Experts Team) was established between ISO / IEC and ITU-T, and standardization is proceeding as the Versatile Video Coding (VVC) coding method (hereinafter referred to as VVC). In VVC, it is considered to further divide tiles into rectangles (bricks) consisting of multiple block rows. A slice is then configured to contain one or more bricks. [Prior art documents] [Patent documents]

[0007] [Patent Document 1] JP 2014-11638 A Summary of the Invention [Problem to be solved by the invention]

[0008] In VVC, the bricks that make up a slice can be derived in advance, and the number of basic block rows contained in the brick can also be derived from other syntax. Therefore, the number of entry_point_offset_minus1 indicating the start position of the basic block row belonging to the slice can be derived without using num_entry_point_offset. Therefore, num_entry_point_offset is a redundant syntax. In the present invention, a technology is provided that reduces the amount of code in the bitstream by reducing redundant syntax. [Means for solving the problem]

[0009] One aspect of the present invention is an image decoding device that decodes an image including a rectangular area including one or more block rows each including a plurality of blocks from a bit stream obtained by encoding the image, the image decoding device including: a decoding means for decoding, from the bit stream, a first flag related to enabling parallel processing; first information used to identify a rectangular area to be processed first among a plurality of rectangular areas included in a slice of the image; second information used to identify a rectangular area to be processed last among the plurality of rectangular areas; and third information corresponding to the number of blocks in a vertical direction of the rectangular area in the image; The present invention is characterized in that, in a state where a second flag decoded from the set and which is a flag related to the slice mode indicates that a mode in which the slice is rectangular is used, and a target slice included in the image contains multiple rectangular areas in the horizontal or vertical direction, the present invention further comprises an identification means for identifying the number of pieces of information for identifying the starting position of the coded data of the block row for the target slice based on the first information, the second information, and the third information, and the decoding means decodes the coded data of the block row based at least on the number of pieces of information for identifying the starting position identified by the identification means and the information for identifying the starting position. Effect of the Invention

[0010] According to the configuration of the present invention, it is possible to reduce the amount of code in a bit stream by reducing redundant syntax. [Brief description of the drawings]

[0011] [Figure 1] FIG. 1 is a block diagram showing an example of the functional configuration of an image encoding device. [Diagram 2] FIG. 2 is a block diagram showing an example of the functional configuration of an image decoding device. [Diagram 3] 4 is a flowchart of an input image encoding process performed by the image encoding device. [Figure 4] 13 is a flowchart of a bitstream decoding process performed by an image decoding device. [Diagram 5] FIG. 2 is a block diagram showing an example of the hardware configuration of a computer apparatus. [Figure 6] FIG. 1 is a diagram showing an example of a bitstream format. [Figure 7] FIG. 13 is a diagram showing an example of dividing an input image. [Figure 8] FIG. 13 is a diagram showing an example of dividing an input image. [Figure 9] FIG. 1 is a diagram showing the relationship between tiles and slices. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0012] Hereinafter, the embodiments will be described in detail with reference to the attached drawings. Note that the following embodiments do not limit the invention according to the claims. Although the embodiments describe a number of features, not all of these features are essential to the invention, and the features may be combined in any manner. Furthermore, in the attached drawings, the same reference numbers are used for the same or similar configurations, and duplicated descriptions are omitted.

[0013] [First embodiment] First, an example of the functional configuration of the image encoding device according to this embodiment will be described with reference to the block diagram of FIG. 1. An input image to be encoded is input to the image division unit 102. The input image may be an image of each frame constituting a moving image, or may be a still image. The image division unit 102 divides the input image into "one or more tiles." A tile is a set of continuous basic blocks covering a rectangular area in the input image. The image division unit 102 further divides each tile into one or more bricks. A brick is a rectangular area (a rectangular area including one or more block rows consisting of multiple blocks whose size is equal to or smaller than a tile) consisting of one or more rows of basic blocks (basic block rows) in a tile. The image division unit 102 further divides the input image into slices consisting of "one or more tiles" or "one or more bricks in one tile." A slice is a basic unit of encoding, and header information such as information indicating the type of slice is added to each slice. An example of dividing an input image into four tiles, four slices, and eleven bricks is shown in FIG. 7. The top left tile is divided into one brick, the bottom left tile into two bricks, the top right tile into five bricks, and the bottom right tile into three bricks. The left slice is configured to include three bricks, the top right slice into two bricks, the center right slice into three bricks, and the bottom right slice into three bricks. The image division unit 102 outputs information about the size of each of the tiles, bricks, and slices divided in this way as division information.

[0014] The block dividing unit 103 divides the image of the basic block row (basic block row image) output from the image dividing unit 102 into a plurality of basic blocks, and outputs the image (block image) in units of basic blocks to the subsequent stage.

[0015] The prediction unit 104 divides an image in units of basic blocks into sub-blocks, and performs intra-prediction, which is intra-frame prediction, and inter-prediction, which is inter-frame prediction, on a sub-block basis to generate a predicted image. Intra-prediction across bricks (intra-prediction using pixels of blocks of other bricks) and prediction of motion vectors across bricks (prediction of motion vectors using motion vectors of blocks of other bricks) are not performed. Furthermore, the prediction unit 104 calculates and outputs a prediction error from the input image and the predicted image. The prediction unit 104 also outputs information required for prediction (prediction information), such as a sub-block division method, a prediction mode, and motion vectors, together with the prediction error.

[0016] The transform / quantization unit 105 performs orthogonal transform on the prediction error in units of subblocks to obtain transform coefficients, and quantizes the obtained transform coefficients to obtain quantized coefficients. The inverse quantization / inverse transform unit 106 inverse quantizes the quantized coefficients output from the transform / quantization unit 105 to reproduce the transform coefficients, and further performs inverse orthogonal transform on the reproduced transform coefficients to reproduce the prediction error.

[0017] The frame memory 108 functions as a memory for storing the reconstructed image. The image reconstruction unit 107 generates a predicted image by appropriately referring to the frame memory 108 based on the prediction information output from the prediction unit 104, and generates and outputs a reconstructed image from the predicted image and the input prediction error.

[0018] The in-loop filter unit 109 performs in-loop filtering such as deblocking filtering and sample adaptive offset on the reconstructed image, and outputs an image that has been subjected to the in-loop filtering (filtered image).

[0019] The encoding unit 110 generates code data (encoded data) by encoding the quantization coefficients output from the transform / quantization unit 105 and the prediction information output from the prediction unit 104, and outputs the generated code data.

[0020] The integrated encoding unit 111 generates header code data using the division information output from the image division unit 102, and generates and outputs a bit stream including the generated header code data and the code data output from the encoding unit 110. The control unit 199 controls the operation of the entire image encoding device, and controls the operation of each functional unit of the image encoding device.

[0021] Next, a description will be given of the coding process for an input image by the image coding device having the configuration shown in Fig. 1. In this embodiment, for ease of explanation, only intra-prediction coding process will be described, but the present invention is not limited to this and can also be applied to inter-prediction coding process. Furthermore, in order to provide a specific explanation, this embodiment will be described assuming that the block division unit 103 divides the basic block row image output from the image division unit 102 into units of "basic blocks having a size of 64 x 64 pixels".

[0022] The image division unit 102 divides the input image into tiles and bricks. An example of division of an input image by the image division unit 102 is shown in Fig. 8. In this embodiment, as shown in Fig. 8(a), an input image having a size of 1152 x 1152 pixels is divided into nine tiles (each tile has a size of 384 x 384 pixels). An ID (tile ID) is assigned to each tile in raster order starting from the top left, with the tile ID of the top left tile being 0 and the tile ID of the bottom right tile being 8.

[0023] Also, an example of dividing an input image into tiles, bricks, and slices is shown in FIG. 8(b). As shown in FIG. 8(b), the tiles with tile ID=0 and tile ID=7 are divided into two bricks (each brick has a size of 384×192 pixels). The tile with tile ID=2 is divided into two bricks (the upper brick has a size of 384×128 pixels, and the lower brick has a size of 384×256 pixels). The tile with tile ID=3 is divided into three bricks (each brick has a size of 384×128 pixels). The tiles with tile ID=1, 4, 5, 6, and 8 are not divided into bricks (equivalent to dividing one tile into one brick), and as a result, tile=brick. An ID is assigned to each brick in order from top to bottom in the raster order of the tiles. The BID shown in FIG. 8(b) is the brick ID. The input image is also divided into slices including bricks corresponding to BID=0 to 2, a slice including a brick corresponding to BID=3, a slice including a brick corresponding to BID=4, a slice including bricks corresponding to BID=5 to 8 and 10 to 12, and a slice including bricks corresponding to BID=9 and 13. An ID is also assigned to each slice in order from top to bottom among the slices in raster order; for example, slice 0 refers to the slice with ID=0, and slice 4 refers to the slice with ID=4.

[0024] The image division unit 102 then outputs information about the size of each of the divided tiles, bricks, and slices as division information to the integrated encoding unit 111. The image division unit 102 also divides each brick into basic block row images, and outputs the divided basic block row images to the block division unit 103.

[0025] The block division unit 103 divides the basic block row image output from the image division unit 102 into a plurality of basic blocks, and outputs block images (64×64 pixels) which are images in units of basic blocks to a prediction unit 104 at a subsequent stage.

[0026] The prediction unit 104 divides an image in units of basic blocks into sub-blocks, determines an intra-prediction mode such as horizontal prediction or vertical prediction for each sub-block, and generates a predicted image from the determined intra-prediction mode and encoded pixels. Furthermore, the prediction unit 104 calculates a prediction error from the input image and the predicted image, and outputs the calculated prediction error to the transformation and quantization unit 105. The prediction unit 104 also outputs information such as the sub-block division method and the intra-prediction mode as prediction information to the encoding unit 110 and the image reproduction unit 107.

[0027] The transform / quantization unit 105 performs orthogonal transform (orthogonal transform processing corresponding to the size of the subblock) on the prediction error output from the prediction unit 104 in units of subblocks to obtain transform coefficients (orthogonal transform coefficients). The transform / quantization unit 105 then quantizes the obtained transform coefficients to obtain quantization coefficients. The transform / quantization unit 105 then outputs the obtained quantization coefficients to the coding unit 110 and the inverse quantization / inverse transform unit 106.

[0028] The inverse quantization and inverse transform unit 106 inverse quantizes the quantized coefficients output from the transform and quantization unit 105 to regenerate the transform coefficients, and further performs inverse orthogonal transform on the regenerated transform coefficients to regenerate the prediction errors. The inverse quantization and inverse transform unit 106 then outputs the regenerated prediction errors to the image reproduction unit 107.

[0029] The image reproduction unit 107 generates a predicted image by appropriately referring to the frame memory 108 based on the prediction information output from the prediction unit 104, and generates a reproduced image from the predicted image and the prediction error input from the inverse quantization and inverse transform unit 106. Then, the image reproduction unit 107 stores the generated reproduced image in the frame memory 108.

[0030] The in-loop filter unit 109 reads the reconstructed image from the frame memory 108, and performs in-loop filter processing such as deblocking filtering and sample adaptive offset on the read reconstructed image. Then, the in-loop filter unit 109 stores (restores) the image that has been subjected to the in-loop filter processing in the frame memory 108 again.

[0031] The coding unit 110 generates coded data by entropy coding the quantized coefficients output from the transform / quantization unit 105 and the prediction information output from the prediction unit 104. There is no particular specification as to the entropy coding method, but Golomb coding, arithmetic coding, Huffman coding, etc. can be used. The coding unit 110 then outputs the generated coded data to the integrated coding unit 111.

[0032] The integrated encoding unit 111 generates header code data by using the division information output from the image division unit 102, and generates and outputs a bit stream by multiplexing the generated header code data and the code data output from the encoding unit 110. The output destination of the bit stream is not limited to a specific output destination, and the bit stream may be output (stored) in an internal or external memory of the image encoding device, or may be transmitted to an external device that can communicate with the image encoding device via a network such as a LAN or the Internet.

[0033] Next, an example of the format of the bitstream (VVC-based coded data coded by the image coding device) output by the integrated coding unit 111 is shown in FIG. 6. The bitstream in FIG. 6 includes a sequence parameter set (SPS) which is header information including information related to coding of a sequence. The bitstream in FIG. 6 also includes a picture parameter set (PPS) which is header information including information related to coding of a picture. The bitstream in FIG. 6 also includes a slice header (SLH) which is header information including information related to coding of a slice. The bitstream in FIG. 6 also includes coded data of each brick (brick 0 to brick (N-1) in FIG. 6).

[0034] The SPS includes image size information and basic block data partition information. The PPS includes tile data partition information, which is partition information for tiles, brick data partition information, which is partition information for bricks, slice data partition information 0, which is partition information for slices, and basic block row data synchronization information. The SLH includes slice data partition information 1 and basic block row data position information.

[0035] First, the SPS will be described. The SPS includes pic_width_in_luma_samples, which is information 601, and pic_height_in_luma_samples, which is information 602, as image size information. pic_width_in_luma_samples represents the horizontal size (number of pixels) of the input image, and pic_height_in_luma_samples represents the vertical size (number of pixels) of the input image. In this embodiment, since the input image of FIG. 8 is used as the input image, pic_width_in_luma_samples=1152 and pic_height_in_luma_samples=1152. The SPS also includes log2_ctu_size_minus2, which is information 603, as basic block data division information. log2_ctu_size_minus2 represents the size of the basic block. The number of pixels in the vertical and horizontal directions of the basic block is expressed as 1<<(log2_ctu_size_minus2+2). In this embodiment, the size of a basic block is 64×64 pixels, so the value of log2_ctu_size_minus2 is 4.

[0036] Next, the PPS will be described. The PPS includes information 604 to 607 as tile data division information. Information 604 is single_tile_in_pic_flag, which indicates whether the input image is divided into multiple tiles and encoded. When single_tile_in_pic_flag=1, it indicates that the input image is not divided into multiple tiles and encoded. On the other hand, when single_tile_in_pic_flag=0, it indicates that the input image is divided into multiple tiles and encoded.

[0037] Information 605 is information that is included in the tile data division information when single_tile_in_pic_flag = 0. Information 605 is uniform_tile_spacing_flag that indicates whether each tile has the same size. When uniform_tile_spacing_flag = 1, it indicates that each tile has the same size, and when uniform_tile_spacing_flag = 0, it indicates that tiles of different sizes exist.

[0038] Information 606 and information 607 are information to be included in the tile data division information when uniform_tile_spacing_flag=1. Information 606 is tile_cols_width_minus1 indicating (number of basic blocks in the horizontal direction of the tile-1). Information 607 is tile_rows_height_minus1 indicating (number of basic blocks in the vertical direction of the tile-1). The number of tiles in the horizontal direction of the input image is obtained as a quotient when the number of basic blocks in the horizontal direction of the input image is divided by the number of basic blocks in the horizontal direction of the tile. If a remainder occurs as a result of this division, the number obtained by adding 1 to the quotient is set as the "number of tiles in the horizontal direction of the input image". The number of tiles in the vertical direction of the input image is obtained as a quotient when the number of basic blocks in the vertical direction of the input image is divided by the number of basic blocks in the vertical direction of the tile. If a remainder occurs as a result of this division, the number obtained by adding 1 to the quotient is set as the "number of tiles in the vertical direction of the input image". The total number of tiles in the input image can be calculated by multiplying the number of tiles in the horizontal direction of the input image by the number of tiles in the vertical direction of the input image.

[0039] Note that when uniform_tile_spacing_flag=0, tiles of different sizes are included, so the number of tiles in the horizontal direction of the input image, the number of tiles in the vertical direction of the input image, and the vertical and horizontal sizes of each tile are coded.

[0040] The PPS also includes information 608 to 613 as brick data division information. Information 608 is brick_splitting_present_flag. When brick_splitting_present_flag=1, it indicates that one or more tiles in the input image are divided into multiple bricks. On the other hand, when brick_splitting_present_flag=0, it indicates that each tile in the input image is composed of a single brick.

[0041] Information 609 is information included in the brick data division information when brick_splitting_present_flag=1. Information 609 is brick_split_flag[] indicating whether or not each tile is divided into multiple bricks. brick_split_flag[] indicating whether the i-th tile is divided into multiple bricks is represented as brick_split_flag[i]. When brick_split_flag[i]=1, it indicates that the i-th tile is divided into multiple bricks, and when brick_split_flag[i]=0, it indicates that the i-th tile is composed of a single brick.

[0042] Information 610 is uniform_brick_spacing_flag[i] indicating whether the size of each brick constituting the i-th tile is the same when brick_split_flag[i]=1. When brick_split_flag[i]=0 for all i, information 610 is not included in the brick data division information. Information 610 includes uniform_brick_spacing_flag[i] for i that satisfies brick_split_flag[i]=1. When uniform_brick_spacing_flag[i]=1, it indicates that the size of each brick constituting the i-th tile is the same. On the other hand, when uniform_brick_spacing_flag[i]=0, it indicates that there is a brick that is different in size from the others among the bricks constituting the i-th tile.

[0043] Information 611 is information that is included in the brick data division information when uniform_brick_spacing_flag[i] = 1. The information 611 is brick_height_minus1[i] that indicates (the number of basic blocks in the vertical direction of the brick in the i-th tile-1).

[0044] The number of basic blocks in the vertical direction of a brick can be obtained by dividing the number of pixels in the vertical direction of the brick by the number of pixels in the vertical direction of the basic block (64 pixels in this embodiment). The number of bricks constituting a tile is obtained as the quotient when the number of basic blocks in the vertical direction of the tile is divided by the number of basic blocks in the vertical direction of the brick. If a remainder occurs as a result of this division, the quotient plus 1 is set as the "number of bricks constituting a tile". For example, suppose that the number of basic blocks in the vertical direction of a tile is 10, and the value of brick_height_minus1 is 2. In this case, the tile is divided into four bricks, starting from the top: a brick with 3 basic block rows, a brick with 3 basic block rows, a brick with 3 basic block rows, and a brick with 1 basic block row.

[0045] The information 612 is num_brick_rows_minus1[i] indicating (the number of bricks constituting the i-th tile-1) for i that satisfies uniform_brick_spacing_flag[i]=0.

[0046] In this embodiment, when uniform_brick_spacing_flag[i]=0, num_brick_rows_minus1[i] indicating (the number of bricks constituting the i-th tile-1) is included in the brick data division information. However, this is not limited to this.

[0047] For example, assume that the number of bricks constituting the i-th tile is 2 or more at the time of brick_split_flag[i]=1. Then, num_brick_rows_minus2[i] indicating (the number of bricks constituting the tile-2) may be encoded instead of num_brick_rows_minus1[i]. In this way, the number of bits of the syntax indicating the number of bricks constituting the tile can be reduced. For example, when the tile is composed of two bricks and num_brick_rows_minus1[i] is Golomb-encoded, 3-bit data of "010" indicating "1" is encoded. On the other hand, when num_brick_rows_minus2[i] indicating (the number of bricks constituting the tile-2) is Golomb-encoded, 1-bit data of "0" indicating 0 is encoded.

[0048] Information 613 is brick_row_height_minus1[i][j] indicating (the number of basic blocks in the vertical direction of the jth brick in the ith tile-1) for i that satisfies uniform_brick_spacing_flag[i]=0. The number of brick_row_height_minus1[i][j] is coded as many times as num_brick_rows_minus1[i]. When the above-mentioned num_brick_rows_minus2[i] is used, the number of brick_row_height_minus1[i][j] is coded as many times as num_brick_rows_minus2[i]+1. The number of basic blocks in the vertical direction of the brick at the bottom edge in the tile can be found by subtracting the sum of "brick_row_height_minus1+1" from the number of basic blocks in the vertical direction of the tile. For example, assume that the number of basic blocks in the vertical direction of a tile = 10, num_brick_rows_minus1 = 3, and brick_row_height_minus1 = 2, 1, 2. In this case, the number of basic blocks in the vertical direction of the bottom edge brick in that tile is 10 - (3 + 2 + 3) = 2.

[0049] Furthermore, the PPS includes information 614 to 618 as slice data division information 0. Information 614 is single_brick_per_slice_flag. When single_brick_per_slice_flag=1, it indicates that all slices in the input image are composed of a single brick. On the other hand, when single_brick_per_slice_flag=0, it indicates that one or more slices in the input image are composed of multiple bricks. In other words, it indicates that each slice is composed of only one brick.

[0050] Information 615 is rect_slice_flag, which is information included in slice data division information 0 when single_brick_per_slice_flag=0. rect_slice_flag indicates whether the tiles included in the slice are in raster order or rectangular. FIG. 9(a) shows the relationship between tiles and slices when rect_slice_flag=0, indicating that tiles in the slice are coded in raster order. On the other hand, FIG. 9(b) shows the relationship between tiles and slices when rect_slice_flag=1, indicating that multiple tiles in the slice are rectangular.

[0051] Information 616 is num_slices_in_pic_minus1, which is information included in slice data division information 0 when rect_slice_flag = 1 and single_brick_per_slice_flag = 0. num_slices_in_pic_minus1 indicates (the number of slices in the input image - 1).

[0052] The information 617 is top_left_brick_idx[i] which indicates the index of the top left brick of each slice in the input image (the i-th slice).

[0053] Information 618 is bottom_right_brick_idx_delta[i] indicating the difference between the index of the top left brick and the index of the bottom right brick in the i-th slice in the input image. Here, the "top left brick in the i-th slice in the input image" is the first brick to be processed in the slice. Also, the "bottom right brick in the i-th slice in the input image" is the last brick to be processed in the slice.

[0054] Here, the range of i is 0 to num_slices_in_pic_minus1. However, since the index of the top left brick of the first slice in a frame is determined to be 0, top_left_brick_idx[0] of the first slice is not coded. In this embodiment, the range of i of top_left_brick_idx[i] and bottom_right_brick_idx_delta[i] is set to 0 to num_slices_in_pic_minus1, but is not limited to this. For example, the bottom right brick of the last slice (the num_slices_in_pic_minus1-th slice) is determined to be the brick with the largest BID. Therefore, bottom_right_brick_idx_delta[num_slices_in_pic_minus1] does not need to be coded. Furthermore, the bricks included in slices other than the last slice are already identified by the top_left_brick_idx[i] and bottom_right_brick_idx_delta[i] of slices other than the last slice. Therefore, the bricks contained in the last slice can be identified as all the bricks that are not contained in the previous slices. In this case, the top left brick of the last slice can be identified as the brick with the smallest BID among the remaining bricks. Therefore, top_left_brick_idx[num_slices_in_pic_minus1] does not need to be coded either. This allows further reduction in the number of bits required for the header.

[0055] The PPS also contains coded information 619 as basic block row data synchronization information. The information 619 is entropy_coding_sync_enabled_flag. When entropy_coding_sync_enabled_flag=1, a table of occurrence probabilities at the time of processing a basic block at a specific position in the adjacent basic block row above is applied to the leftmost block. This enables parallel processing of entropy encoding and decoding on a basic block row basis.

[0056] Next, the SLH will be described. The SLH contains coded information 620-621 as slice data division information 1. Information 620 is slice_address that is included in slice data division information 1 when rect_slice_flag=1 or the number of bricks in the input image is two or more. When rect_slice_flag=0, slice_address indicates the BID at the beginning of the slice, and when rect_slice_flag=1, it indicates the number of the current slice.

[0057] Information 621 is num_bricks_in_slice_minus1 that is included in slice data division information 1 when rect_slice_flag = 0 and single_brick_per_slice_flag = 0. num_bricks_in_slice_minus1 indicates (the number of bricks in a slice - 1).

[0058] The SLH includes information 622 as basic block row data position information. The information 622 is entry_point_offset_minus1[]. When entropy_coding_sync_enabled_flag=1, entry_point_offset_minus1[] is coded and included in the basic block row data position information by the number of (the number of basic block rows in the slice-1).

[0059] entry_point_offset_minus1[] indicates the entry point of the coded data of the basic block row, i.e., the start position of the coded data of the basic block row. entry_point_offset_minus1[j-1] indicates the entry point of the coded data of the jth basic block row. The start position of the coded data of the 0th basic block row is omitted because it is the same as the start position of the coded data of the slice to which the basic block row belongs. Then, {the size of the coded data of the (j-1)th basic block row - 1} is encoded as entry_point_offset_minus1[j-1].

[0060] Next, the encoding process of an input image by the image encoding device of this embodiment (the generation process of a bit stream having the configuration in FIG. 6) will be described with reference to the flowchart in FIG.

[0061] First, in step S301, the image division unit 102 divides an input image into tiles, bricks, and slices. Then, the image division unit 102 outputs information about the size of each of the divided tiles, bricks, and slices as division information to the integrated encoding unit 111. The image division unit 102 also divides each brick into basic block row images, and outputs the divided basic block row images to the block division unit 103.

[0062] In step S302, the block division unit 103 divides the basic block row image into a plurality of basic blocks, and outputs block images, which are images in units of basic blocks, to the prediction unit 104 at the subsequent stage.

[0063] In step S303, the prediction unit 104 divides the basic block-based image output from the block division unit 103 into sub-blocks, determines an intra-prediction mode for each sub-block, and generates a predicted image from the determined intra-prediction mode and encoded pixels. Furthermore, the prediction unit 104 calculates a prediction error from the input image and predicted image, and outputs the calculated prediction error to the transformation and quantization unit 105. The prediction unit 104 also outputs information such as the sub-block division method and the intra-prediction mode as prediction information to the encoding unit 110 and the image reproduction unit 107.

[0064] In step S304, the transform / quantization unit 105 performs orthogonal transform on the prediction error output from the prediction unit 104 in units of sub-blocks to obtain transform coefficients (orthogonal transform coefficients). The transform / quantization unit 105 then quantizes the obtained transform coefficients to obtain quantization coefficients. The transform / quantization unit 105 then outputs the obtained quantization coefficients to the coding unit 110 and the inverse quantization / inverse transform unit 106.

[0065] In step S305, the inverse quantization / inverse transform unit 106 inverse quantizes the quantized coefficients output from the transform / quantization unit 105 to regenerate the transform coefficients, and then performs inverse orthogonal transform on the regenerated transform coefficients to regenerate the prediction errors. The inverse quantization / inverse transform unit 106 then outputs the regenerated prediction errors to the image reproduction unit 107.

[0066] In step S306, the image reproduction unit 107 generates a predicted image by appropriately referring to the frame memory 108 based on the prediction information output from the prediction unit 104, and generates a reproduced image from the predicted image and the prediction error input from the inverse quantization / inverse transform unit 106. Then, the image reproduction unit 107 stores the generated reproduced image in the frame memory 108.

[0067] In step S307, the encoding unit 110 performs entropy encoding on the quantized coefficients output from the transform / quantization unit 105 and the prediction information output from the prediction unit 104 to generate coded data.

[0068] Here, when entropy_coding_sync_enabled_flag=1, the occurrence probability table at the time when the basic block at a predetermined position in the adjacent basic block row above is processed is applied before processing the leftmost basic block in the next basic block row. In this embodiment, the description will be given assuming that entropy_coding_sync_enabled_flag=1.

[0069] In step S308, the control unit 199 determines whether or not the coding of all basic blocks in the slice has been completed. If the result of this determination is that the coding of all basic blocks in the slice has been completed, the process proceeds to step S309. On the other hand, if there are basic blocks in the slice that have not yet been coded (uncoded basic blocks), the process proceeds to step S303 to code the uncoded basic blocks.

[0070] In step S309, the integrated encoding unit 111 generates header code data using the division information output from the image division unit 102, and generates and outputs a bit stream including the generated header code data and the code data output from the encoding unit 110.

[0071] 8, in the tile data division information, single_tile_in_pic_flag is 0, and uniform_tile_spacing_flag is 1. In addition, tile_cols_width_minus1 is 5, and tile_rows_height_minus1 is 5.

[0072] The brick_splitting_present_flag in the brick data splitting information is 1. Also, the tiles corresponding to tile IDs = 1, 4, 5, 6, and 8 are not split into bricks. However, brick_split_flag[1], brick_split_flag[4], brick_split_flag[5], brick_split_flag[6], and brick_split_flag[8] are 0. Also, the tiles corresponding to tile IDs = 0, 2, 3, and 7 are split into bricks. However, brick_split_flag[0], brick_split_flag[2], brick_split_flag[3], and brick_split_flag[7] are 1.

[0073] Additionally, the tiles corresponding to tile ID=0, 3, and 7 are all divided into bricks of the same size. However, uniform_brick_spacing_flag[0], uniform_brick_spacing_flag[3], and uniform_brick_spacing_flag[7] are all set to 1. For the tile corresponding to tile ID=2, the size of the brick with BID=3 is different from the size of the brick with BID=4. However, uniform_brick_spacing_flag[2] is set to 0.

[0074] brick_height_minus1[0] is 2, brick_height_minus1[3] is 1, and brick_height_minus1[7] is 2. Note that brick_height_minus1 is encoded when uniform_brick_spacing is 1.

[0075] brick_row_height_minus1[2][0] is 1. Note that if you code the above syntax for num_brick_rows_minus2[2] instead of num_brick_rows_minus1[2], the value will be 0.

[0076] In this embodiment, since one slice includes multiple bricks, single_brick_per_slice_flag in slice data division information 0 is 0. In this embodiment, since a slice contains multiple tiles in a rectangle, rect_slice_flag is 1. As shown in FIG. 8B, since the number of slices in the input image is 5, num_slices_in_pic_minus1 is 4.

[0077] For slice 0, top_left_brick_idx[0] is not coded because 0 is trivial, and bottom_right_brick_idx_delta[0] is 2 (= 2-0). For slice 1, top_left_brick_idx[1] is 3, and bottom_right_brick_idx_delta[1] is 0 (= 3-3). For slice 2, top_left_brick_idx[2] is 4, and bottom_right_brick_idx_delta[2] is 0 (= 4-4). For slice 3, top_left_brick_idx[3] is 5, and bottom_right_brick_idx_delta[3] is 7 (= 12-5). For slice 4, top_left_brick_idx[4] is 9, and bottom_right_brick_idx_delta[4] is 4 (= 13-9). As described above, top_left_brick_idx[4] and bottom_right_brick_idx_delta[4] do not need to be coded.

[0078] Furthermore, for entry_point_offset_minus1[], (the size of the coded data of the (j-1)th basic block row in the slice - 1) sent from the encoding unit 110 is encoded as entry_point_offset_minus1[j-1]. The number of entry_point_offset_minus1[] in a slice is equal to (the number of basic block rows in the slice - 1). In this embodiment, since bottom_right_brick_idx_delta[0] is 2, it can be seen that slice 0 consists of bricks with BID = 0 to 2.

[0079] Here, brick_height_minus1[0] is 2, and tile_rows_height_minus1 is 5. As a result, the number of basic block rows of the brick corresponding to BID=0 and the number of basic block rows of the brick corresponding to BID=1 are both brick_height_minus1[0]+1=3. Furthermore, the number of basic block rows of the brick corresponding to BID=2 is tile_rows_height_minus1+1=6. Therefore, the number of basic block rows of slice 0 is 3+3+6=12. Therefore, the range of j is 0 to 10.

[0080] For slice 1, top_left_brick_idx[1] is 3 and bottom_right_brick_idx_delta[1] is 0, so it is composed of a brick with BID=3. Also, because brick_row_height_minus1[2][0] is 1, the number of basic block rows in slice 1 (brick with BID=3) is 2. Therefore, the range of j is 0 only.

[0081] For slice 2, top_left_brick_idx[2] is 4 and bottom_right_brick_idx_delta[2] is 0, so it consists of a brick with BID=4. In addition, brick_row_height_minus1[2][0] is 1, num_brick_rows_minus1[2] is 1, and tile_rows_height_minus1 is 5. Therefore, the number of basic block rows in slice 2 (brick with BID=4) is {(tile_rows_height_minus1+1)-(brick_row_height_minus1[2][0]+1)}=4. Therefore, the range of j is 0 to 2.

[0082] Slice 3 is composed of bricks with BIDs 5 to 8 and 10 to 12. Because tile_rows_height_minus1 is 5, we can see that the number of basic block rows of the bricks with BIDs 8 and 10 is 6 (=tile_rows_height_minus1+1). Because brick_height_minus1[3] is 1, we can see that the number of basic block rows of the bricks with BIDs 5 to 7 is 2 (=brick_height_minus1[3]+1). Because brick_height_minus1[7] is 2, we can see that the number of basic block rows of the bricks with BIDs 11 and 12 is 3 (=brick_height_minus1[7]+1). Therefore, we can see that the total number of basic block rows of all the bricks that make up slice 3 is 2+2+2+6+6+3+3=24. Therefore, the range of j is 0 to 22.

[0083] Slice 4 is composed of a brick with BID=9 and a brick with BID=13. Because tile_rows_height_minus1 is 5, the number of basic block rows of each of the bricks with BID=9 and BID=13 is 6 (=tile_rows_height_minus1+1). Therefore, the total number of basic block rows of all bricks that make up slice 4 is 6+6=12. Therefore, the range of j is 0 to 10.

[0084] By such a process, the number of basic block rows in each slice is determined. In this embodiment, since the number of entry_point_offset_minus1 can be derived from other syntax, it is not necessary to encode num_entry_point_offset and include it in the header as in the conventional method. Therefore, according to this embodiment, the amount of data in the bit stream can be reduced.

[0085] In step S310, the control unit 199 judges whether or not the coding of all basic blocks in the input image has been completed. If the result of this judgment is that the coding of all basic blocks in the input image has been completed, the process proceeds to step S311. On the other hand, if there are basic blocks remaining in the input image that have not yet been coded, the process proceeds to step S303, and the subsequent processes are performed on the basic blocks that have not yet been coded.

[0086] In step S311, the in-loop filter unit 109 performs in-loop filtering on the reconstructed image generated in step S306, and outputs the image that has been subjected to the in-loop filtering.

[0087] Thus, according to this embodiment, it is not necessary to encode information indicating how many pieces of information indicating the starting positions of the code data of a basic block row that a brick has and include that information in the bit stream, and it is possible to generate a bit stream from which that information can be derived.

[0088] [Second embodiment] In this embodiment, an image decoding device that decodes a bit stream generated by the image encoding device according to the first embodiment will be described. Note that the configuration of the bit stream and other elements common to the first embodiment are the same as those described in the first embodiment, and therefore will not be described here.

[0089] An example of the functional configuration of the image decoding device according to this embodiment will be described with reference to the block diagram of FIG. 2. The separation decoding unit 202 acquires a bit stream generated by the image encoding device according to the first embodiment. The method of acquiring the bit stream is not limited to a specific acquisition method. For example, the bit stream may be acquired directly or indirectly from the image encoding device via a network such as a LAN or the Internet, or a bit stream stored inside or outside the image decoding device may be acquired. The separation decoding unit 202 then separates information on the decoding process and coded data on coefficients from the acquired bit stream and sends them to the decoding unit 203. The separation decoding unit 202 also decodes coded data of the header of the bit stream. In this embodiment, the separation decoding unit 202 decodes header information on the division of an image, such as the size of tiles, bricks, slices, and basic blocks, to generate division information, and outputs the generated division information to the image reproduction unit 205. That is, the separation decoding unit 202 performs the reverse operation of the integrated encoding unit 111 in FIG. 1.

[0090] The decoding unit 203 reproduces the quantized coefficients and prediction information by decoding the coded data output from the separate decoding unit 202. The inverse quantization / inverse transform unit 204 performs inverse quantization on the quantized coefficients to generate transform coefficients, and performs inverse orthogonal transform on the generated transform coefficients to reproduce the prediction error.

[0091] The frame memory 206 is a memory for storing image data of a reconstructed picture. The image reconstruction unit 205 generates a predicted image by appropriately referring to the frame memory 206 based on the input prediction information. The image reconstruction unit 205 then generates a reconstructed image from the generated predicted image and the prediction error reconstructed by the inverse quantization and inverse transform unit 204. The image reconstruction unit 205 then identifies the positions of tiles, bricks, and slices in the input image for the reconstructed image based on the partition information input from the separation decoding unit 202, and outputs the identified positions.

[0092] The in-loop filter unit 207 performs in-loop filtering such as deblocking filtering on the reconstructed image, and outputs the in-loop filtered image, similar to the in-loop filter unit 109. The control unit 299 controls the operation of the entire image decoding device, and controls the operation of each functional unit of the image decoding device.

[0093] Next, a bitstream decoding process by the image decoding device having the configuration shown in Fig. 2 will be described. In the following, a bitstream is input to the image decoding device on a frame-by-frame basis, but a bitstream of a still image for one frame may be input to the image decoding device. In addition, in this embodiment, for ease of explanation, only intra-prediction decoding process will be described, but the present invention is not limited to this and can also be applied to inter-prediction decoding process.

[0094] The separate decoding unit 202 separates information related to the decoding process and coded data related to coefficients from the input bit stream and sends it to the decoding unit 203. The separate decoding unit 202 also decodes coded data in the header of the bit stream. More specifically, the separate decoding unit 202 decodes basic block data partition information, tile data partition information, brick data partition information, slice data partition information 0, basic block row data synchronization information, basic block row data position information, etc. in FIG. 6 to generate partition information. The separate decoding unit 202 then outputs the generated partition information to the image reproduction unit 205. The separate decoding unit 202 also reproduces coded data in units of basic blocks of picture data and outputs it to the decoding unit 203.

[0095] The decoding unit 203 reproduces the quantization coefficients and the prediction information by decoding the coded data output from the separate decoding unit 202. The reproduced quantization coefficients are output to the inverse quantization and inverse transform unit 204, and the reproduced prediction information is output to the image reproduction unit 205.

[0096] The inverse quantization and inverse transform unit 204 performs inverse quantization on the input quantized coefficients to generate transform coefficients, and performs inverse orthogonal transform on the generated transform coefficients to reconstruct the prediction error. The reconstructed prediction error is output to the image reconstruction unit 205.

[0097] The image reproduction unit 205 generates a predicted image by appropriately referring to the frame memory 206 based on the prediction information input from the separate decoding unit 202. The image reproduction unit 205 then generates a reconstructed image from the generated predicted image and the prediction error reconstructed by the inverse quantization and inverse transform unit 204. The image reproduction unit 205 then specifies the shape of tiles, bricks, and slices as shown in FIG. 7 and their positions in the input image for the reconstructed image based on the partition information input from the separate decoding unit 202, and outputs (stores) them in the frame memory 206. The images stored in the frame memory 206 are used as references for prediction.

[0098] The in-loop filter unit 207 performs in-loop filtering such as deblocking filtering on the reconstructed image read from the frame memory 206 , and outputs (stores) the image that has been subjected to the in-loop filtering processing to the frame memory 206 .

[0099] The control unit 299 outputs the reproduced image stored in the frame memory 206. The output destination of the reproduced image is not limited to a specific output destination. For example, the control unit 299 may output the reproduced image to a display device included in the image decoding device and cause the display device to display the reproduced image. Furthermore, for example, the control unit 299 may transmit the reproduced image to an external device via a network such as a LAN or the Internet.

[0100] Next, the bitstream decoding process (bitstream decoding process having the configuration in FIG. 6) performed by the image decoding device according to this embodiment will be described with reference to the flowchart in FIG.

[0101] In step S401, the separation and decoding unit 202 separates information related to decoding processing and coded data related to coefficients from the input bit stream and sends them to the decoding unit 203. Also, the separation and decoding unit 202 decodes the coded data of the header of the bit stream. More specifically, the separation and decoding unit 202 decodes the basic block data division information, tile data division information, block data division information, slice data division information, basic block row data synchronization information, basic block row data position information, etc. in FIG. 6 to generate division information. Then, the separation and decoding unit 202 outputs the generated division information to the image reproduction unit 205. Also, the separation and decoding unit 202 reproduces the coded data of the basic block unit of the picture data and outputs it to the decoding unit 203.

[0102] In this embodiment, the division of the input image, which is the source of the bit stream encoding, is the division shown in FIG. 8. Information related to the input image, which is the source of the bit stream encoding, and its division can be derived from the division information.

[0103] From pic_width_in_luma_samples included in the image size information, it can be specified that the horizontal size (width) of the input image is 1152 pixels. Also, from pic_height_in_luma_samples included in the image size information, it can be specified that the vertical size (height) of the input image is 1152 pixels.

[0104] Also, since log2_ctu_size_minus2 = 4 in the basic block data division information, the size of the basic block can be derived as 64×64 pixels from 1<<log2_ctu_size_minus2+2.

[0105] Also, since single_tile_in_pic_flag = 0 in the tile data division information, it can be specified that the input image is divided into a plurality of tiles. And since uniform_tile_spacing_flag = 1, it can be specified that each tile has the same size (excluding the edges).

[0106] In addition, since tile_cols_width_minus1=5 and tile_rows_height_minus1=5, it can be determined that each tile is composed of 6×6 basic blocks. In other words, it can be determined that each tile is composed of 384×384 pixels. Since the input image is 1152×1152 pixels, it can be seen that the input image is divided into nine tiles, three in the horizontal direction and three in the vertical direction, and then coded.

[0107] Furthermore, since the brick data division information indicates that brick_splitting_present_flag=1, it can be specified that at least one tile in the input image is divided into multiple bricks.

[0108] Also, brick_split_flag[1], brick_split_flag[4], brick_split_flag[5], brick_split_flag[6], and brick_split_flag[8] are 0. This makes it possible to specify that the tiles corresponding to tile IDs = 1, 4, 5, 6, and 8 are not divided into bricks. In this embodiment, since the number of basic block rows of all tiles is 6, it can be seen that the number of basic block rows of bricks of tiles corresponding to tile IDs = 1, 4, 5, 6, and 8 is 6.

[0109] On the other hand, brick_split_flag[0], brick_split_flag[2], brick_split_flag[3], and brick_split_flag[7] are all 1. This allows us to specify that the tiles corresponding to tile IDs = 0, 2, 3, and 7 are divided into bricks. In addition, uniform_brick_spacing_flag[0], uniform_brick_spacing_flag[3], and uniform_brick_spacing_flag[7] are all 1. This allows us to specify that the tiles corresponding to tile IDs = 0, 3, and 7 are all divided into bricks of the same size.

[0110] Moreover, brick_height_minus1[0] and brick_height_minus1[7] are both 2. Therefore, it can be specified that the number of basic blocks in the vertical direction of the brick in both the tile corresponding to tile ID=0 and the tile corresponding to tile ID=7 is 3. It can also be specified that the number of bricks in both the tile corresponding to tile ID=0 and the tile corresponding to tile ID=7 is 2 (= number of basic block rows in the tile (6) / number of basic blocks in the vertical direction of the brick in the tile (3)).

[0111] Also, brick_height_minus1[3] is 1. Therefore, it can be determined that the number of basic blocks in the vertical direction of the brick in the tile corresponding to tile ID=3 is 2. It can also be determined that the number of bricks in the tile corresponding to tile ID=3 is 3 (= number of basic block rows in the tile (6) / number of basic blocks in the vertical direction of the brick in the tile (2)).

[0112] For the tile corresponding to tile ID=2, since num_brick_rows_minus1[2]=1, it can be specified that the tile is composed of two bricks. Furthermore, uniform_brick_spacing_flag[2]=0. This makes it possible to specify that the tile corresponding to tile ID=2 has a brick of a different size from the others. Furthermore, brick_row_height_minus1[2][0]=1, brick_row_height_minus1[2][1]=3, and the number of basic blocks in the vertical direction of all tiles is 6. This makes it possible to specify that the tile corresponding to tile ID=2 is composed of, from the top, a brick with two basic blocks in the vertical direction and a brick with four basic blocks in the vertical direction. Note that brick_row_height_minus1[2][1]=3 may not be coded. If there are two bricks in a tile, the height of the second brick can be calculated from the height of the tile and the height of the first brick in the tile (brick_row_height_minus1[2][0]=1).

[0113] Also, since single_brick_per_slice_flag=0 in slice data division information 0, it can be specified that at least one slice is composed of multiple bricks. In this embodiment, when uniform_brick_spacing_flag[i]=0, num_brick_rows_minus1[i] indicating (the number of bricks constituting the i-th tile-1) is included in the brick data division information. However, this is not limited to this.

[0114] For example, assume that the number of bricks constituting the i-th tile is 2 or more when brick_split_flag[i] = 1. Then, num_brick_rows_minus2[i] indicating (the number of bricks constituting the tile-2) may be decoded instead of num_brick_rows_minus1[i]. In this way, it is possible to decode a bitstream with a reduced number of syntax bits indicating the number of bricks constituting the tile.

[0115] Next, calculate the coordinates of the top-left and bottom-right boundaries of each brick. The coordinates are expressed as the horizontal and vertical positions of the basic block, with the top-left corner of the input image as the origin. For example, the coordinates of the top-left boundary of the third basic block from the left and the second basic block from the top are (3,2) and (4,3), respectively.

[0116] The coordinates of the upper left boundary of the brick with BID=0 in the tile corresponding to tile ID=0 are (0,0). The number of basic block rows of the brick with BID=0 is 3, and the number of basic blocks in the horizontal direction of all tiles is 6, so the coordinates of the lower right boundary are (3,3).

[0117] The coordinates of the upper left boundary of the brick with BID=1 in the tile corresponding to tile ID=0 are (0,3). The number of basic block rows of the brick with BID=1 is 3, and the number of basic blocks in the horizontal direction of all tiles is 6, so the coordinates of the lower right boundary are (6,6).

[0118] The coordinates of the upper left boundary of the tile (brick with BID=2) corresponding to tile ID=1 are (6,0). The number of basic block rows of the brick with BID=2 is 6, and the number of basic blocks in the horizontal direction of all tiles is 6, so the coordinates of the lower right boundary are (12,6).

[0119] The coordinates of the upper left boundary of the brick with BID=3 in the tile corresponding to tile ID=2 are (12,0). The number of basic block rows of the brick with BID=3 is 2, so the coordinates of the lower right boundary are (18,2).

[0120] The coordinates of the upper left boundary of the brick with BID=4 in the tile corresponding to tile ID=2 are (12,2). The number of basic block rows of the brick with BID=4 is 4, so the coordinates of the lower right boundary are (18,6).

[0121] The coordinates of the upper left boundary of the brick with BID=5 in the tile corresponding to tile ID=3 are (0,6). The number of basic block rows of the brick with BID=5 is 2, so the coordinates of the lower right boundary are (6,8).

[0122] The coordinates of the upper left boundary of the brick with BID=6 in the tile corresponding to tile ID=3 are (0,8). The number of basic block rows of the brick with BID=6 is 2, so the coordinates of the lower right boundary are (6,10).

[0123] The coordinates of the upper left boundary of the brick with BID=7 in the tile corresponding to tile ID=3 are (0,10). Since the number of basic block rows of the brick with BID=7 is 2, the coordinates of the lower right boundary are (6,12).

[0124] The coordinates of the upper left boundary of the tile corresponding to tile ID=4 (brick with BID=8) are (6,6). The number of basic block rows of the brick with BID=8 is 6, and the number of basic blocks in the horizontal direction of all tiles is 6, so the coordinates of the lower right boundary are (12,12).

[0125] The coordinates of the upper left boundary of the tile corresponding to tile ID=5 (brick with BID=9) are (12,6). The number of basic block rows of the brick with BID=9 is 6, and the number of basic blocks in the horizontal direction of all tiles is 6, so the coordinates of the lower right boundary are (18,12).

[0126] The coordinates of the upper left boundary of the tile corresponding to tile ID=6 (brick with BID=10) are (0,12). The number of basic block rows of the brick with BID=10 is 6, and the number of basic blocks in the horizontal direction of all tiles is 6, so the coordinates of the lower right boundary are (6,18).

[0127] The coordinates of the upper left boundary of the brick with BID=11 in the tile corresponding to tile ID=7 are (6,12). The number of basic block rows of the brick with BID=11 is 3, and the number of basic blocks in the horizontal direction of all tiles is 6, so the coordinates of the lower right boundary are (12,15).

[0128] The coordinates of the upper left boundary of the brick with BID=12 in the tile corresponding to tile ID=7 are (6,15). The number of basic block rows of the brick with BID=12 is 3, and the number of basic blocks in the horizontal direction of all tiles is 6, so the coordinates of the lower right boundary are (12,18).

[0129] The coordinates of the upper left boundary of the tile corresponding to tile ID=8 (brick with BID=13) are (12,12). The number of basic block rows of the brick with BID=13 is 6, and the number of basic blocks in the horizontal direction of all tiles is 6, so the coordinates of the lower right boundary are (18,18).

[0130] Next, the bricks included in each slice are identified. Because num_slices_in_pic_minus1=4, it can be determined that the number of slices in the input image is 5. In addition, it is possible to determine the corresponding brick from the slice_address of the slice to be processed. That is, when slice_address is N, it can be determined that the slice to be processed is slice N.

[0131] For slice 0, bottom_right_brick_idx_delta[0] is 2. This allows us to identify the bricks included in slice 0 as those included in the rectangular area surrounded by the coordinates of the upper left boundary of the brick with BID=0 and the coordinates of the lower right boundary of the brick with BID=2. The coordinates of the upper left boundary of the brick with BID=0 are (0,0), and the coordinates of the lower right boundary of the brick with BID=2 are (12,6), so we can identify the bricks included in slice 0 as those with BID=0 to 2.

[0132] For slice 1, since top_left_brick_idx[1] is 3 and bottom_right_brick_idx_delta[1] is 0, it can be identified that the brick contained in slice 1 is the brick with BID=3.

[0133] For slice 2, since top_left_brick_idx[2] is 4 and bottom_right_brick_idx_delta[2] is 0, it can be identified that the brick contained in slice 2 is the brick with BID=4.

[0134] In the case of slice 3, top_left_brick_idx[3] is 5 and bottom_right_brick_idx_delta[3] is 7. The coordinates of the upper left boundary of the brick with BID=5 are (0,6) and the coordinates of the lower right boundary of the brick with BID=12 are (12,18), so it is possible to identify the bricks included in the area with the upper left coordinates (0,6) and the lower right coordinates (12,18) as being included in slice 3. As a result, it is possible to identify the bricks corresponding to BID=5 to 8 and 10 to 12 as being included in slice 3. The coordinates of the lower right boundary of the brick corresponding to BID=9 are (18,12), and the coordinates of the lower right boundary of the brick corresponding to BID=13 are (18,18), both of which are outside the range of slice 3, and therefore are determined not to belong to slice 3.

[0135] In the case of slice 4, top_left_brick_idx[4] is 9 and bottom_right_brick_idx_delta[4] is 4. The coordinates of the upper left boundary of the brick with BID=9 are (12,6), and the coordinates of the lower right boundary of the brick with BID=13 are (18,18). Therefore, the bricks included in the area with the upper left coordinates (12,6) and the lower right coordinates (18,18) can be identified as being included in slice 4. As a result, the bricks corresponding to BID=9 and 13 can be identified as being included in slice 4. Here, the coordinates of the upper left boundary of the brick corresponding to BID=10 are (0,12), the coordinates of the upper left boundary of the brick corresponding to BID=11 are (6,12), and the coordinates of the upper left boundary of the brick corresponding to BID=12 are (6,15). Therefore, since all of them are outside the range of slice 4, they are determined not to belong to slice 4.

[0136] In this embodiment, the bricks included in slice 4, which is the last slice of the input image, are identified from top_left_brick_idx[4] and bottom_right_brick_idx_delta[4], but this is not limiting. It has already been derived that the bricks included in slices 0 to 3 are bricks with BID=0 to 2, BID=3, BID=4, BID=5 to 8, and 10 to 12, respectively, and it has already been derived that the input image is composed of 14 bricks with BID=0 to 13. Therefore, it is possible to identify that the bricks included in the last slice are the remaining bricks with BID=9 and 13. Therefore, even if top_left_brick_idx[4] and bottom_right_brick_idx_delta[4] are not coded, it is possible to identify the bricks included in the last slice. In this way, it is possible to decode a bitstream with a reduced bit amount in the header portion.

[0137] Also, in the basic block row data synchronization information, entropy_coding_sync_enabled_flag = 1. This indicates that entry_point_offset_minus1[j-1], which indicates (the size of the coded data of the (j-1)th basic block row in the slice - 1), is coded in the bitstream. The number of such entries is the number of basic block rows in the slice being processed - 1.

[0138] As described above, in this embodiment, since it is possible to identify the bricks that belong to each slice, entry_point_offset_minus1[j] is coded as a number obtained by subtracting 1 from the total number of basic block rows of the bricks that belong to the slice.

[0139] For slice 0, the sum of the number of basic block rows of the brick corresponding to BID=0 (3) + the number of basic block rows of the brick corresponding to BID=1 (3) + the number of basic block rows of the brick corresponding to BID=2 (6) is (3+3+6=12)-1=11. Therefore, for slice 0, an entry_point_offset_minus1[] of 11 is coded. In this case, the range of j is 0 to 10.

[0140] For slice 1, the number of basic block rows of the brick corresponding to BID=3 is (2)-1=1. Therefore, for slice 1, an entry_point_offset_minus1[] of 1 is coded. In this case, the range of j is 0 only.

[0141] For slice 2, the number of basic block rows of the brick corresponding to BID=4 is (4)-1=3. Therefore, for slice 3, an entry_point_offset_minus1[] of 3 is encoded. In this case, the range of j is 0 to 2.

[0142] For slice 3, the sum of the numbers of basic block rows of bricks corresponding to BID=5 to 8 and 10 to 12 is (2+2+2+6+6+3+3=24)-1=23. Therefore, for slice 3, an entry_point_offset_minus1[] of 23 is coded. In this case, the range of j is 0 to 22.

[0143] For slice 4, the sum of the number of basic block rows of the brick corresponding to BID=9 (6) + the number of basic block rows of the brick corresponding to BID=13 (6) is (6+6=12)-1=11. Therefore, for slice 4, an entry_point_offset_minus1[] of 11 is encoded. In this case, the range of j is 0 to 10.

[0144] As a result, the number of entry_point_offset_minus1 can be derived from other syntaxes without encoding num_entry_point_offset as in the past. Since the start position of the data of each basic block row is known, parallel decoding can be performed for each basic block row. The partition information derived by the separation decoding unit 202 is sent to the image reproduction unit 205, and is used to identify the position of the processing target in the input image in step S404.

[0145] In step S402, the decoding unit 203 reproduces the quantized coefficients and the prediction information by decoding the coded data separated by the separation decoding unit 202. In step S403, the inverse quantization / inverse transform unit 204 performs inverse quantization on the input quantized coefficients to generate transform coefficients, and performs inverse orthogonal transform on the generated transform coefficients to reproduce the prediction error.

[0146] In step S404, the image reproduction unit 205 generates a predicted image by appropriately referring to the frame memory 206 based on the prediction information input from the decoding unit 203. The image reproduction unit 205 then generates a reconstructed image from the generated predicted image and the prediction error reconstructed by the inverse quantization and inverse transform unit 204. The image reproduction unit 205 then identifies the positions of the tiles, bricks, and slices in the input image based on the partition information input from the separation decoding unit 202, synthesizes the reconstructed image at those positions, and outputs (stores) the image in the frame memory 206.

[0147] In step S405, the control unit 299 judges whether or not all basic blocks of the input image have been decoded. If the result of this judgment is that all basic blocks of the input image have been decoded, the process proceeds to step S406. On the other hand, if there are basic blocks remaining in the input image that have not yet been decoded, the process proceeds to step S402, where the decoding process is performed on the basic blocks that have not yet been decoded.

[0148] In step S406, the in-loop filter unit 207 performs in-loop filtering on the reconstructed image read from the frame memory 206, and outputs (stores) the image that has been subjected to the in-loop filtering in the frame memory 206.

[0149] In this way, according to this embodiment, it is possible to decode an input image from a "bitstream that does not contain information indicating how many pieces of information indicating the starting positions of basic block rows that a brick has been encoded" generated by the image encoding device according to the first embodiment.

[0150] The image encoding device according to the first embodiment and the image decoding device according to the second embodiment may be separate devices, or the image encoding device according to the first embodiment and the image decoding device according to the second embodiment may be integrated into one device.

[0151] [Third embodiment] 1 and 2 may be implemented by hardware, or part of them may be implemented by software. In the latter case, the functional units except for the frame memory 108 and the frame memory 206 may be implemented by software (computer program). A computer device capable of executing such a computer program is applicable to the image encoding device and the image decoding device described above.

[0152] An example of the hardware configuration of a computer device applicable to the above-mentioned image encoding device and image decoding device will be described with reference to the block diagram of Fig. 5. Note that the hardware configuration shown in Fig. 5 is merely an example of the hardware configuration of a computer device applicable to the above-mentioned image encoding device and image decoding device, and can be appropriately changed / modified.

[0153] The CPU 501 executes various processes using computer programs and data stored in the RAM 502 and the ROM 503. As a result, the CPU 501 controls the operation of the entire computer device, and executes or controls each process described as being performed by the image encoding device and the image decoding device above. In other words, the CPU 501 can function as each functional unit (except for the frame memory 108 and the frame memory 206) shown in Figures 1 and 2.

[0154] The RAM 502 has an area for storing computer programs and data loaded from the ROM 503 or the external storage device 506, and an area for storing data received from the outside via the I / F 507. The RAM 502 also has a work area used when the CPU 501 executes various processes. In this way, the RAM 502 can provide various areas as needed. The ROM 503 stores setting data, startup programs, and the like for the computer device.

[0155] The operation unit 504 is a user interface such as a keyboard, a mouse, and a touch panel screen, and the user can input various instructions to the CPU 501 by operating it.

[0156] Display unit 505 is configured with a liquid crystal screen, a touch panel screen, or the like, and can display the results of processing by CPU 501 as images, characters, etc. Note that display unit 505 may be a device such as a projector that projects images and characters.

[0157] The external storage device 506 is a large-capacity information storage device such as a hard disk drive device, etc. The external storage device 506 stores an OS (operating system), as well as computer programs and data for causing the CPU 501 to execute or control each of the processes described above as being performed by the image encoding device and image decoding device.

[0158] The computer programs stored in the external storage device 506 include computer programs for causing the CPU 501 to execute or control the functions of each functional unit in Figures 1 and 2 except for the frame memory 108 and the frame memory 206. The data stored in the external storage device 506 includes the information described above as known information, as well as various information related to encoding and decoding.

[0159] Computer programs and data stored in the external storage device 506 are loaded into the RAM 502 as appropriate under the control of the CPU 501 and are processed by the CPU 501 .

[0160] The frame memory 108 in the image encoding device in FIG. 1 and the frame memory 206 in the image encoding device in FIG.

[0161] The I / F 507 is an interface for performing data communication with an external device. For example, when the computer device is applied to an image encoding device, the image encoding device can output the generated bit stream to the outside via the I / F 507. When the computer device is applied to an image decoding device, the image decoding device can receive the bit stream via the I / F 507. The image decoding device can transmit the result of decoding the bit stream to the outside via the I / F 507. The CPU 501, the RAM 502, the ROM 503, the operation unit 504, the display unit 505, the external storage device 506, and the I / F 507 are all connected to a bus 508.

[0162] Note that the specific numerical values ​​used in the above description are used for the purpose of providing a specific description, and it is not intended that the above embodiments are limited to these numerical values. In addition, some or all of the above-described embodiments may be appropriately combined. In addition, some or all of the above-described embodiments may be selectively used.

[0163] (Other embodiments) The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.

[0164] The invention is not limited to the above-described embodiments, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]

[0165] 102: Image division unit 103: Block division unit 104: Prediction unit 105: Transformation and quantization unit 106: Inverse quantization and inverse transformation unit 107: Image reproduction unit 108: Frame memory 109: In-loop filter unit 110: Encoding unit 111: Integrated encoding unit

Claims

1. 1. An image decoding device that decodes an image including a rectangular area including one or more block rows each including a plurality of blocks from a bit stream obtained by encoding the image, comprising: a decoding means for decoding, from the bitstream, first information indicating an integer value n equal to the number of slices included in the image minus 1, a first flag regarding enabling parallel processing, second information used to identify a rectangular area to be processed first among a plurality of rectangular areas included in a target slice which is an i-th slice (i is an integer value) in the image, third information used to identify a rectangular area to be processed last among the plurality of rectangular areas, and fourth information corresponding to the number of blocks in the vertical direction of the rectangular area in the image; an identification means for identifying, in a state where a value of the first flag is 1, a second flag, which is a flag related to a slice mode and is decoded from a picture parameter set of the bit stream, indicates that a mode in which the slice is rectangular is used, and the target slice included in the image includes a plurality of rectangular regions at least in a horizontal direction or a vertical direction, a number of syntax elements used to identify a start position of coded data of a block row for the target slice, based on at least the second information, the third information, and the fourth information; Equipped with When the target slice in the image is an n-th corresponding slice, the third information is not decoded from the bitstream; The number of the syntax elements specified by the specifying means is included in a slice header of the bitstream.

2. An image decoding device comprising:

2. The decoding means comprises: The image decoding device according to claim 1 , wherein each block row in the rectangular area can be decoded in parallel.

3. The image decoding device according to claim 1 , wherein the first flag is an entropy_coding_sync_enabled_flag.

4. The image decoding device according to claim 1 , wherein the second flag is a rect_slice_flag.

5. The image decoding device as described in claim 1, characterized in that information on the number of the syntax elements is not signaled in the bitstream.

6. The image decoding device of claim 1, wherein the second information is decoded from the picture parameter set of the bitstream.

7. The image decoding device of claim 1, wherein the third information is decoded from the picture parameter set of the bitstream.

8. The image decoding device of claim 1, wherein the fourth information is decoded from the picture parameter set of the bitstream.

9. The image decoding device of claim 1, characterized in that the first rectangular area to be processed among the multiple rectangular areas included in the target slice is the rectangular area in the upper left corner of the multiple rectangular areas, and the last rectangular area to be processed among the multiple rectangular areas included in the target slice is the rectangular area in the lower right corner of the multiple rectangular areas.

10. An image decoding device as described in claim 1, characterized in that the size of each of the multiple blocks forming the block row is determined from fifth information decoded from a sequence parameter set in the bitstream.

11. The image decoding device described in Claim 10, characterized in that the size of each of the multiple blocks is derived by arithmetically shifting 1 to the left with a value that is the result of adding a predetermined value to the fifth information.

12. The image decoding device described in claim 1, characterized in that the image includes a plurality of slices from the 0th slice to the nth slice.

13. The image decoding device described in Claim 12, characterized in that the 0th slice is a slice in the upper left corner of the image, and the nth slice is a slice in the lower right corner of the image.

14. 1. An image decoding method for decoding an image including a rectangular area including one or more block rows each including a plurality of blocks from a bit stream obtained by encoding the image, the method comprising: a decoding process for decoding from the bitstream first information indicating an integer value n equal to the number of slices included in the image minus 1, a first flag regarding enabling parallel processing, second information used to identify a rectangular area to be processed first among a plurality of rectangular areas included in a target slice, which is a slice corresponding to the i-th slice (i is an integer value) in the image, third information used to identify a rectangular area to be processed last among the plurality of rectangular areas, and fourth information corresponding to the number of blocks in the vertical direction of the rectangular area in the image; a specifying step of specifying, for the target slice, the number of syntax elements used to specify a start position of coded data of a block row based on at least the second information, the third information, and the fourth information, in a state in which the value of the first flag is 1, a second flag, which is a flag related to a slice mode and is decoded from a picture parameter set of the bitstream, indicates that a mode in which the slice is rectangular is used, and the target slice included in the image includes a plurality of rectangular regions at least in a horizontal direction or a vertical direction; having When the target slice in the image is an n-th corresponding slice, the third information is not decoded from the bitstream; The number of the syntax elements identified by the identifying step is included in a slice header of the bitstream.

2. An image decoding method comprising:

15. An image coding device that codes an image including a rectangular area including one or more block rows each including a plurality of blocks, comprising: an encoding means for encoding into a bit stream first information indicating an integer value n equal to the number of slices included in the image minus 1, a first flag regarding enabling parallel processing, second information used to identify a rectangular area to be processed first among a plurality of rectangular areas included in a target slice which is an i-th slice (i is an integer value) in the image, third information used to identify a rectangular area to be processed last among the plurality of rectangular areas, and fourth information corresponding to the number of blocks in the vertical direction of the rectangular area in the image; an identification means for identifying, based on at least the second information, the third information, and the fourth information, the number of syntax elements used to identify a start position of coded data of a block row for the target slice, in a state in which a value of the first flag is 1, a second flag, which is a flag related to a slice mode and is coded in a picture parameter set of the bitstream, indicates that a mode in which the slice is rectangular is used, and the target slice included in the image includes a plurality of rectangular areas at least in a horizontal direction or a vertical direction; Equipped with When the target slice in the image is an n-th corresponding slice, the third information is not coded into the bitstream; The number of the syntax elements specified by the specifying means is included in a slice header of the bitstream.

1. An image encoding device comprising:

16. 1. An image coding method for coding an image including a rectangular region including one or more block rows each including a plurality of blocks, comprising: an encoding process for encoding into a bitstream first information indicating an integer value n equal to the number of slices included in the image minus 1, a first flag regarding enabling parallel processing, second information used to identify a rectangular area to be processed first among a plurality of rectangular areas included in a target slice, which is a slice corresponding to the i-th slice (i is a certain integer value) in the image, third information used to identify a rectangular area to be processed last among the plurality of rectangular areas, and fourth information corresponding to the number of blocks in the vertical direction of the rectangular area in the image; a specifying step of specifying, based on at least the second information, the third information, and the fourth information, the number of syntax elements used to specify a start position of coded data of a block row for the target slice, in a state in which a value of the first flag is 1, a second flag, which is a flag related to a slice mode and is coded in a picture parameter set of the bitstream, indicates that a mode in which the slice is rectangular is used, and the target slice included in the image includes a plurality of rectangular regions at least in a horizontal direction or a vertical direction; having When the target slice in the image is an n-th corresponding slice, the third information is not coded into the bitstream; The number of the syntax elements identified by the identifying step is included in a slice header of the bitstream.

13. An image coding method comprising:

17. A computer program for causing a computer to execute the image decoding method according to claim 14.

18. A computer program product for causing a computer to execute the image coding method according to claim 16.

Citation Information

Patent Citations

  • Image encoder, image encoding method and program, image decoder, and image decoding method and program

    JP2014011638A