Image Encoding Device, Image Decoding Device, Image Encoding Method, Image Decoding Method

The image decoding apparatus addresses redundant syntax in VVC by specifying block row positions, reducing bitstream code and enhancing encoding efficiency.

JP7703083B2Active Publication Date: 2025-07-04CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024101417
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-06-24
Publication Date
2025-07-04
Estimated Expiration
2039-06-21

AI Technical Summary

Technical Problem

In the VVC coding method, the number of entry_point_offset_minus1 indicating the start position of basic block rows becomes redundant syntax, leading to an increase in the amount of coding in the bitstream.

Method used

An image decoding apparatus that decodes a bitstream by specifying the number of blocks in the vertical direction of a rectangular region and determining the start position of block rows based on first and second flags, reducing the need for redundant syntax encoding.

Benefits of technology

This configuration reduces the amount of code in the bitstream by eliminating redundant syntax, thereby optimizing encoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007703083000001
    Figure 0007703083000001
  • Figure 0007703083000002
    Figure 0007703083000002
  • Figure 0007703083000003
    Figure 0007703083000003
Patent Text Reader

Abstract

To provide a technique that reduces an encoding amount of a bit stream by reducing redundant syntax.SOLUTION: In such a state that a value of a first flag is a first value, a second flag indicates that each rectangular region is composed of one slice, the number of blocks in the vertical direction of the rectangular region corresponding to an object slice of an image is smaller than the number of blocks in the vertical direction of a tile including the rectangular region and the object slice is rectangular, the number of information pieces for specifying a head position of code data of a block line is specified for the object slice on the basis of the first information and the second information. The code data of the block line is decoded on the basis of at least the number of information pieces for specifying the specified head position and information for specifying the head position.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image encoding / decoding technology.

Background Art

[0002] As an encoding method for compressed recording of moving images, the HEVC (High Efficiency Video Coding) encoding method (hereinafter referred to as HEVC) is known. In HEVC, in order to improve the encoding efficiency, a basic block having a size larger than that of a conventional macroblock (16×16 pixels) is adopted. This basic block having a large size is called a CTU (Coding Tree Unit), and its size is up to 64×64 pixels at maximum. A CTU is further divided into sub-blocks that are units for performing prediction and conversion.

[0003] Also in HEVC, it is possible to divide a picture into a plurality of tiles or slices and perform encoding. There is little data dependency between each tile or slice, and encoding / decoding processing can be performed in parallel. One of the great advantages of tile and slice division is that processing can be executed in parallel using a multi-core CPU or the like, and the processing time can be shortened.

[0004] Also, each slice is encoded by the conventional binary arithmetic coding method adopted in HEVC. That is, each syntax element is binarized to generate a binary signal. Each syntax element is given a generation probability in advance as a table (hereinafter referred to as a generation probability table), and the binary signal is arithmetically encoded based on the generation probability table. This generation probability table is used as decoding information for decoding subsequent codes during decoding. During encoding, it is used as encoding information for subsequent encoding. And every time encoding is performed, the generation probability table is updated based on statistical information as to whether the encoded binary signal is a symbol with a higher generation probability.

[0005] In addition, HEVC has a technique called Wavefront Parallel Processing (hereinafter referred to as WPP) for processing entropy encoding and decoding in parallel. In WPP, by applying the probability table at the time of encoding the block at a specified position in advance to the block at the left end of the next row, it is possible to suppress a decrease in encoding efficiency and perform parallel encoding processing of blocks in row units. In order to enable parallel processing in block row units, the entry_point_offset_minus1 indicating the start position of each block row in the bitstream and the num_entry_point_offsets indicating the number thereof are encoded in the slice header. Patent Document 1 discloses a technique related to WPP.

[0006] In recent years, activities have been started to standardize an even more efficient coding method as a successor to HEVC. The JVET (Joint Video Experts Team) was established between ISO / IEC and ITU-T, and standardization is underway for the VVC (Versatile Video Coding) coding method (hereinafter referred to as VVC). In VVC, it is being considered to divide a tile into rectangles (bricks) composed of a plurality of block rows. And a slice is configured to include one or more bricks.

Prior Art Documents

Patent Documents

[0007]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0008] In VVC, the blocks that make up a slice can be derived in advance, and furthermore, the number of basic block rows included in those blocks can also be derived from other syntax. Therefore, it is possible to derive the number of entry_point_offset_minus1 indicating the start position of the basic block rows belonging to the slice without using num_entry_point_offset. Therefore, num_entry_point_offset becomes redundant syntax. The present invention provides a technique for reducing the amount of coding of a bitstream by reducing redundant syntax.

Means for Solving the Problems

[0009] One aspect of the present invention is an image decoding apparatus that decodes an image from a bitstream obtained by encoding an image including a rectangular region including one or more block rows composed of a plurality of blocks and a tile including a plurality of blocks, the method comprising: decoding, from the bitstream, first information for specifying the number of blocks in the vertical direction of the rectangular region in the image, a first flag related to enabling parallel processing, and a second flag indicating whether each rectangular region consists of one slice; decoding means for decoding, from a slice header of the bitstream, second information indicating an ID of a rectangular region corresponding to a target slice in the image; specifying means for specifying, based on the first information and the second information, the number of information for specifying the start position of the coded data of the block row for the target slice in a state where the value of the first flag is a first value, the second flag indicates that each rectangular region consists of one slice, the number of blocks in the vertical direction of the rectangular region corresponding to the target slice of the image is smaller than the number of blocks in the vertical direction of the tile including the rectangular region, and the target slice is rectangular; and the decoding means decodes the coded data of the block row based at least on the number of information for specifying the start position specified by the specifying means and the information for specifying the start position.

Effects of the Invention

[0010] According to the configuration of the present invention, the amount of code of the bitstream can be reduced by reducing redundant syntax.

Brief Description of Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Embodiments for Carrying Out the Invention

[0012] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the invention according to the claims. Although a plurality of features are described in the embodiments, not all of these plurality of features are essential to the invention, and the plurality of features may be arbitrarily combined. Further, in the accompanying drawings, the same or similar configurations are denoted by the same reference numerals, and redundant descriptions are omitted.

[0013] [First Embodiment] First, a functional configuration example of the image encoding apparatus according to the present embodiment will be described with reference to the block diagram of FIG. 1. An input image to be encoded is input to the image division unit 102. The input image may be an image of each frame constituting a moving image or a still image. The image division unit 102 divides the input image into "one or a plurality of tiles". A tile is a set of consecutive basic blocks that covers a rectangular area within the input image. The image division unit 102 further divides each tile into one or a plurality of blocks. A block is a rectangular area (a rectangular area including one or more block rows of a plurality of blocks having a size equal to or smaller than that of the tile) composed of one or a plurality of rows of basic blocks (basic block rows) within the tile. The image division unit 102 further divides the input image into slices composed of "one or a plurality of tiles" or "one or more blocks within one tile". A slice is a basic unit of encoding, and header information such as information indicating the type of slice is added for each slice. An example of dividing the input image into 4 tiles, 4 slices, and 11 blocks is shown in FIG. 7. The upper left tile is divided into 1 block, the lower left tile is divided into 2 blocks, the upper right tile is divided into 5 blocks, and the lower right tile is divided into 3 blocks. And the left slice is configured to include 3 blocks, the upper right slice is configured to include 2 blocks, the right center slice is configured to include 3 blocks, and the lower right slice is configured to include 3 blocks. The image division unit 102 outputs information regarding the size as division information for each of the tiles, blocks, and slices thus divided.

[0014] The block division unit 103 divides the image of the basic block row (basic block row image) output from the image division unit 102 into a plurality of basic blocks, and outputs the image of the basic block unit (block image) to the subsequent stage.

[0015] The prediction unit 104 divides the image in basic block units into sub-blocks, performs intra prediction which is in-frame prediction in sub-block units, inter prediction which is inter-frame prediction, etc., and generates a predicted image. Intra prediction across blocks (intra prediction using pixels of blocks in other blocks) and prediction of motion vectors across blocks (prediction of motion vectors using motion vectors of blocks in other blocks) are not performed. Further, the prediction unit 104 calculates and outputs a prediction error from the input image and the predicted image. Also, the prediction unit 104 outputs information necessary for prediction (prediction information), such as information on sub-block division methods, prediction modes, motion vectors, etc., together with the prediction error.

[0016] The transform and quantization unit 105 orthogonally transforms the prediction error in sub-block units to obtain transform coefficients, and quantizes the obtained transform coefficients to obtain quantization coefficients. The inverse quantization and inverse transform unit 106 inverse quantizes the quantization coefficients output from the transform and quantization unit 105 to reproduce the transform coefficients, and further inverse orthogonally transforms the reproduced transform coefficients to reproduce the prediction error.

[0017] The frame memory 108 functions as a memory for storing the reproduced image. The image reproduction unit 107 appropriately refers to the frame memory 108 based on the prediction information output from the prediction unit 104 to generate a predicted image, and generates and outputs a reproduced image from the predicted image and the input prediction error.

[0018] The in-loop filter unit 109 performs in-loop filter processing such as deblocking filters and sample adaptive offsets on the reproduced image, and outputs the image (filtered image) subjected to the in-loop filter processing.

[0019] The encoding unit 110 generates coded data (encoded data) by encoding the quantization coefficients output from the transform and quantization unit 105 and the prediction information output from the prediction unit 104, and outputs the generated coded data.

[0020] The integrated encoding unit 111 generates header encoded data using the segmentation information output from the image segmentation unit 102, and generates and outputs a bitstream including the generated header encoded data and the encoded data output from the encoding unit 110. The control unit 199 controls the operation of the entire image encoding apparatus and controls the operation of each functional unit of the above-described image encoding apparatus.

[0021] Next, the encoding process for the input image by the image encoding apparatus having the configuration shown in FIG. 1 will be described. In the present embodiment, for the sake of simplicity of explanation, only the process of intra prediction encoding will be described, but the present invention is not limited thereto and is also applicable to the process of inter prediction encoding. Further, in the present embodiment, for the purpose of specific explanation, the block segmentation unit 103 will be described as dividing the basic block row image output from the image segmentation unit 102 in units of "basic blocks having a size of 64×64 pixels".

[0022] The image segmentation unit 102 divides the input image into tiles and bricks. An example of dividing the input image by the image segmentation unit 102 is shown in FIG. 8. In FIG. 8, the input image is divided into four tiles and ten bricks. In the present embodiment, an input image having a size of 1536×1024 pixels is divided into four tiles (the size of one tile is 768×512 pixels).

[0023] The upper left tile is not divided into bricks (equivalent to dividing one tile into one brick), and as a result, tile = brick. The lower left tile is divided into two bricks (the height of one brick is 256 pixels), and the upper right tile is divided into four bricks (the height of one brick is 128 pixels). Also, the lower right tile is divided into three bricks (the heights of the respective bricks are 192 pixels, 128 pixels, and 192 pixels in order from the top).

[0024] Also, each block within a tile is assigned an ID in ascending order from the top in the raster-ordered tile. The BID shown in FIG. 8 is the ID of the block. In this embodiment, it is assumed that each slice is composed of only one block. That is, the block with BID = 0 corresponds to slice 0, the block with BID = 1 corresponds to slice 1, and so on, where the slice and the block have the same ID associated with each other.

[0025] Then, for each of the divided tiles, blocks, and slices, the image division unit 102 outputs information regarding the size as division information to the integrated encoding unit 111. Also, the image division unit 102 divides each block into basic block row images and outputs the divided basic block row images to the block division unit 103.

[0026] The block division unit 103 divides the basic block row image output from the image division unit 102 into a plurality of basic blocks and outputs block images (64×64 pixels), which are images in basic block units, to the subsequent prediction unit 104.

[0027] The prediction unit 104 divides the image in basic block units into sub-blocks, determines an intra prediction mode such as horizontal prediction or vertical prediction in sub-block units, and generates a prediction image from the determined intra prediction mode and the encoded pixels. Further, the prediction unit 104 calculates a prediction error from the input image and the prediction image and outputs the calculated prediction error to the conversion / quantization unit 105. Also, the prediction unit 104 outputs information such as the sub-block division method and the intra prediction mode as prediction information to the encoding unit 110 and the image reproduction unit 107.

[0028] The conversion / quantization unit 105 performs orthogonal transformation (orthogonal transformation process corresponding to the size of the sub-block) on the prediction error output from the prediction unit 104 in sub-block units to obtain transformation coefficients (orthogonal transformation coefficients). Then, the conversion / quantization unit 105 quantizes the obtained transformation coefficients to obtain quantization coefficients. Then, the conversion / quantization unit 105 outputs the obtained quantization coefficients to the encoding unit 110 and the inverse quantization / inverse transformation unit 106.

[0029] The inverse quantization and inverse transformation unit 106 inverse quantizes the quantized coefficients output from the transformation and quantization unit 105 to reproduce the transformation coefficients, and further inverse orthogonally transforms the reproduced transformation coefficients to reproduce the prediction error. Then, the inverse quantization and inverse transformation unit 106 outputs the reproduced prediction error to the image reproduction unit 107.

[0030] The image reproduction unit 107 generates a predicted image by appropriately referring to the frame memory 108 based on the prediction information output from the prediction unit 104, and generates a reproduced image from the predicted image and the prediction error input from the inverse quantization and inverse transformation unit 106. Then, the image reproduction unit 107 stores the generated reproduced image in the frame memory 108.

[0031] The in-loop filter unit 109 reads out the reproduced image from the frame memory 108, and performs in-loop filter processing such as a deblocking filter and a sample adaptive offset on the read-out reproduced image. Then, the in-loop filter unit 109 stores (re-stores) the image subjected to the in-loop filter processing in the frame memory 108 again.

[0032] The encoding unit 110 generates coded data by entropy encoding the quantized coefficients output from the transformation and quantization unit 105 and the prediction information output from the prediction unit 104. Although the method of entropy encoding is not particularly specified, Golomb encoding, arithmetic encoding, Huffman encoding, etc. can be used. Then, the encoding unit 110 outputs the generated coded data to the integrated encoding unit 111.

[0033] The integrated encoding unit 111 generates header coded data using the segmentation information output from the image segmentation unit 102, and generates and outputs a bitstream by multiplexing the generated header coded data and the coded data output from the encoding unit 110. The output destination of the bitstream is not limited to a specific output destination, and it may be output (stored) in a memory inside or outside the image encoding device, or transmitted to an external device capable of communicating with the image encoding device via a network such as a LAN or the Internet.

[0034] Next, FIG. 6 shows an example of the format of the bitstream output by the integrated encoding unit 111 (the coded data by VVC encoded by the image encoding apparatus). The bitstream in FIG. 6 includes a sequence parameter set (SPS), which is header information including information related to the encoding of the sequence. The bitstream in FIG. 6 also includes a picture parameter set (PPS), which is header information including information related to the encoding of the picture. The bitstream in FIG. 6 further includes a slice header (SLH), which is header information including information related to the encoding of the slice. The bitstream in FIG. 6 also includes the coded data of each block (blocks 0 to (N - 1) in FIG. 6).

[0035] The SPS includes image size information and basic block data division information. The PPS includes tile data division information, which is tile division information, block data division information, which is block division information, slice data division information 0, which is slice division information, and basic block row data synchronization information. The SLH includes slice data division information 1 and basic block row data position information.

[0036] First, the SPS will be described. The SPS includes pic_width_in_luma_samples which is information 601 as the image size information and pic_height_in_luma_samples which is information 602. pic_width_in_luma_samples represents the horizontal size (number of pixels) of the input image, and pic_height_in_luma_samples represents the vertical size (number of pixels) of the input image. In this embodiment, since the input image in FIG. 8 is used as the input image, pic_width_in_luma_samples = 1536 and pic_height_in_luma_samples = 1024. The SPS also includes log2_ctu_size_minus2 which is information 603 as the basic block data division information. log2_ctu_size_minus2 represents the size of the basic block. The number of pixels in the vertical and horizontal directions of the basic block is indicated by 1<<(log2_ctu_size_minus2 + 2). In this embodiment, since the size of the basic block is 64×64 pixels, the value of log2_ctu_size_minus2 is 4.

[0037] Next, the PPS will be described. The PPS includes information 604 - 607 as the tile data division information. Information 604 is single_tile_in_pic_flag which indicates whether the input image is divided into multiple tiles and encoded. When single_tile_in_pic_flag = 1, it indicates that the input image is not divided into multiple tiles and encoded. On the other hand, when single_tile_in_pic_flag = 0, it indicates that the input image is divided into multiple tiles and encoded.

[0038] Information 605 is information included in tile data division information when single_tile_in_pic_flag = 0. Information 605 is uniform_tile_spacing_flag indicating whether each tile has the same size. When uniform_tile_spacing_flag = 1, it indicates that each tile has the same size, and when uniform_tile_spacing_flag = 0, it indicates that there are tiles with different sizes.

[0039] Information 606 and Information 607 are information included in tile data division information when uniform_tile_spacing_flag = 1. Information 606 is tile_cols_width_minus1 indicating (the number of basic blocks in the horizontal direction of the tile - 1). Information 607 is tile_rows_height_minus1 indicating (the number of basic blocks in the vertical direction of the tile - 1). The number of tiles in the horizontal direction of the input image is obtained as the quotient when the number of basic blocks in the horizontal direction of the input image is divided by the number of basic blocks in the horizontal direction of the tile. If there is a remainder in this division, the number obtained by adding 1 to the quotient is taken as "the number of tiles in the horizontal direction of the input image". Also, the number of tiles in the vertical direction of the input image is obtained as the quotient when the number of basic blocks in the vertical direction of the input image is divided by the number of basic blocks in the vertical direction of the tile. If there is a remainder in this division, the number obtained by adding 1 to the quotient is taken as "the number of tiles in the vertical direction of the input image". Also, the total number of tiles in the input image can be obtained by calculating the number of tiles in the horizontal direction of the input image × the number of tiles in the vertical direction of the input image.

[0040] Note that when uniform_tile_spacing_flag = 0, since tiles with different sizes from others are included, the number of tiles in the horizontal direction of the input image, the number of tiles in the vertical direction of the input image, and the vertical and horizontal sizes of each tile are coded.

[0041] PPS also includes information 608 to 613 as brick data splitting information. Information 608 is brick_splitting_present_flag. When brick_splitting_present_flag = 1, it indicates that one or more tiles in the input image are split into multiple bricks. On the other hand, when brick_splitting_present_flag = 0, it indicates that each tile in the input image is composed of a single brick.

[0042] Information 609 is information included in the brick data splitting information when brick_splitting_present_flag = 1. Information 609 is brick_split_flag[] which indicates whether each tile is split into multiple bricks. brick_split_flag[] indicating whether the i-th tile is split into multiple bricks is denoted as brick_split_flag[i]. When brick_split_flag[i] = 1, it indicates that the i-th tile is split into multiple bricks, and when brick_split_flag[i] = 0, it indicates that the i-th tile is composed of a single brick.

[0043] The information 610 is the uniform_brick_spacing_flag[i] that indicates whether the sizes of the respective bricks constituting the i-th tile are the same when brick_split_flag[i]=1. If brick_split_flag[i]=0 for all i, the information 610 is not included in the brick data division information. The information 610 includes the uniform_brick_spacing_flag[i] for i that satisfies brick_split_flag[i]=1. When uniform_brick_spacing_flag[i]=1, it indicates that the sizes of the respective bricks constituting the i-th tile are the same. On the other hand, when uniform_brick_spacing_flag[i]=0, it indicates that there is a brick with a size different from the others among the bricks constituting the i-th tile.

[0044] The information 611 is the information included in the brick data division information when uniform_brick_spacing_flag[i]=1. The information 611 is the brick_height_minus1[i] that indicates (the number of basic blocks in the vertical direction of the brick in the i-th tile - 1).

[0045] Note that the number of basic blocks in the vertical direction of a brick can be obtained by dividing the number of pixels in the vertical direction of the brick by the number of pixels in the vertical direction of a basic block (64 pixels in this embodiment). Also, the number of bricks constituting a tile is obtained as the quotient when the number of basic blocks in the vertical direction of the tile is divided by the number of basic blocks in the vertical direction of a brick. If there is a remainder after this division, the number obtained by adding 1 to the quotient is taken as the "number of bricks constituting the tile". For example, assume that the number of basic blocks in the vertical direction of a tile is 10 and the value of brick_height_minus1 is 2. At this time, this tile is divided into four bricks in order from the top: a brick with 3 basic block rows, a brick with 3 basic block rows, a brick with 3 basic block rows, and a brick with 1 basic block row.

[0046] The information 612 is num_brick_rows_minus1[i] which indicates (the number of bricks constituting the i-th tile - 1) for i satisfying uniform_brick_spacing_flag[i]=0.

[0047] In this embodiment, when uniform_brick_spacing_flag[i]=0, num_brick_rows_minus1[i] which indicates (the number of bricks constituting the i-th tile - 1) is included in the brick data division information. However, it is not limited to this.

[0048] For example, assume that when brick_split_flag[i]=1, the number of bricks constituting the i-th tile is 2 or more, and num_brick_rows_minus2[i] which indicates (the number of bricks constituting the tile - 2) may be encoded instead of num_brick_rows_minus1[i]. By doing so, the number of bits of the syntax indicating the number of bricks constituting the tile can be reduced. For example, when the tile is composed of 2 bricks and num_brick_rows_minus1[i] is Golomb-encoded, 3-bit data of "010" indicating "1" is encoded. On the other hand, when num_brick_rows_minus2[i] which indicates (the number of bricks constituting the tile - 2) is Golomb-encoded, 1-bit data of "0" indicating 0 is encoded.

[0049] The information 613 is brick_row_height_minus1[i][j] which indicates (the number of basic blocks in the vertical direction of the j-th brick in the i-th tile - 1) for i that satisfies uniform_brick_spacing_flag[i]=0. brick_row_height_minus1[i][j] is encoded by the number of num_brick_rows_minus1[i]. The number of basic blocks in the vertical direction of the bottom brick in the tile can be obtained by subtracting the sum of "brick_row_height_minus1 + 1" from the number of basic blocks in the vertical direction of the tile. For example, assume that the number of basic blocks in the vertical direction of the tile = 10, num_brick_rows_minus1 = 3, brick_row_height_minus1 = 2, 1, 2. At this time, the number of basic blocks in the vertical direction of the bottom brick in the tile is 10 - (3 + 2 + 3) = 2.

[0050] In addition, the PPS contains information 614 to 618 as slice data division information 0. The information 614 is single_brick_per_slice_flag. When single_brick_per_slice_flag = 1, it indicates that all slices in the input image are composed of a single brick. That is, each slice is composed of only one brick. On the other hand, when single_brick_per_slice_flag = 0, it indicates that one or more slices in the input image are composed of multiple bricks.

[0051] The information 615 is rect_slice_flag, which is the information included in the slice data division information 0 when single_brick_per_slice_flag = 0. rect_slice_flag indicates whether the tiles included in the slice are in raster order or rectangular. Fig. 9(a) shows the relationship between the tiles and the slice when rect_slice_flag = 0, indicating that the tiles within the slice are encoded in raster order. On the other hand, Fig. 9(b) shows the relationship between the tiles and the slice when rect_slice_flag = 1, indicating that multiple tiles within the slice are rectangular.

[0052] The information 616 is num_slices_in_pic_minus1, which is the information included in the slice data division information 0 when rect_slice_flag = 1 and single_brick_per_slice_flag = 0. num_slices_in_pic_minus1 indicates (the number of slices in the input image - 1).

[0053] The information 617 is top_left_brick_idx[i] that indicates the index of the top - left brick of each slice (the i - th slice) in the input image.

[0054] The information 618 is bottom_right_brick_idx_delta[i] that indicates the difference between the index of the top - left brick and the index of the bottom - right brick in each slice in the input image. However, since the index of the top - left brick of the first slice in the input image is determined to be 0, top_left_brick_idx[0] of the first slice is not encoded.

[0055] In addition, information 619 is encoded and included in PPS as basic block row data synchronization information. Information 619 is the entropy_coding_sync_enabled_flag. When entropy_coding_sync_enabled_flag = 1, the probability table at the time of processing the basic block at a predetermined position in the basic block row adjacent above is applied to the leftmost block. Thereby, parallel processing of entropy encoding / decoding is enabled for each basic block row.

[0056] Next, SLH will be described. Information 620 to 621 is encoded and included in SLH as slice data division information 1. Information 620 is the slice_address that is included in the slice data division information 1 when rect_slice_flag = 1 or the number of bricks in the input image is 2 or more. The slice_address indicates the BID at the head of the slice when rect_slice_flag = 0, and indicates the number of the current slice when rect_slice_flag = 1.

[0057] Information 621 is the num_bricks_in_slice_minus1 that is included in the slice data division information 1 when rect_slice_flag = 0 and single_brick_per_slice_flag = 0. The num_bricks_in_slice_minus1 indicates (the number of bricks in the slice - 1).

[0058] Information 622 is included in SLH as basic block row data position information. Information 622 is the entry_point_offset_minus1[]. The entry_point_offset_minus1[] is encoded and included in the basic block row data position information by the number of (the number of basic block rows in the slice - 1) when entropy_coding_sync_enabled_flag = 1.

[0059] entry_point_offset_minus1[] represents the entry point of the code data of the basic block line, that is, the starting position of the code data of the basic block line. entry_point_offset_minus1[j - 1] indicates the entry point of the code data of the j-th basic block line. Since the starting position of the code data of the 0-th basic block line is the same as the starting position of the code data of the slice to which the basic block line belongs, it is omitted. And, {the size of the code data of the (j - 1)-th basic block line - 1} is encoded as entry_point_offset_minus1[j - 1].

[0060] Next, the encoding process of the input image by the image encoding device in this embodiment (the generation process of the bitstream having the configuration shown in FIG. 6) will be described according to the flowchart of FIG. 3.

[0061] First, in step S301, the image division unit 102 divides the input image into tiles, bricks, and slices. Then, for each of the divided tiles, bricks, and slices, the image division unit 102 outputs information regarding the size as division information to the integrated encoding unit 111. Further, the image division unit 102 divides each brick into basic block line images, and outputs the divided basic block line images to the block division unit 103.

[0062] In step S302, the block division unit 103 divides the basic block line image into a plurality of basic blocks, and outputs block images, which are images in units of basic blocks, to the subsequent prediction unit 104.

[0063] In step S303, the prediction unit 104 divides the image in units of basic blocks output from the block division unit 103 into sub-blocks, determines the intra prediction mode in units of sub-blocks, and generates a predicted image from the determined intra prediction mode and the encoded pixels. Further, the prediction unit 104 calculates a prediction error from the input image and the predicted image, and outputs the calculated prediction error to the conversion / quantization unit 105. Also, the prediction unit 104 outputs information such as the sub-block division method and the intra prediction mode as prediction information to the encoding unit 110 and the image reproduction unit 107.

[0064] In step S304, the conversion / quantization unit 105 orthogonally transforms the prediction error output from the prediction unit 104 in units of sub-blocks to obtain conversion coefficients (orthogonal conversion coefficients). Then, the conversion / quantization unit 105 quantizes the obtained conversion coefficients to obtain quantization coefficients. Then, the conversion / quantization unit 105 outputs the obtained quantization coefficients to the encoding unit 110 and the inverse quantization / inverse conversion unit 106.

[0065] In step S305, the inverse quantization / inverse conversion unit 106 inverse quantizes the quantization coefficients output from the conversion / quantization unit 105 to reproduce the conversion coefficients, and further inverse orthogonally transforms the reproduced conversion coefficients to reproduce the prediction error. Then, the inverse quantization / inverse conversion unit 106 outputs the reproduced prediction error to the image reproduction unit 107.

[0066] In step S306, the image reproduction unit 107 generates a predicted image by appropriately referring to the frame memory 108 based on the prediction information output from the prediction unit 104, and generates a reproduced image from the predicted image and the prediction error input from the inverse quantization / inverse conversion unit 106. Then, the image reproduction unit 107 stores the generated reproduced image in the frame memory 108.

[0067] In step S307, the encoding unit 110 generates coded data by entropy encoding the quantization coefficients output from the conversion / quantization unit 105 and the prediction information output from the prediction unit 104.

[0068] Here, when entropy_coding_sync_enabled_flag = 1, the probability occurrence table at the time of processing the basic block at a predetermined position in the basic block row adjacent above is applied before processing the basic block at the left end of the next basic block row. In the present embodiment, it is described assuming that entropy_coding_sync_enabled_flag = 1.

[0069] In step S308, the control unit 199 determines whether or not the encoding of all the basic blocks in the slice has been completed. As a result of this determination, if the encoding of all the basic blocks in the slice has been completed, the process proceeds to step S309. On the other hand, if there remains an uncoded basic block (uncoded basic block) among the basic blocks in the slice, the process proceeds to step S303 to encode the uncoded basic block.

[0070] In step S309, the integrated encoding unit 111 generates header encoded data using the segmentation information output from the image segmentation unit 102, and generates and outputs a bit stream including the generated header encoded data and the encoded data output from the encoding unit 110.

[0071] When the input image is divided as shown in FIG. 8, single_tile_in_pic_flag of the tile data segmentation information is 0, and uniform_tile_spacing_flag is 1. Also, tile_cols_width_minus1 is 11, and tile_rows_height_minus1 is 7.

[0072] brick_splitting_present_flag of the brick data segmentation information is 1. Since the upper left tile is not divided into bricks, brick_split_flag[0] is 0. However, since the lower left tile, the upper right tile, and the lower right tile are all divided into bricks, brick_split_flag[1], brick_split_flag[2], and brick_split_flag[3] are 1.

[0073] Also, since the upper-right tile and the lower-left tile are both divided into bricks of the same size, uniform_brick_spacing_flag[1] and uniform_brick_spacing_flag[2] are 1. For the lower-right tile, since the size of the brick with BID = 8 is different from the sizes of the bricks with BID = 7 and BID = 9, uniform_brick_spacing_flag[3] is 0.

[0074] brick_height_minus1[1] is 1 and brick_height_minus1[2] is 3. brick_row_height_minus1[3][0] is 2 and brick_row_height_minus1[3][1] is 1. Note that the value of num_brick_rows_minus1[3] is 2. When encoding the above-mentioned syntax of num_brick_rows_minus2[3] instead of num_brick_rows_minus1[3], the value becomes 1.

[0075] Also, as described above, in this embodiment, since each slice is assumed to be composed of only one brick, single_brick_per_slice_flag of slice data division information 0 is 1.

[0076] Also, as slice data division information 1, first, the first BID in the slice is encoded as slice_address. As described above, in this embodiment, since each slice is assumed to be composed of only one brick, 0 is encoded as slice_address for slice 0, 1 for slice 1, and N for slice N.

[0077] Also, regarding the basic block row data position information, the size of the coded data of the ((j - 1)-th basic block row in the slice sent from the coding unit 110 minus 1 is coded as entry_point_offset_minus1[j - 1]. The number of entry_point_offset_minus1[] in the slice is equal to (the number of basic block rows in the slice minus 1). In this embodiment, the upper left tile consists of a single block, and the number of basic block rows of that is tile_rows_height_minus1 + 1 = 8. Thus, the range of j is 0 to 6. The number of basic block rows in the slice (block) of the lower left tile can be obtained as brick_height_minus1[2] + 1 = 4. The number of basic block rows in the slice (block) of the upper right tile can be obtained as brick_height_minus1[1] + 1 = 2. For each slice (block) of the lower right tile, since brick_row_height_minus1[3][0] to [3][1] and the number of basic block rows of the tile is 8, they are 3, 2, 3 (= 8 - 3 - 2) in order from the top.

[0078] By such processing, the number of basic block rows in each slice is determined. In this embodiment, since the number of entry_point_offset_minus1 can be derived from other syntax, there is no need to code num_entry_point_offset and include it in the header as in the prior art. Thus, according to this embodiment, the data amount of the bit stream can be reduced.

[0079] In step S310, the control unit 199 determines whether the coding of all the basic blocks in the input image has been completed. As a result of this determination, if the coding of all the basic blocks in the input image has been completed, the process proceeds to step S311. On the other hand, if there are still basic blocks in the input image that have not been coded yet, the process proceeds to step S303, and the subsequent processing is performed on the basic blocks that have not been coded yet.

[0080] In step S311, the in-loop filter unit 109 performs in-loop filter processing on the reproduced image generated in step S306, and outputs the image subjected to the in-loop filter processing.

[0081] As described above, according to this embodiment, it is not necessary to encode information indicating how many pieces of information indicating the head position of the code data of the basic block rows included in the block are encoded and include it in the bit stream, and a bit stream from which the information can be derived can be generated.

[0082] [Second Embodiment] In this embodiment, an image decoding apparatus that decodes the bit stream generated by the image encoding apparatus according to the first embodiment will be described. Note that since requirements common to the first embodiment, such as the configuration of the bit stream, are as described in the first embodiment, the description thereof will be omitted.

[0083] A functional configuration example of the image decoding apparatus according to this embodiment will be described with reference to the block diagram of FIG. 2. The separation decoding unit 202 acquires the bit stream generated by the image encoding apparatus according to the first embodiment. The method of acquiring the bit stream is not limited to a specific acquisition method. For example, the bit stream may be acquired directly or indirectly from the image encoding apparatus via a network such as a LAN or the Internet, or the bit stream stored inside or outside the image decoding apparatus may be acquired. Then, the separation decoding unit 202 separates code data related to decoding processing and coefficients from the acquired bit stream, and sends it to the decoding unit 203. Further, the separation decoding unit 202 decodes the code data of the header of the bit stream. In this embodiment, header information related to the division of the image, such as the size of the tile, block, slice, and basic block, is decoded to generate division information, and the generated division information is output to the image reproduction unit 205. That is, the separation decoding unit 202 performs an operation reverse to that of the integrated encoding unit 111 in FIG. 1.

[0084] The decoding unit 203 decodes the coded data output from the separation decoding unit 202 to reproduce quantization coefficients and prediction information. The inverse quantization / inverse transformation unit 204 performs inverse quantization on the quantization coefficients to generate transformation coefficients, and performs inverse orthogonal transformation on the generated transformation coefficients to reproduce the prediction error.

[0085] The frame memory 206 is a memory for storing the image data of the reproduced picture. The image reproduction unit 205 appropriately refers to the frame memory 206 based on the input prediction information to generate a predicted image. Then, the image reproduction unit 205 generates a reproduced image from the generated predicted image and the prediction error reproduced by the inverse quantization / inverse transformation unit 204. Then, the image reproduction unit 205 specifies and outputs the positions of the tile, block, and slice in the input image of the reproduced image based on the division information input from the separation decoding unit 202.

[0086] The in-loop filter unit 207 performs in-loop filter processing such as a deblocking filter on the reproduced image, similar to the in-loop filter unit 109 described above, and outputs the image subjected to the in-loop filter processing. The control unit 299 controls the operation of the entire image decoding device and controls the operation of each functional unit of the image decoding device described above.

[0087] Next, the decoding process of the bitstream by the image decoding device having the configuration shown in FIG. 2 will be described. Hereinafter, it will be described assuming that the bitstream is input to the image decoding device in units of frames, but it may be possible to input the bitstream of a still image for one frame to the image decoding device. Further, in this embodiment, for the sake of easy explanation, only the intra prediction decoding process will be described, but the present invention is not limited thereto and is also applicable to the inter prediction decoding process.

[0088] The separation and decoding unit 202 separates information related to decoding processing and coded data related to coefficients from the input bit stream and sends them to the decoding unit 203. Also, the separation and decoding unit 202 decodes the coded data of the header of the bit stream. More specifically, the separation and decoding unit 202 decodes the basic block data division information, tile data division information, block data division information, slice data division information 0, basic block row data synchronization information, basic block row data position information, etc. in FIG. 6 to generate division information. Then, the separation and decoding unit 202 outputs the generated division information to the image reproduction unit 205. Also, the separation and decoding unit 202 reproduces the coded data in units of basic blocks of picture data and outputs it to the decoding unit 203.

[0089] The decoding unit 203 decodes the coded data output from the separation and decoding unit 202 to reproduce quantization coefficients and prediction information. The reproduced quantization coefficients are output to the inverse quantization and inverse transformation unit 204, and the reproduced prediction information is output to the image reproduction unit 205.

[0090] The inverse quantization and inverse transformation unit 204 performs inverse quantization on the input quantization coefficients to generate transformation coefficients, and performs inverse orthogonal transformation on the generated transformation coefficients to reproduce prediction errors. The reproduced prediction errors are output to the image reproduction unit 205.

[0091] The image reproduction unit 205 appropriately refers to the frame memory 206 based on the prediction information input from the separation and decoding unit 202 to generate a predicted image. Then, the image reproduction unit 205 generates a reproduced image from the generated predicted image and the prediction errors reproduced by the inverse quantization and inverse transformation unit 204. Then, for the reproduced image, the image reproduction unit 205 specifies, based on the division information input from the separation and decoding unit 202, the shapes of tiles, blocks, slices and their positions in the input image as shown in FIG. 7, for example, and outputs (stores) them in the frame memory 206. The image stored in the frame memory 206 is used for reference during prediction.

[0092] The in-loop filter unit 207 performs in-loop filter processing such as deblocking filtering on the reproduced image read from the frame memory 206, and outputs (stores) the image subjected to the in-loop filter processing to the frame memory 206.

[0093] The control unit 299 outputs the reproduced image stored in the frame memory 206. The output destination of the reproduced image is not limited to a specific output destination. For example, the control unit 299 may output the reproduced image to a display device included in the image decoding apparatus and cause the display device to display the reproduced image. Also, for example, the control unit 299 may transmit the reproduced image to an external device via a network such as a LAN or the Internet.

[0094] Next, the decoding process of the bitstream by the image decoding apparatus according to the present embodiment (the decoding process of the bitstream having the configuration of FIG. 6) will be described with reference to the flowchart of FIG. 4.

[0095] In step S401, the separation decoding unit 202 separates information related to the decoding process and coded data related to coefficients from the input bitstream, and sends it to the decoding unit 203. Also, the separation decoding unit 202 decodes the coded data of the header of the bitstream. More specifically, the separation decoding unit 202 decodes the basic block data division information, tile data division information, block data division information, slice data division information, basic block row data synchronization information, basic block row data position information, etc. in FIG. 6 to generate division information. Then, the separation decoding unit 202 outputs the generated division information to the image reproduction unit 205. Also, the separation decoding unit 202 reproduces the coded data of the basic block unit of the picture data and outputs it to the decoding unit 203.

[0096] In the present embodiment, the division of the input image which is the encoding source of the bitstream is the division shown in FIG. 8. Information related to the input image which is the encoding source of the bitstream and its division can be derived from the division information.

[0097] From pic_width_in_luma_samples included in the image size information, it can be identified that the horizontal size (width) of the input image is 1536 pixels. Also, from pic_height_in_luma_samples included in the image size information, it can be identified that the vertical size (height) of the input image is 1024 pixels.

[0098] Also, since log2_ctu_size_minus2 = 4 in the basic block data partition information, the size of the basic block can be derived as 64×64 pixels from 1<<log2_ctu_size_minus2+2.

[0099] Also, since single_tile_in_pic_flag = 0 in the tile data partition information, it can be identified that the input image is divided into multiple tiles. And since uniform_tile_spacing_flag = 1, it can be identified that each tile has the same size (excluding the edges).

[0100] Also, since tile_cols_width_minus1 = 11 and tile_rows_height_minus1 = 7, it can be identified that each tile is composed of 12×8 basic blocks. That is, it can be identified that each tile is composed of 768×512 pixels. Since the input image is 1536×1024 pixels, it can be seen that the input image is divided into 4 tiles, 2 horizontally and 2 vertically, and encoded.

[0101] Also, since brick_splitting_present_flag = 1 in the brick data partition information, it can be identified that at least one tile in the input image is divided into multiple bricks.

[0102] Also, brick_split_flag[0]=0 and brick_split_flag[1] to [3]=1. Thus, it can be specified that the top-left tile is composed of a single brick, and the other three tiles are composed of multiple bricks.

[0103] Also, uniform_brick_spacing_flag[1] to [2]=1 and uniform_brick_spacing_flag[3]=0. Thus, it can be specified that for the bottom-left tile and the top-right tile, except for the edges, they are composed of bricks of the same size, and for the bottom-right tile, it can be specified that there are bricks of different sizes from the others.

[0104] Also, brick_height_minus1[1]=1 and brick_height_minus1[2]=3. Thus, it can be specified that the number of basic block rows of the bricks in the top-right tile is 2, and the number of basic block rows of the bricks in the bottom-left tile is 4. That is, it can be specified that the top-right tile contains 4 bricks with a vertical size of 128 pixels, and the bottom-left tile contains 2 bricks with a vertical size of 256 pixels. Also, regarding the bottom-right tile, since num_brick_rows_minus1[3] is 2, it can be specified that the tile is composed of 3 bricks. Also, brick_row_height_minus1[3][0]=2, brick_row_height_minus1[3][1]=1, and the number of basic vertical blocks of the bottom-right tile is 8. Thus, it can be specified that the bottom-right tile is composed of a brick with 3 basic vertical blocks, a brick with 2 basic vertical blocks, and a brick with 3 basic vertical blocks from top to bottom.

[0105] Also, since single_brick_per_slice_flag = 1 in the slice data division information 0, it is specified that all slices in the input image are composed of a single brick. In this embodiment, when uniform_brick_spacing_flag[i] = 0, num_brick_rows_minus1[i] indicating (the number of bricks constituting the i-th tile - 1) is included in the brick data division information. However, it is not limited to this.

[0106] For example, when brick_split_flag[i] = 1, it is assumed that the number of bricks constituting the i-th tile is 2 or more, and num_brick_rows_minus2[i] indicating (the number of bricks constituting the tile - 2) may be decoded instead of num_brick_rows_minus1[i]. By doing so, it is possible to decode a bit stream with the number of bits of the syntax indicating the number of bricks constituting the tile reduced.

[0107] Next, determine the coordinates of the upper-left and lower-right boundaries of each block. The coordinates are based on the upper-left corner of the input image as the origin and are represented by the horizontal and vertical positions of the basic blocks. For example, the coordinates of the upper-left boundary of the third basic block from the left and the second basic block from the top are (3, 2), and the coordinates of the lower-right boundary are (4, 3). The coordinates of the upper-left boundary of the block with BID = 0 are (0, 0). Since the number of basic block rows of the block with BID = 0 is 8 and the number of basic blocks in the horizontal direction of all tiles is 12, the coordinates of the lower-right boundary are (12, 8). For each block with BID = 1 to 4 belonging to the upper-right tile, since the number of basic block rows of the block is 2, the coordinates of the upper-left boundary are (12, 0), (12, 2), (12, 4), (12, 6). Similarly, for each block with BID = 1 to 4 belonging to the upper-right tile, the coordinates of the lower-right boundary are (24, 2), (24, 4), (24, 6), (24, 8). For each block with BID = 5 to 6 belonging to the lower-left tile, since the number of basic block rows of the block is 4, the coordinates of the upper-left boundary are (0, 8), (0, 12). Similarly, for each block with BID = 5 to 6 belonging to the lower-left tile, the coordinates of the lower-right boundary are (12, 12), (12, 16). For each block with BID = 7 to 9 belonging to the lower-right tile, since the number of basic block rows of each block is 3, 2, 3 respectively, the coordinates of the upper-left boundary are (12, 8), (12, 11), (12, 13). Similarly, for each block with BID = 7 to 9 belonging to the lower-right tile, the coordinates of the lower-right boundary are (24, 11), (24, 13), (24, 16).

[0108] Also, entropy_coding_sync_enabled_flag = 1 in the basic block row data synchronization information. From this, it can be seen that entry_point_offset_minus1[j - 1] indicating (the size of the coded data of the (j - 1)-th basic block row in the slice - 1) is encoded in the bitstream. The number of entry_point_offset_minus1[] is equal to (the number of basic block rows in the slice to be decoded next - 1). In this embodiment, since the slice is composed of only one block, the number of basic block rows in the slice to be processed is the same as the number of basic block rows in the block, and the correspondence between the slice and the block can be obtained from the value of slice_address. The slice (block) with slice_address being N contains the block with BID = N. Since the number of basic block rows in each block has already been derived, the number of entry_point_offset_minus1[] can be derived from other syntaxes without encoding num_entry_point_offset as in the prior art. And since the start position of the data for each basic block row is known, the decoding process can be performed in parallel for each basic block row.

[0109] In this way, various information such as the input image which is the source of the bitstream encoding and the information related to its division can be derived from the division information decoded by the separation decoding unit 202. The division information derived from the separation decoding unit 202 is sent to the image reproduction unit 205 and is used to specify the position of the processing target in the input image in step S404.

[0110] In step S402, the decoding unit 203 decodes the coded data separated by the separation decoding unit 202 and reproduces the quantization coefficients and prediction information. In step S403, the inverse quantization and inverse transformation unit 204 performs inverse quantization on the input quantization coefficients to generate transformation coefficients, and performs inverse orthogonal transformation on the generated transformation coefficients to reproduce the prediction error.

[0111] In step S404, the image playback unit 205 appropriately refers to the frame memory 206 based on the prediction information input from the separation decoding unit 202 to generate a predicted image. Then, the image playback unit 205 generates a reproduced image from the generated predicted image and the prediction error reproduced by the inverse quantization / inverse transformation unit 204. Then, for the reproduced image, the image playback unit 205 specifies the positions of the tiles and blocks in the input image based on the division information input from the separation decoding unit 202, synthesizes them at those positions, and outputs (stores) them in the frame memory 206.

[0112] In step S405, the control unit 299 determines whether all the basic blocks of the input image have been decoded. As a result of this determination, if all the basic blocks of the input image have been decoded, the process proceeds to step S406. On the other hand, if there are still basic blocks in the input image that have not been decoded yet, the process proceeds to step S402 to perform decoding processing on the basic blocks that have not been decoded yet.

[0113] In step S406, the in-loop filter unit 207 performs in-loop filter processing on the reproduced image read from the frame memory 206, and outputs (stores) the image subjected to the in-loop filter processing in the frame memory 206.

[0114] Thus, according to the present embodiment, it is possible to decode an input image from a "bit stream that does not include information indicating how many pieces of information indicating the head positions of the basic block rows included in a block are encoded" generated by the image encoding apparatus according to the first embodiment.

[0115] Note that the image encoding apparatus according to the first embodiment and the image decoding apparatus according to the second embodiment may be separate apparatuses. Also, the image encoding apparatus according to the first embodiment and the image decoding apparatus according to the second embodiment may be integrated into one apparatus.

[0116] [Third Embodiment] Each functional unit shown in FIG. 1 or FIG. 2 may be implemented in hardware, but a part of it may also be implemented in software. In the latter case, each functional unit except for the frame memory 108 and the frame memory 206 may be implemented in software (computer program). A computer device capable of executing such a computer program is applicable to the above-described image encoding device and image decoding device.

[0117] An example of the hardware configuration of a computer device applicable to the above-described image encoding device and image decoding device will be described with reference to the block diagram of FIG. 5. Note that the hardware configuration shown in FIG. 5 is merely an example of the hardware configuration of a computer device applicable to the above-described image encoding device and image decoding device, and can be appropriately changed / modified.

[0118] The CPU 501 executes various processes using computer programs and data stored in the RAM 502 and the ROM 503. Thereby, the CPU 501 controls the operation of the entire computer device, and executes or controls each process described as being performed by the above-described image encoding device and image decoding device. That is, the CPU 501 can function as each functional unit (excluding the frame memory 108 and the frame memory 206) shown in FIG. 1 and FIG. 2.

[0119] The RAM 502 has an area for storing computer programs and data loaded from the ROM 503 and the external storage device 506, and an area for storing data received from the outside via the I / F 507. The RAM 502 also has a work area used when the CPU 501 executes various processes. In this way, the RAM 502 can appropriately provide various areas. The ROM 503 stores setting data of the computer device, a startup program, and the like.

[0120] The operation unit 504 is a user interface such as a keyboard, a mouse, or a touch panel screen, and various instructions can be input to the CPU 501 by the user's operation.

[0121] The display unit 505 is composed of a liquid crystal screen, a touch panel screen, etc., and can display the processing results by the CPU 501 in the form of images, characters, etc. Note that the display unit 505 may be a device such as a projector that projects images and characters.

[0122] The external storage device 506 is a large-capacity information storage device such as a hard disk drive device. In the external storage device 506, an OS (operating system), and computer programs and data for causing the CPU 501 to execute or control the various processes described above as performed by the above-described image encoding device and image decoding device are stored.

[0123] The computer programs stored in the external storage device 506 include computer programs for causing the CPU 501 to execute or control the functions of each functional unit excluding the frame memory 108 and the frame memory 206 in FIGS. 1 and 2. Further, the data stored in the external storage device 506 includes those described as known information in the above description and various information related to encoding and decoding.

[0124] The computer programs and data stored in the external storage device 506 are appropriately loaded into the RAM 502 according to the control by the CPU 501 and become the processing targets by the CPU 501.

[0125] The frame memory 108 included in the image encoding device in FIG. 1 and the frame memory 206 included in the image encoding device in FIG. 2 can be implemented using memory devices such as the above-described RAM 502 and external storage device 506.

[0126] I / F 507 is an interface for data communication with an external device. For example, when a computer device is applied to an image encoding device, the image encoding device can output the generated bit stream externally via I / F 507. Also, when a computer device is applied to an image decoding device, the image decoding device can receive the bit stream via I / F 507. Further, the image decoding device can transmit externally via I / F 507 the result of decoding the bit stream. The CPU 501, RAM 502, ROM 503, operation unit 504, display unit 505, external storage device 506, and I / F 507 are all connected to the bus 508.

[0127] Note that the specific numerical values used in the above description are for the purpose of specific explanation, and it is not intended that each of the above embodiments be limited to these numerical values. Also, some or all of the above-described embodiments may be combined as appropriate. Further, some or all of the above-described embodiments may be selectively used.

[0128] (Other Embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. Also, it can be realized by a circuit (for example, ASIC) that realizes one or more functions.

[0129] The invention is not limited to the above embodiments, and various changes and modifications are possible without departing from the spirit and scope of the invention. Accordingly, claims are attached to disclose the scope of the invention.

Explanation of Reference Numerals

[0130] 102: Image segmentation unit 103: Block segmentation unit 104: Prediction unit 105: Transformation and quantization unit 106: Inverse quantization and inverse transformation unit 107: Image reproduction unit 108: Frame memory 109: In-loop filter unit 110: Encoding unit 111: Integrated encoding unit

Claims

1. An image decoding apparatus for decoding an image from a bitstream obtained by encoding an image including a plurality of rectangular regions, wherein each of the plurality of rectangular regions includes one or more block rows composed of a plurality of blocks, and the image decoding apparatus includes: decoding means for decoding, from the bitstream, first information for specifying the number of vertical blocks of a certain rectangular region among the plurality of rectangular regions in the image, decoding a first flag regarding activation of parallel processing from the bitstream, decoding a second flag from a picture parameter set of the bitstream, and decoding second information indicating an ID of the certain rectangular region corresponding to a target slice in the image from a slice header of the bitstream; when the value of the second flag is 1, the second flag indicates that each rectangular region consists of one slice; even when the second flag indicates that each rectangular region consists of one slice, the second information can be decoded from the slice header; the image decoding apparatus further includes: specific means for specifying, based on the first information and the second information, the number of syntax elements used to specify the start position of the coded data of the block row for the target slice in a state where the value of the first flag is 1, the second flag indicates that each rectangular region consists of one slice, and the target slice is rectangular; the size of each of the plurality of blocks constituting the block row is determined from third information decoded from a sequence parameter set of the bitstream, and the number of the syntax elements specified by the specific means is included in the slice header of the bitstream; An image decoding apparatus characterized by the above.

2. The image decoding apparatus according to claim 1, wherein the first flag is an entropy_coding_sync_enabled_flag.

3. The image decoding apparatus according to claim 1, wherein when the value of the first flag is 1, the leftmost block uses probability information of a block at a predetermined position in an adjacent block row above. Claim 4. The image decoding apparatus according to claim 1, wherein the size of each of the plurality of blocks is determined by arithmetically shifting 1 to the left by the result of the sum of a predetermined value and the value of the third information. Claim 5. The image decoding apparatus according to claim 1, wherein each of the plurality of blocks constituting the block row corresponds to a CTU (Coding Tree Unit). Claim 6. The image decoding apparatus according to claim 1, wherein the syntax element is entry_point_offset_minus1. Claim 7. An image decoding method for decoding an image from a bitstream obtained by encoding an image including a plurality of rectangular regions, wherein each of the plurality of rectangular regions includes one or more block rows composed of a plurality of blocks, and the image decoding method includes: a decoding step of decoding, from the bitstream, first information for specifying the number of blocks in the vertical direction of a certain rectangular region among the plurality of rectangular regions in the image, decoding, from the bitstream, a first flag regarding enabling parallel processing, decoding a second flag from a picture parameter set of the bitstream, and decoding, from a slice header of the bitstream, second information indicating an ID of the certain rectangular region corresponding to a target slice in the image; when the value of the second flag is 1, the second flag indicates that each rectangular region consists of one slice; even when the second flag indicates that each rectangular region consists of one slice, the second information can be decoded from the slice header; the image decoding method further includes: a specifying step of specifying, based on the first information and the second information, the number of syntax elements used to specify the head position of the coded data of the block row for the target slice in a state where the value of the first flag is 1, the second flag indicates that each rectangular region consists of one slice, and the target slice is rectangular; the size of each of the plurality of blocks constituting the block row is determined from third information decoded from a sequence parameter set of the bitstream, and the specified number of the syntax elements is included in the slice header of the bitstream characterized in that it is an image decoding method. Claim 8. A computer program for causing a computer to execute the image decoding method according to claim 7.

Citation Information

Patent Citations

  • Image encoder, image encoding method and program, image decoder, and image decoding method and program

    JP2014011638A

  • Video coding device and video decoding device

    WO2019078169A1