Image encoding device, image decoding device, image encoding method, and image decoding method

By determining the rectangularity and grid size of the sub-picture in the VVC encoding method, encoding and decoding is performed only when the maximum number is not less than 2, the problem of syntax redundancy of the sub-picture is solved, and the code quantity is reduced and the encoding efficiency is improved.

CN120263975APending Publication Date: 2025-07-04CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510485396.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-09-11
Filing Date
2020-08-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the VVC encoding method, there is a redundant part in the syntax of the sub-picture, resulting in an increase in the code volume.

Method used

Through the image encoding device and the decoding device, the determining component and the acquisition component respectively, determine the rectangularity, grid size and number of the sub-picture, and only encode and decode when the maximum number is not less than 2, reducing redundant information.

Benefits of technology

The code amount of generated bitstreams is effectively reduced and the encoding efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263975A_ABST
    Figure CN120263975A_ABST
Patent Text Reader

Abstract

The invention relates to an image encoding apparatus, an image decoding apparatus, an image encoding method, and an image decoding method. The image encoding apparatus includes an encoding means for dividing an image into a plurality of sub-pictures and encoding each of the sub-pictures such that each of the sub-pictures can be independently decoded, the image encoding apparatus including: a first decision means for deciding whether each of the sub-pictures forming the image is defined as only one rectangle; a second determination means for determining the number of basic pixels, which is the vertical / horizontal size of the mesh forming the sub-picture; an assigning means for assigning the number to each grid divided by the number of basic pixels; and a third determination means for determining the maximum number of sub-pictures, which is the number of numbers allocated by the allocation means, and which are grid sets to which the same number is allocated, and encoding a value obtained by subtracting 2 from the maximum number determined by the third determination means.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] (This application is a divisional application of an application with an application date of August 21, 2020, an application number of 2020800638779, and an invention title of "Image Encoding Device and Image Decoding Device and Their Methods and Storage Medium".) Technical Field

[0002] The present invention relates to an image encoding device, an image encoding method and program, an image decoding device, an image decoding method and program, and image code data. More specifically, it relates to an image encoding method / decoding method capable of dividing a picture into rectangles and extracting independent code data. Background Art

[0003] As an encoding method for compression recording of moving images, the HEVC (High Efficiency Video Coding) encoding method (hereinafter referred to as HEVC) is known. In HEVC, in order to improve the encoding efficiency, basic blocks larger than a conventional macroblock (16×16 pixels) are adopted. The basic block with a large size is called a CTU (Coding Tree Unit), and its size is up to 64×64 pixels at maximum. The CTU is further divided into sub-blocks as units for prediction or transformation.

[0004] In addition, in HEVC, a picture can be divided into multiple tiles or slices and encoded. Tiles or slices have small data dependencies, and encoding / decoding processing can be performed in parallel. A great advantage of tile or slice division is that processing can be performed in parallel by a multi-core CPU or the like to shorten the processing time. Patent Document 1 discloses a technique related to tiles and slices.

[0005] In recent years, international standardization activities for a more efficient encoding method as a successor to HEVC have been started. JVET (Joint Video Experts Team) has been established between ISO / IEC and ITU-T, and the VVC (Versatile Video Coding) encoding method (hereinafter referred to as VVC) has been standardized. In VVC, there are sub-pictures configured to include one or more than one slice and form a rectangle. A picture can be divided into one or more than one sub-picture, and each sub-picture can be processed as independent code data.

[0006] Citation List

[0007] Patent Document

[0008] Patent Document 1: Japanese Unexamined Patent Application Publication No. 2014-11638 Summary of the Invention

[0009] Problems to be Solved by the Invention

[0010] Regarding sub - pictures of VVC, values representing the maximum number of sub - pictures, basic pixel counts such as the vertical / horizontal sizes of a grid defining the boundary positions of each sub - picture, the ID of each sub - picture, control flags corresponding to each sub - picture, etc. are defined as syntax.

[0011] There are still many redundant parts in the syntax regarding sub - pictures, resulting in an increase in the code amount.

[0012] Therefore, the present invention is made to solve the above problems and provides a technique for reducing the code amount of the generated bitstream by eliminating the redundancy of the syntax regarding sub - pictures.

[0013] Solution to the problem

[0014] To solve the above problems, an image encoding device according to the present invention has the following configuration. That is, the image encoding device is an image encoding device including an encoding component for dividing an image into a plurality of sub - pictures and encoding each sub - picture so that each sub - picture can be independently decoded, and is characterized by including: a first determination component for determining whether each sub - picture forming the image is defined as only one rectangle; a second determination component for determining the basic pixel count, which is the vertical / horizontal size of the grid forming the sub - picture; an assignment component for assigning numbers to each grid divided by the basic pixel count; and a third determination component for determining the maximum number of the sub - pictures, where the maximum number of the sub - pictures is the number of the numbers assigned by the assignment component, wherein the sub - pictures are sets of grids assigned the same number, and encodes the value obtained by subtracting 2 from the maximum number determined by the third determination component.

[0015] In addition, an image decoding device according to the present invention has the following configuration. That is, the image decoding device is an image decoding device including an image decoding component capable of decoding a bitstream including code data obtained by encoding each of a plurality of sub - pictures, and is characterized by including: a first acquisition component for acquiring from the bitstream information indicating whether each sub - picture forming the image includes only one rectangle; a second acquisition component for acquiring from the bitstream the maximum number of the sub - pictures; and a third acquisition component for acquiring from the bitstream the basic pixel count, which is the vertical / horizontal size of the grid forming the sub - picture, wherein the value obtained by adding 2 to the value acquired by the second acquisition component is set as the maximum number of the sub - pictures, the number of vertical / horizontal grids forming the image is acquired according to the vertical / horizontal size of the grid acquired by the third acquisition component, and information indicating which sub - picture each grid belongs to is acquired.

[0016] In addition, the image encoding device according to the present invention has the following configuration. That is to say, the image encoding device is an image encoding device including an encoding component for dividing an image into a plurality of sub-pictures and encoding each sub-picture so that each sub-picture can be independently decoded, characterized in that it includes: a first determination component for determining whether each sub-picture forming the image is defined as only one rectangle; a second determination component for determining a basic number of pixels, which is the vertical / horizontal size of the grid forming the sub-picture; an allocation component for allocating numbers to each grid divided by the basic number of pixels; and a third determination component for determining a maximum number of the sub-pictures, which is the number of the numbers allocated by the allocation component, wherein the sub-picture is a set of grids allocated the same number, and the number and the basic number of pixels are encoded only when the maximum number determined by the third determination component is not less than 2.

[0017] In addition, the image decoding device according to the present invention has the following configuration. That is to say, the image decoding device is an image decoding device including an image decoding component capable of decoding a bitstream including code data obtained by encoding each of a plurality of sub-pictures, characterized in that it includes: a first acquisition component for acquiring from the bitstream information indicating whether each sub-picture forming the image includes only one rectangle; a second acquisition component for acquiring from the bitstream the maximum number of the sub-pictures; and a third acquisition component for acquiring from the bitstream a basic number of pixels, which is the vertical / horizontal size of the grid forming the sub-picture, wherein the sub-picture is a set of grids allocated the same number, and the number and the basic number of pixels are acquired only when the maximum number determined by the second acquisition component is not less than 2.

[0018] Effects of the Invention

[0019] According to the present invention, the syntax regarding the sub-pictures forming an image can be effectively encoded and the encoding can be efficiently improved.

[0020] From the following description in conjunction with the drawings, other features and advantages of the present invention will become apparent. Note that throughout the drawings, the same reference numerals denote the same or similar components. Description of the Drawings

[0021] The drawings included in the specification and constituting a part of the specification illustrate embodiments of the present invention and are used together with the specification to explain the principles of the present invention.

[0022] Figure 1 is a block diagram showing the configuration of the image encoding device;

[0023] Figure 2 is a block diagram showing the configuration of an image decoding device;

[0024] Figure 3 is a flowchart showing the image encoding process in an image encoding device;

[0025] Figure 4 is a flowchart showing the image decoding process in an image decoding device;

[0026] Figure 5 is a block diagram showing an example of the hardware configuration;

[0027] Figure 6 is a diagram showing the bitstream configuration;

[0028] Figure 7 is a diagram showing an example of image segmentation;

[0029] Figure 8 is a diagram showing an example of image segmentation;

[0030] Figure 9 is a diagram showing the bitstream configuration;

[0031] Figure 10 is a diagram showing an example of image segmentation; and

[0032] Figure 11 is a diagram showing the bitstream configuration.

[0033] List of Reference Numerals

[0034] 101, 112, 201, 208: Terminals

[0035] 102: Image segmentation unit

[0036] 103: Block segmentation unit

[0037] 104: Prediction unit

[0038] 105: Transform / quantization unit

[0039] 106, 204: Inverse quantization / inverse transform unit

[0040] 107, 205: Image reproduction unit

[0041] 108, 206: Frame memory

[0042] 109, 207: In-loop filter unit

[0043] 110: Encoding unit

[0044] 111: Integrated encoding unit

[0045] 202: Separation decoding unit

[0046] 203: Decoding unit Detailed implementation manners

[0047] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments are not intended to limit the scope of the claimed invention. In the embodiments, multiple features are described, but the invention that requires all of these features is not limited, and multiple of these features can be appropriately combined. In addition, in the accompanying drawings, the same reference numerals are given to the same or similar configurations, and redundant descriptions thereof are omitted.

[0048] (First Embodiment)

[0049] Now, embodiments of the present invention will be described with reference to the accompanying drawings. Figure 1 is a block diagram showing an image coding device according to this embodiment. Refer to Figure 1 , reference numeral 101 represents a terminal configured to input image data.

[0050] Reference numeral 102 represents an image segmentation unit that divides the input image into one or more slice rows or one or more slice columns. Each slice is a set of consecutive basic blocks that cover a rectangular area in the image. The image segmentation unit 102 also divides a slice into one or more bricks. Each brick is a rectangle formed by one or more rows of basic blocks, and these rows of basic blocks are rows of basic blocks in the slice. The image segmentation unit 102 divides the image into strips, and each strip is formed by one or more slices in the image or one or more bricks in a slice. A strip is the basic unit of coding, and header information such as information indicating the strip type is added to each strip.

[0051] In this embodiment, as Figure 8 shown, the picture is divided into two sub - pictures. In addition, for simplicity of description, the sub - pictures, strips, slices, and bricks have the same vertical / horizontal size and the same position. However, the present invention is not limited thereto.

[0052] Reference numeral 103 represents a block segmentation unit that divides the row - image of basic blocks output from the image segmentation unit 102 into multiple basic blocks and outputs the image of each basic block to the subsequent stage.

[0053] Reference numeral 104 denotes a prediction unit that determines sub-block division of the image data of each basic block and performs intra prediction as intra-frame prediction or inter prediction as inter-frame prediction for each sub-block, thereby generating predicted image data. Intra prediction across blocks or motion vector prediction is not performed. Further, the prediction unit 104 calculates a prediction error based on the input image data and the predicted image data, and outputs the prediction error. Additionally, the prediction unit 104 outputs information required for prediction, for example, information such as sub-block division, prediction mode, and motion vector, as well as the prediction error. The information required for prediction will hereinafter be referred to as prediction information.

[0054] Reference numeral 105 denotes a transform / quantization unit that orthogonally transforms the prediction error based on sub-blocks to obtain transform coefficients, and further quantizes the transform coefficients to obtain quantization coefficients. Reference numeral 106 denotes an inverse quantization / inverse transform unit that inverse-quantizes the quantization coefficients output from the transform / quantization unit 105 to reproduce the transform coefficients, and also inverse-orthogonally transforms the transform coefficients to reproduce the prediction error. Reference numeral 108 denotes a frame memory that stores the reproduced image data.

[0055] Reference numeral 107 denotes an image reproduction unit. The image reproduction unit 107 generates predicted image data by appropriately referring to the frame memory 108 based on the prediction information output from the prediction unit 104, generates reproduced image data based on the predicted image data and the input prediction error, and outputs the reproduced image data. Reference numeral 109 denotes an in-loop filter unit. The in-loop filter unit 109 performs in-loop filtering processing such as deblocking filtering processing or sample adaptive offset on the reproduced image, and outputs the image that has undergone the filtering processing.

[0056] Reference numeral 110 denotes an encoding unit. The encoding unit 110 generates code data by encoding the quantization coefficients output from the transform / quantization unit 105 and the prediction information output from the prediction unit 104, and outputs the code data. Reference numeral 111 denotes an integrated encoding unit. The integrated encoding unit 111 receives segmentation information from the image segmentation unit 102 and generates header code data. The integrated encoding unit 111 also forms a bitstream by combining the header code data with the code data output from the encoding unit 110, and outputs the bitstream. Reference numeral 112 denotes a terminal that outputs the bitstream generated by the integrated encoding unit 111 to the outside.

[0057] The image encoding operation of the image encoding device according to the first embodiment will be described below. In this embodiment, frame input moving image data is used. However, still image data corresponding to one frame may also be input. In addition, in this embodiment, for ease of description, only the intra prediction encoding process will be described. However, the present invention is not limited thereto, and it may also be applied to the inter prediction encoding process. In addition, in this embodiment, for the sake of description, it is assumed that the block segmentation unit 103 divides the image into basic blocks each having a size of 64×64 pixels. However, the present invention is not limited thereto.

[0058] In this embodiment, first, the image segmentation unit 102 divides one frame of image data input from the terminal 101 into sub-pictures, as Figure 8 shown. In this embodiment, the case where one image data is divided into two sub-pictures and formed by 18 grids will be described. The vertical / horizontal size of one grid is 256×256 pixels, and the vertical / horizontal size of the sub-picture is 768×768 pixels. For each grid of the left sub-picture, 0 is set as the index value. For each grid of the right sub-picture, 1 is set as the index value. As described above, in this embodiment, one image data is divided into a plurality of grids and divided into a plurality of sub-pictures based on the plurality of grids. In addition, each grid is assigned an index value as a number, and the grids forming one sub-picture are assigned the same index value.

[0059] In this embodiment, each sub-picture is formed by one strip and one slice / one block. Information about the size of the sub-picture or the strip is sent to the integrated encoding unit 111 as segmentation information. In addition, each block is divided into basic block row images, and sent to the block segmentation unit 103, where each basic block row image is the image data of each basic block row.

[0060] The block segmentation unit 103 divides the input basic block row image into a plurality of basic blocks, and outputs the image of each basic block to the prediction unit 104. In this embodiment, the image of each basic block with a size of 64×64 pixels is output.

[0061] The prediction unit 104 performs prediction processing on the image data of each basic block input from the block segmentation unit 103. More specifically, the sub-block segmentation for dividing the basic block into finer sub-blocks is determined, and the intra prediction mode such as horizontal prediction or vertical prediction is determined based on the sub-blocks.

[0062] Predicted image data is generated based on the determined intra prediction mode and the encoded pixels. In addition, a prediction error is generated according to the input image data and the predicted image data, and output to the transform / quantization unit 105. In addition, information about the sub-block segmentation or the intra prediction mode is output to the encoding unit 110 and the image reproduction unit 107 as prediction information.

[0063] The transform / quantization unit 105 performs orthogonal transform / quantization on the input prediction error to generate quantization coefficients. First, an orthogonal transform process corresponding to the size of the sub-block is performed to generate orthogonal transform coefficients. Next, the orthogonal transform coefficients are quantized to generate quantization coefficients. The generated quantization coefficients are output to the encoding unit 110 and the inverse quantization / inverse transform unit 106.

[0064] The inverse quantization / inverse transform unit 106 inverse-quantizes the input quantization coefficients to reproduce the transform coefficients, and further inverse-orthogonally transforms the reproduced transform coefficients to reproduce the prediction error. The reproduced prediction error is output to the image reproduction unit 107.

[0065] The image reproduction unit 107 reproduces a predicted image by appropriately referring to the frame memory 108 based on the prediction information input from the prediction unit 104. Then, image data is reproduced according to the reproduced predicted image and the reproduced prediction error input from the inverse quantization / inverse transform unit 106. The image data is input and stored in the frame memory 108.

[0066] The in-loop filter unit 109 reads out the reproduced image from the frame memory 108 and performs in-loop filtering processes such as deblocking filtering. The image that has undergone the filtering process is input to the frame memory 108 again and stored again.

[0067] The encoding unit 110 performs entropy encoding on the quantization coefficients generated by the transform / quantization unit 105 and the prediction information input from the prediction unit 104 based on blocks, thereby generating code data. The method of entropy encoding is not particularly specified, and Golomb coding, arithmetic coding, Huffman coding, etc. can be used. The generated code data is output to the integrated encoding unit 111.

[0068] The integrated encoding unit 111 receives segmentation information from the image segmentation unit 102, generates encoded data for the header, and multiplexes the code data and the like input from the encoding unit 110, thereby forming a bitstream. Finally, the bitstream is output from the terminal 112 to the outside.

[0069] The format of the encoded data using VVC encoded by the image encoding device according to this embodiment is as Figure 6 shown. In the Figure 6 encoded data shown, first, there is a sequence parameter set as header information including information on sequence encoding. Subsequently, there are a picture parameter set as header information including information on picture encoding, a slice header as header information including information on the encoding of each slice, and the encoded data of each block.

[0070] In the sequence parameter set, pic_width_in_luma_samples and pic_height_in_luma_samples exist as image size information. These represent the number of pixels in the horizontal direction and the number of pixels in the vertical direction with respect to the image luminance. In this embodiment, since Figure 8 the image shown in is encoded, pic_width_in_luma_samples is 1536, and pic_height_in_luma_samples is 768. Additionally, as basic block data segmentation information, there is log2_ctu_size_minus2 which represents the size of the basic block. The number of vertical and horizontal pixels of the basic block is represented by 1<<(log2_ctu_size_minus2 + 2). In this embodiment, since the basic block has 64×64 pixels, the value of log2_ctu_size_minus2 is 4.

[0071] Furthermore, as sub - picture information, there is subpics_present_flag which is information indicating whether sub - picture segmentation exists. If subpics_present_flag is 1, this indicates that the image is segmented into one or more sub - pictures. If subpics_present_flag is 0, this indicates that the image is not segmented.

[0072] If subpics_present_flag is 1, the following information is also encoded. That is, the information encoded at this time includes, for example, max_subpics_minus2 which represents the maximum number of sub - pictures. Additionally, the information encoded at this time includes information such as subpic_grid_col_width_minus1 and subpic_grid_row_height_minus1 which represent the basic number of pixels of the grid. Furthermore, the information encoded at this time includes information such as subpic_grid_idx[i][j] which is an index value. Furthermore, the information encoded at this time includes information such as subpic_treated_as_pic_flag[i] and loop_filter_across_subpic_enabled_flag[i] regarding filtering.

[0073] Traditionally, the maximum number of sub - pictures is expressed as "the maximum number - 1". In this case, the maximum number of sub - pictures is 1, that is, there is one sub - picture, and the design is such that the code amount is minimized in the state where the image is not actually divided. However, in this embodiment, the maximum number of sub - pictures is expressed as "the maximum number - 2". If the image is divided into two sub - pictures, the value of max_subpics_minus2 is 0, and it can be represented with the minimum code amount.

[0074] The picture parameter set includes slice data information, block data information, strip data information, and basic block data information. The strip header includes strip data information, etc., and is followed by the encoded data of each block.

[0075] Figure 3 is a flowchart showing the encoding process in the image encoding device according to the first embodiment. First, in step S301, the image segmentation unit 102 divides the image into slices, blocks, and strips as described above, and sends the segmentation information to the integrated encoding unit 111. In step S309, the integrated encoding unit 111 converts the segmentation information into header information and encodes it into a bitstream. The image segmentation unit 102 also divides the image into sub - pictures and sends these sub - pictures to the block segmentation unit 103.

[0076] In step S302, the block segmentation unit 103 divides the basic block row image into basic blocks. In step S303, the prediction unit 104 performs prediction processing on the image data of each basic block generated in step S302, thereby generating sub - block segmentation information, prediction information such as intra - prediction mode, and prediction image data. The prediction unit 104 also calculates the prediction error based on the input image data and the prediction image data.

[0077] In step S304, the transform / quantization unit 105 orthogonally transforms the prediction error calculated in step S303 to generate transform coefficients, and performs quantization to generate quantization coefficients.

[0078] In step S305, the inverse quantization / inverse transform unit 106 inverse - quantizes the quantization coefficients generated in step S304, thereby reproducing the transform coefficients. The inverse quantization / inverse transform unit 106 also performs inverse orthogonal transformation on the transform coefficients, thereby reproducing the prediction error.

[0079] In step S306, the image reproduction unit 107 reproduces the prediction image based on the prediction information generated in step S303. The image reproduction unit 107 also reproduces the image data based on the reproduced prediction image and the prediction error generated in step S305.

[0080] In step S307, the encoding unit 110 encodes the prediction information generated in step S303 and the quantization coefficients generated in step S304 to generate code data.

[0081] In step S308, the image encoding device determines whether the encoding of all the basic blocks in the strip is completed. If the encoding is completed, the process proceeds to step S309. Otherwise, the process returns to step S302 to process the next basic block.

[0082] In step S309, the integration encoding unit 111 generates header information based on the segmentation information sent from the image segmentation unit 102 and performs encoding.

[0083] Based on Figure 8 The detailed example in this embodiment will be described based on the image segmentation shown. In this embodiment, since the image is segmented into two sub - pictures, the subpics_present_flag is 1. Next, for max_subpics_minus2 representing the maximum number of sub - pictures, the value of "the maximum number - 2" is set, that is, "0" in this embodiment. If this information is encoded by Golomb coding, the 1 - bit data "0" representing "0" is encoded. Compared with the case of encoding the 3 - bit data "010" representing the value "1", this can improve the encoding efficiency.

[0084] subpic_grid_col_width_minus1 is 63, and subpic_grid_row_height_minus1 is also 63. These two values are defined in units of four pixels, and more specifically, these are the values obtained by subtracting 1 from the value obtained by dividing the basic number of pixels of the grid by 4.

[0085] There is subpic_grid_idx[i][j] which is the index value of the sub - picture. Here, [i][j] represents the position of the grid, where i represents the row direction and j represents the column direction. In this embodiment, the range of the row direction is from 0 to 2, and the range of the column direction is from 0 to 5, and [0][0], [0][1], [0][2], [1][0], [1][1], [1][2], [2][0], [2][1] and [2][2] are 0. Additionally, [0][3], [0][4], [0][5], [1][3], [1][4], [1][5], [2][3], [2][4] and [2][5] are 1. The code data of the strip forming the sub - picture is combined based on these sub - picture information to create the code data of the sub - picture.

[0086] In step S310, the image encoding device determines whether the encoding of all basic blocks in the frame has ended. If the encoding has ended, the process proceeds to step S311. Otherwise, the process returns to step S302 to process the next basic block.

[0087] In step S311, the in-loop filter unit 109 performs in-loop filtering processing on the image data reproduced in step S306 to generate an image that has undergone filtering processing, and the process ends.

[0088] With the above configuration and operations, particularly in step S309, if an image is segmented into sub-pictures, the information representing the maximum number of sub-pictures is encoded by "maximum number - 2", thereby effectively encoding the syntax regarding sub-pictures.

[0089] Note that in this embodiment, a sub-picture is formed by a stripe, a slice, or a block. However, the present invention is not limited thereto. Figure 7 An example in which a stripe is segmented into multiple slices is shown. More specifically, the stripes with index 0 and index 3 are each formed by four slices, the stripes with index 1 and index 2 are each formed by two slices, and the stripe with index 4 is formed by six slices. This embodiment can also be applied to this sub-picture configuration. Note that in this case, a slice and a block have the same basic pixel size.

[0090] Note that in this embodiment, the value obtained by subtracting 2 from the maximum number of sub-pictures is encoded. However, the present invention is not limited thereto. Instead, max_subpics_minus1 obtained by subtracting 1 from the maximum number of sub-pictures can be encoded, and when the value is 1 or greater, that is, only when the maximum number of sub-pictures is 2 or greater, the following parameters can be encoded. That is, parameters such as subpic_grid_col_width_minus1, subpic_grid_row_height_minus1, and subpic_grid_idx[i][j] can be encoded. That is, if the maximum number is 1, the parameters can be omitted so that these parameters are not encoded.

[0091] The format of the encoded data at this time is as Figure 9 shown. In this case, when the maximum number of sub-pictures is 1, the amount of information of subpic_grid_col_width_minus1, subpic_grid_row_height_minus1, and subpic_grid_idx[0][0] can be reduced. This enables effective encoding.

[0092] Note that in the above description, when the value of max_subpics_minus1 is 1, i.e., only when the maximum number of sub-pictures is 2 or greater, the following parameters are encoded. That is, when the maximum number of sub-pictures is 2 or greater, the parameters subpic_grid_col_width_minus1, subpic_grid_row_height_minus1, and subpic_grid_idx[i][j] are encoded. In addition, only when the maximum number of sub-pictures is 2 or greater, the following two values can be encoded. That is, only when the value of max_subpics_minus1 is 1 or greater, the subpic_treated_as_flag and loop_filter_across_subpic_enabled_flag can be encoded. This enables further reduction of the amount of information and efficient encoding.

[0093] In addition, in this embodiment, an example of dividing an image into two sub-pictures has been described. Figure 10 Examples of dividing an image into, for example, three sub-pictures are shown. The value that can be maximally represented by Ceil(log2(max_subpics_minus1 + 1)) bits can be set as subpic_grid_idx. When using a value of max_subpics_minus1 or less, effective encoding is possible. More specifically, since the image is divided into three parts, the maximum number of sub-pictures is 3. Therefore, max_subpics_minus1 is 2 obtained by subtracting 1 from 3. The number of bits at this time is 2 as a result of the calculation of Ceil(log2(2 + 1)), and a value that can be represented by 2 bits can be set. More specifically, four types of values 0, 1, 2, and 3 can be set. When the maximum value that can be set is limited to max_subpics_minus1, as Figure 10 shown, these values are limited to 0, 1, and 2, and the syntax regarding sub-pictures can be correctly encoded.

[0094] In addition, in this embodiment, a grid is defined based on four pixels. However, a grid can be defined based on CTU rows / columns. If the basic number of pixels of a CTU is 64×64, the value of subpic_grid_col_width_minus1 is 3, and the value of subpic_grid_row_height_minus1 is also 3. The information of the encoding object is small, and the amount of code can be reduced.

[0095] Note that, in this embodiment, subpic_grid_idx, which is the ID representing the position of the rectangular area of the sub-picture, is encoded in the sequence parameter set. However, the present invention is not limited to this. Figure 11 FIG. shows the configuration of the bitstream, where subpic_idx, which is the ID of the corresponding sub-picture, is changed to be encoded in the slice header that is the header of each slice. When such a configuration is adopted, the number of information representing the position of the rectangular area of the sub-picture can be reduced from the number of grids to the number of slices. In Figure 8 the example shown, the number of grids is 18, and the number of sub-pictures = the number of slices. Therefore, since the number of slices is 2, the number of information of the encoding object can be reduced from 18 to 2.

[0096] In addition, subpic_idx is arranged in the slice header. However, the present invention is not limited to this, and subpic_idx can be arranged in the block code data. This can reduce the code amount as in the case where subpic_idx is stored in the slice header.

[0097] (Second Embodiment)

[0098] Figure 2 FIG. is a block diagram showing the configuration of an image decoding apparatus according to a second embodiment of the present invention. The image decoding apparatus according to this embodiment can decode a bitstream including code data generated by encoding each of a plurality of sub-pictures. In this embodiment, decoding the encoded data generated in the first embodiment will be described below as an example.

[0099] Reference numeral 201 denotes a terminal for inputting the encoded bitstream. Reference numeral 202 denotes a separation decoding unit that separates code data related to information and coefficients regarding decoding processing from the bitstream and sends these to the decoding unit 203. In addition, the separation decoding unit 202 decodes the code data present in the header portion of the bitstream. In this embodiment, the segmentation information is generated by decoding header information regarding image segmentation (such as the sizes of slices, blocks, strips, and basic blocks, etc.) and is output to the image reproduction unit 205. The separation decoding unit 202 performs an operation opposite to that of Figure 1 the operation of the integration encoding unit 111 shown in FIG.

[0100] Reference numeral 203 denotes a decoding unit that decodes the code data output from the separation decoding unit 202 to reproduce quantization coefficients and prediction information. Reference numeral 204 denotes an inverse quantization / inverse transformation unit that inverse-quantizes the quantization coefficients to obtain transform coefficients and performs an inverse orthogonal transformation on the transform coefficients to reproduce a prediction error. Reference numeral 206 denotes a frame memory 206 that stores the image data of the reproduced picture. Reference numeral 205 denotes an image reproduction unit. The image reproduction unit 205 generates predicted image data by appropriately referring to the frame memory 206 based on the input prediction information. Then, the image reproduction unit 205 generates reproduced image data based on the predicted image data and the prediction error reproduced by the inverse quantization / inverse transformation unit 204. The positions of slices, blocks, and stripes in the frame are specified based on the segmentation information input from the separation decoding unit 202, and the generated reproduced image data is output.

[0101] Reference numeral 207 denotes an in-loop filter unit. Similar to the above-described in-loop filter unit 109 shown in Figure 1 the in-loop filter unit 207 performs in-loop filtering processing such as deblocking filtering processing on the reproduced image and outputs the image that has undergone the in-loop filtering processing. Reference numeral 208 denotes a terminal that outputs the reproduced image data to the outside.

[0102] The image decoding operation of the image decoding device according to this embodiment will be described below. In this embodiment, a bitstream generated in the first embodiment based on a frame input is used. However, a still image bitstream corresponding to one frame may be input. Further, in this embodiment, for ease of description, only the intra prediction decoding process will be described. However, the present invention is not limited thereto, and it can also be applied to the inter prediction decoding process.

[0103] In Figure 2 the bitstream corresponding to one frame input from the terminal 201 is input to the separation decoding unit 202. The separation decoding unit 202 separates the code data related to the information and coefficients regarding the decoding process from the bitstream and decodes the code data present in the header portion of the bitstream. More specifically, segmentation information is generated by decoding the basic block data segmentation information, slice data segmentation information, block data segmentation information, stripe data segmentation information, basic block row data synchronization information, and basic block row data position information shown in Figure 6 and is sent to the image reproduction unit 205. Next, the code data of each basic block of the picture data is reproduced and output to the decoding unit 203.

[0104] The decoding unit 203 decodes the code data to reproduce quantization coefficients and prediction information. The reproduced quantization coefficients are output to the inverse quantization / inverse transformation unit 204, and the reproduced prediction information is output to the image reproduction unit 205.

[0105] The inverse quantization / inverse transformation unit 204 inverse quantizes the input quantized coefficients to generate orthogonal transform coefficients, and performs an inverse orthogonal transformation to reproduce the prediction error. The reproduced prediction error is output to the image reproduction unit 205.

[0106] The image reproduction unit 205 reproduces a predicted image by appropriately referring to the frame memory 206 based on the prediction information input from the decoding unit 203. Image data is reproduced based on the predicted image and the prediction error input from the inverse quantization / inverse transformation unit 204. For example, as Figure 8 shown, the shape and position of slices, strips, and blocks in a frame are specified based on the segmentation information input from the separation decoding unit 202, and the generated reproduced image data is input and stored in the frame memory 206. The stored image data is used as a reference in prediction.

[0107] Similar to Figure 1 the in-loop filter unit 109 shown, the in-loop filter unit 207 reads out the reproduced image from the frame memory 206, and performs in-loop filtering processing such as deblocking filtering processing. The image that has undergone the filtering processing is input to the frame memory 206 again. The reproduced image stored in the frame memory 206 is finally output from the terminal 208 to the outside.

[0108] Figure 4 is a flowchart showing the image decoding process in the image decoding device according to the second embodiment. First, in step S401, the separation decoding unit 202 separates the code data related to the information and coefficients regarding the decoding process from the bitstream, and decodes the code data in the header part. The separation decoding unit 202 decodes Figure 6 the slice data information, block data information, strip data information, etc. shown, generates the information for decoding, and sends it to the image reproduction unit 205. The segmentation of the image stored in the bitstream in this embodiment is as Figure 8 shown.

[0109] From the values of pic_width_in_luma_samples and pic_height_in_luma_samples of the image size information, it is obtained that the image has 1536×768 pixels.

[0110] Next, since the value of log2_ctu_size_minus2 of the basic block data segmentation information is 4, the size of the basic block is obtained as 64×64 pixels from 1<<log2_ctu_size_minus2+2.

[0111] Next, obtain the sub-picture segmentation information. First, since the subpics_present_flag is 1, it can be determined that the image is segmented into sub-pictures. Next, obtain the information of the maximum number of sub-pictures represented by max_subpics_minus2. In this embodiment, when the value 0 is obtained and 2 is added to this value, the maximum number of sub-pictures is obtained as 2. After that, obtain the basic pixel information of the grid forming the sub-pictures represented by subpic_grid_col_width_minus1 and subpic_grid_row_height_minus1. Then, derive the number of vertical and horizontal grids forming the image, represented by NumSubPicGridRows and NumSubPicGridCols, through a predetermined calculation method. In this embodiment, subpic_grid_col_width_minus1 can be obtained as 63, and subpic_grid_row_height_minus1 can be obtained as 63.

[0112] Calculate NumSubPicGridRows and NumSubPicGridCols through the following equations.

[0113] NumSubPicGridCols =

[0114] (pic_width_in_luma_samples +

[0115] subpic_grid_col_width_minus1 * 4 + 3) /

[0116] (subpic_grid_col_width_minus1 * 4 + 4)

[0117] NumSubPicGridRows =

[0118] (pic_height_in_luma_samples +

[0119] subpic_grid_row_height_minus1 * 4 + 3) /

[0120] (subpic_grid_row_height_minus1 * 4 + 4)

[0121] In this embodiment, NumSubPicGridCols is derived as 6, and NumSubPicGridRows is derived as 3.

[0122] Thereafter, the index information indicating which sub-picture each grid represented by subpic_grid_idx[i][j] belongs to is obtained as many as the necessary quantity represented by NumSubpicGridRows and NumSubpicGridCols. In this embodiment, this quantity is 18. As for the values, 0 is obtained for [0][0], [0][1], [0][2], [1][0], [1][1], [1][2], [2][0], [2][1], and [2][2]. Additionally, 1 is obtained for [0][3], [0][4], [0][5], [1][3], [1][4], [1][5], [2][3], [2][4], and [2][5]. Based on this information, it can be restored that the image is segmented into two sub-pictures, and the two sub-pictures are segmented as shown in Figure 8 Thus, the bitstream that has been effectively encoded with the syntax of the sub-picture in the first embodiment can be decoded.

[0123] After obtaining the information, information such as slice data segmentation information and block data segmentation information required for decoding is obtained. The segmentation information derived from the separation decoding unit 202 is sent to the image reproduction unit 205 and is used to specify the position of the data in the image that will be processed in step S404.

[0124] In step S402, the decoding unit 203 decodes the code data separated in step S401 and reproduces the quantization coefficients and prediction information. In step S403, the inverse quantization / inverse transformation unit 204 inverse quantizes the quantization coefficients to obtain transform coefficients, and also performs an inverse orthogonal transformation to reproduce the prediction error.

[0125] In step S404, the image reproduction unit 205 reproduces the prediction information and the prediction image generated in step S403. The image reproduction unit 205 also reproduces the image data based on the reproduced prediction image and the prediction error generated in step S404. The reproduced image data is synthesized into the appropriate position in the image based on the segmentation information generated in step S401.

[0126] In step S405, the image decoding device determines whether the decoding of all basic blocks in the frame has ended. If the decoding has ended, the process proceeds to step S406. Otherwise, the process returns to step S402 to process the next basic block.

[0127] In step S406, the in-loop filter unit 207 performs in-loop filtering processing on the image data reproduced in step S404 to generate an image that has undergone filtering processing, and the process ends.

[0128] Using the above configuration and operations, the bitstream generated by effectively encoding the syntax of the sub - pictures generated in the first embodiment can be decoded.

[0129] Note that in this embodiment, a value obtained by subtracting 2 from the maximum number of sub - pictures is acquired. However, the present invention is not limited thereto. Instead, max_subpics_minus1 obtained by subtracting 1 from the maximum number of sub - pictures can be acquired, and when this value is 1 or greater, that is, only when the maximum number of sub - pictures is 2 or greater, the following information can be acquired. That is, subpic_grid_col_width_minus1, subpic_grid_row_height_minus1, and subpic_grid_idx[i][j] can be acquired. That is, if the maximum number is 1, these information may not be acquired. This enables the decoding of a bitstream with reduced redundant sub - picture information.

[0130] Note that in the above description, when the value of max_subpics_minus1 is 1 or greater, that is, only when the maximum number of sub - pictures is 2 or greater, the following information is acquired. That is, when the maximum number of sub - pictures is 2 or greater, the information of subpic_grid_col_width_minus1, subpic_grid_row_height_minus1, and subpic_grid_idx[i][j] is acquired. When the maximum number of sub - pictures is 2 or greater, the following two information can also be acquired. That is, only when the value of max_subpics_minus1 is 1 or greater, the two values of subpic_treated_as_flag and loop_filter_across_subpic_enabled_flag can also be acquired. This enables the decoding of a bitstream with further reduced redundant sub - picture information.

[0131] In addition, in this embodiment, an example of dividing an image into two sub - pictures has been described. However, the present invention is not limited thereto. For example, in Figure 10 the case where the shown image is divided into three sub - pictures, the bitstream created by an encoding method for setting the maximum value of subpic_grid_idx to max_subpics_minus1 or less can also be decoded.

[0132] In addition, in this embodiment, a four-pixel defined grid is used. However, a CTU row / column defined grid can be used. In this case, when restoring the values of subpic_grid_col_width_minus1 and subpic_grid_row_height_minus1, the bitstream with reduced code amount can also be decoded by calculating based on the number of basic pixels of the CTU.

[0133] Note that, in this embodiment, subpic_grid_idx, which is the ID representing the position of the rectangular area of the subpicture, is obtained from the sequence parameter set. However, the present invention is not limited to this. Figure 11 The configuration of the bitstream is shown, where subpic_idx, which is the ID of the corresponding subpicture, is changed to be encoded in the slice header that is the header of each slice. Figure 11 The configuration of the bitstream is shown, where subpic_idx, which is the ID of the corresponding subpicture, is changed to be encoded in the slice header that is the header of each slice. When such a configuration is adopted, the number of information representing the position of the rectangular area of the subpicture can be reduced from the number of grids to the number of slices. In Figure 8 In the example shown, the number of grids is 18, and the number of subpictures = the number of slices. Therefore, since the number of slices is 2, decoding can be performed by obtaining two subpic_idx and calculating the same information as 18 subpic_grid_idx.

[0134] In addition, the embodiment where subpic_idx is arranged in the slice header has been described. However, the present invention is not limited to this, and even if subpic_idx is arranged in the block code data, the same information as subpic_grid_idx can be decoded.

[0135] (Third Embodiment)

[0136] In the above embodiment, it is described assuming that Figure 1 the processing unit shown in 1 or 2 is formed by hardware. However, the processing performed by the processing unit shown in these drawings can be configured by a computer program.

[0137] Figure 5 It is a block diagram showing an example of the configuration of the hardware of a computer applicable to the image processing apparatus according to each embodiment.

[0138] The CPU 501 controls the entire computer using computer programs and data stored in the RAM 502 or the ROM 503, and also executes the above-mentioned various processes as the processes to be performed by the image processing apparatus according to the above embodiment. That is, the CPU 501 serves as Figure 1 or Figure 2The processing unit shown.

[0139] RAM 502 includes areas configured to temporarily store computer programs and data loaded from an external storage device 506 and data acquired from the outside via an I / F (interface) 507. In addition, RAM 502 includes a work area used by the CPU 501 to perform various processes. That is, for example, RAM 502 can be allocated as a frame memory, or various other types of areas can be appropriately provided.

[0140] ROM 503 stores setting data of the computer, a boot program, etc. The operation unit 504 is formed by a keyboard, a mouse, etc. When a user of the computer operates the operation unit 504, various instructions can be input into the CPU 501. The display unit 505 displays the processing result of the CPU 501. In addition, the display unit 505 is formed by, for example, a liquid crystal display.

[0141] The external storage device 506 is a mass information storage device represented by a hard disk drive. The external storage device 506 stores an OS (operating system) and a computer program configured to enable the CPU 501 to implement Figure 1 or Figure 2 the functions of the unit shown. The external storage device 506 can also store each image data to be processed.

[0142] Under the control of the CPU 501, the computer programs and data stored in the external storage device 506 are appropriately loaded into the RAM 502 and processed by the CPU 501. The I / F 507 can connect to a network such as a LAN or the Internet or other devices such as a projection device or a display device. The computer can acquire or transmit various information via the I / F 507. Reference numeral 508 denotes a bus connecting the above units.

[0143] For the operation formed by the above configuration, the CPU 501 plays a main role in controlling the operation described with reference to the above flowchart.

[0144] Other embodiments

[0145] The present invention can be implemented by the following processing: supplying a program for implementing one or more functions of the above embodiments to a system or device via a network or a storage medium, and causing one or more processors in a computer of the system or device to read and execute the program. The present invention can also be implemented by a circuit (for example, an ASIC) for implementing one or more functions.

[0146] The present invention is not limited to the above embodiments, and various changes and modifications can be made within the spirit and scope of the present invention. Therefore, in order to inform the public of the scope of the present invention, the appended claims are proposed.

[0147] This application claims priority to Japanese Patent Application No. 2019-165580, filed on September 11, 2019, which is hereby incorporated by reference.

Claims

1. An image encoding device for dividing an image into one or more sub - pictures each including basic blocks and encoding the one or more sub - pictures, the image encoding device comprising: A first encoding unit configured to encode first information corresponding to the block size of each basic block into a sequence parameter set of a bitstream, wherein the block size of each basic block is derived by performing an arithmetic left shift of 1 on the sum of the value of the first information and a predetermined integer; A second encoding unit configured to encode a flag related to the presence of information of a sub - picture into the sequence parameter set of the bitstream; A third encoding unit configured to, when the flag has a value indicating the presence of information of the sub - picture, encode a syntax element corresponding to the number of sub - pictures included in the image and second information of an ID related to the sub - pictures in the image into the sequence parameter set of the bitstream; and A fourth encoding unit configured to encode third information for identifying the arrangement of sub - pictures in the image into the sequence parameter set of the bitstream, wherein when the block size derived based on the first information is 64, the value of the third information is equal to the result of subtracting 1 from the value identified in units of 64, wherein, when the value represented by the syntax element is one or more, the third information for identifying the arrangement of sub - pictures in the image is encoded into the sequence parameter set of the bitstream, wherein, when the value represented by the syntax element is not one or more, the third information for identifying the arrangement of sub - pictures in the image is not encoded into the sequence parameter set of the bitstream, and wherein fourth information corresponding to the width of the image is encoded into the bitstream.

2. The device according to claim 1, wherein, The value of the first information is an integer.

3. The device according to claim 1, wherein, The sub - picture is a rectangular region.

4. The device according to claim 1, wherein, The loop_filter_across_subpic_enabled_flag related to filtering can be encoded into the bitstream.

5. The device according to claim 4, wherein, When the value represented by the syntax element is not one or more, the loop_filter_across_subpic_enabled_flag is not encoded into the bitstream.

6. The device according to claim 1, wherein, The fourth information represents the width of the image in units of luminance samples.

7. An image decoding device capable of decoding a bitstream obtained by encoding an image including one or more sub - pictures each including basic blocks, the image decoding device comprising: A first decoding unit configured to decode first information corresponding to the block size of each basic block from a sequence parameter set of the bitstream, wherein the block size of each basic block is derived by performing an arithmetic left shift of 1 on the sum of the value of the first information and a predetermined integer; A second decoding unit configured to decode a flag related to the presence of information of a sub - picture from the sequence parameter set of the bitstream; A third decoding unit configured to, when the flag has a value indicating the presence of information of the sub-picture, decode from the sequence parameter set of the bitstream a syntax element corresponding to the number of sub-pictures included in the image and second information of an ID related to the sub-pictures in the image; and A fourth decoding unit configured to decode from the sequence parameter set of the bitstream third information for identifying the arrangement of the sub-pictures in the image, wherein when the block size derived based on the first information is 64, the value of the third information is equal to the result of subtracting 1 from the value identified in units of 64, wherein when the numerical value represented by the syntax element is one or more, the third information for identifying the arrangement of the sub-pictures in the image is decoded from the sequence parameter set of the bitstream, wherein when the numerical value represented by the syntax element is not one or more, the third information for identifying the arrangement of the sub-pictures in the image is not decoded from the sequence parameter set of the bitstream, and wherein fourth information corresponding to the width of the image is decoded from the bitstream.

8. The apparatus according to claim 7, wherein, The value of the first information is an integer.

9. The device according to claim 7, wherein The sub-picture is a rectangular region.

10. The apparatus according to claim 7, wherein The loop_filter_across_subpic_enabled_flag related to filtering can be decoded from the bitstream.

11. The apparatus according to claim 10, wherein, When the numerical value represented by the syntax element is not one or more, the loop_filter_across_subpic_enabled_flag is not decoded from the bitstream.

12. The device according to claim 7, wherein, The fourth information represents the width of the image in units of luminance samples.

13. An image encoding method for dividing an image into one or more sub-pictures each including basic blocks and encoding the one or more sub-pictures, the image encoding method comprising: Encoding first information corresponding to the block size of each basic block into a sequence parameter set of a bitstream, wherein the block size of each basic block is derived by performing an arithmetic left shift of 1 on the sum of the value of the first information and a predetermined integer; Encoding a flag related to the presence of information of the sub-picture into the sequence parameter set of the bitstream; When the flag has a value indicating the presence of information of the sub-picture, encoding a syntax element corresponding to the number of sub-pictures included in the image and second information of an ID related to the sub-pictures in the image into the sequence parameter set of the bitstream; and Encoding third information for identifying the arrangement of the sub-pictures in the image into the sequence parameter set of the bitstream, wherein when the block size derived based on the first information is 64, the value of the third information is equal to the result of subtracting 1 from the value identified in units of 64, wherein when the numerical value represented by the syntax element is one or more, the third information for identifying the arrangement of the sub-pictures in the image is encoded into the sequence parameter set of the bitstream, Wherein, when the value represented by the syntax element is not one or more than one, the third information for identifying the arrangement of sub - pictures in the image is not encoded into the sequence parameter set of the bitstream, and wherein, fourth information corresponding to the width of the image is encoded into the bitstream.

14. An image decoding method capable of decoding a bitstream obtained by encoding an image including one or more than one sub - pictures each including a basic block, the image decoding method comprising: decoding, from the sequence parameter set of the bitstream, first information corresponding to the block size of each basic block, wherein the block size of each basic block is derived by performing an arithmetic left - shift of 1 on the sum of the value of the first information and a predetermined integer; decoding, from the sequence parameter set of the bitstream, a flag related to the presence of information of the sub - picture; when the flag has a value indicating the presence of information of the sub - picture, decoding, from the sequence parameter set of the bitstream, a syntax element corresponding to the number of sub - pictures included in the image and second information of an ID related to the sub - pictures in the image; and decoding, from the sequence parameter set of the bitstream, third information for identifying the arrangement of sub - pictures in the image, wherein when the block size derived based on the first information is 64, the value of the third information is equal to the result of subtracting 1 from the value identified in units of 64, wherein, when the value represented by the syntax element is one or more than one, decoding the third information for identifying the arrangement of sub - pictures in the image from the sequence parameter set of the bitstream, wherein, when the value represented by the syntax element is not one or more than one, not decoding the third information for identifying the arrangement of sub - pictures in the image from the sequence parameter set of the bitstream, and wherein, decoding fourth information corresponding to the width of the image from the bitstream.

Citation Information

Patent Citations

  • Image encoder, image encoding method and program, image decoder, and image decoding method and program

    JP2014011638A

  • Power system, controller, power management method, program, and power management server

    JP2019165580A