Image encoding device, image decoding device, image encoding method, and image decoding method
By segmenting images into sub-images and encoding the maximum number of sub-images with unique identifiers, the method addresses redundant syntax issues in VVC, resulting in reduced bitstream size.
Patent Information
- Application Number
- CN202510485398.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-11
- Filing Date
- 2020-08-21
- Publication Date
- 2025-07-15
AI Technical Summary
In the VVC encoding method, there is a redundant part in the syntax of the sub-picture, resulting in an increase in the code volume.
By dividing the image into a plurality of sub-pictures and encoding each sub-picture, it is possible to decode independently, and the first determining component determines whether the sub-picture is a rectangle, the second determining component determines the vertical/horizontal size of the grid, assigns the component assigns the number, and encodes when the maximum number determined by the third determining component is not less than 2.
The code amount of generated bitstreams is effectively reduced and the encoding efficiency is improved.
Smart Images

Figure CN120321391A_ABST
Abstract
Description
[0001] (This application is a divisional application of the application with the filing date of August 21, 2020, application number 2020800638779, and invention title "Image encoding device and image decoding device, and their methods and storage media".) Technical Field
[0002] The present invention relates to an image encoding device, an image encoding method and program, an image decoding device, an image decoding method and program, and image code data. More specifically, the present invention relates to an image encoding / decoding method capable of dividing a picture into rectangles and extracting independent code data. Background Art
[0003] As an encoding method for compression recording of moving images, the HEVC (High Efficiency Video Coding) encoding method (hereinafter referred to as HEVC) is known. In HEVC, in order to improve the encoding efficiency, basic blocks larger than conventional macroblocks (16×16 pixels) are adopted. The basic blocks with large sizes are called CTUs (Coding Tree Units), and their maximum size is 64×64 pixels. A CTU is further divided into sub-blocks as units for prediction or transformation.
[0004] In addition, in HEVC, a picture can be divided into a plurality of tiles or slices and encoded. Tiles or slices have small data dependencies, and encoding / decoding processing can be executed in parallel. A great advantage of tile or slice division is that processing can be executed in parallel by a multi-core CPU or the like to shorten the processing time. Patent Document 1 discloses a technique related to tiles and slices.
[0005] In recent years, international standardization activities for more efficient encoding methods as successors to HEVC have been started. JVET (Joint Video Experts Team) has been established between ISO / IEC and ITU-T, and the VVC (Versatile Video Coding) encoding method (hereinafter referred to as VVC) has been standardized. In VVC, there are sub-pictures configured to include one or more slices and forming rectangles. A picture can be divided into one or more sub-pictures, and each sub-picture can be processed as independent code data.
[0006] Citation List
[0007] Patent Document
[0008] Patent Document 1: Japanese Unexamined Patent Application Publication No. 2014-11638 Summary of the Invention
[0009] Problems to be Solved by the Invention
[0010] Regarding sub - pictures of VVC, values representing the maximum number of sub - pictures, basic pixel numbers such as the vertical / horizontal sizes of a grid defining the boundary positions of each sub - picture, the ID of each sub - picture, control flags corresponding to each sub - picture, etc. are defined as syntax.
[0011] There are still many redundant parts in the syntax regarding sub - pictures, resulting in an increase in the code amount.
[0012] Therefore, the present invention is made to solve the above - mentioned problems and provides a technique for reducing the code amount of the generated bit - stream by eliminating the redundancy of the syntax regarding sub - pictures.
[0013] Solutions for solving the problems
[0014] To solve the above - mentioned problems, an image encoding device according to the present invention has the following configuration. That is, the image encoding device is an image encoding device including an encoding component for dividing an image into a plurality of sub - pictures and encoding each sub - picture so that each sub - picture can be independently decoded, and is characterized by including: a first determination component for determining whether each sub - picture forming the image is defined as only one rectangle; a second determination component for determining the basic pixel number, which is the vertical / horizontal size of the grid forming the sub - picture; an allocation component for assigning numbers to each grid divided by the basic pixel number; and a third determination component for determining the maximum number of the sub - pictures, where the maximum number of the sub - pictures is the number of the numbers assigned by the allocation component, wherein the sub - picture is a set of grids assigned the same number, and encodes the value obtained by subtracting 2 from the maximum number determined by the third determination component.
[0015] In addition, an image decoding device according to the present invention has the following configuration. That is, the image decoding device is an image decoding device including an image decoding component capable of decoding a bit - stream including code data obtained by encoding each of a plurality of sub - pictures, and is characterized by including: a first acquisition component for acquiring from the bit - stream information indicating whether each sub - picture forming the image includes only one rectangle; a second acquisition component for acquiring from the bit - stream the maximum number of the sub - pictures; and a third acquisition component for acquiring from the bit - stream the basic pixel number, which is the vertical / horizontal size of the grid forming the sub - picture, wherein the value obtained by adding 2 to the numerical value acquired by the second acquisition component is set as the maximum number of the sub - pictures, the number of vertical / horizontal grids forming the image is acquired according to the vertical / horizontal size of the grid acquired by the third acquisition component, and information indicating which sub - picture each grid belongs to is acquired.
[0016] In addition, the image encoding device according to the present invention has the following configuration. That is to say, the image encoding device is an image encoding device including an encoding component for dividing an image into a plurality of sub-pictures and encoding each sub-picture so that each sub-picture can be independently decoded, characterized in that it includes: a first determination component for determining whether each sub-picture forming the image is defined as only one rectangle; a second determination component for determining a basic number of pixels, which is the vertical / horizontal size of the grid forming the sub-picture; an allocation component for allocating numbers to each grid divided by the basic number of pixels; and a third determination component for determining a maximum number of the sub-pictures, which is the number of the numbers allocated by the allocation component, wherein the sub-picture is a set of grids allocated the same number, and the number and the basic number of pixels are encoded only when the maximum number determined by the third determination component is not less than 2.
[0017] In addition, the image decoding device according to the present invention has the following configuration. That is to say, the image decoding device is an image decoding device including an image decoding component capable of decoding a bitstream including code data obtained by encoding each of a plurality of sub-pictures, characterized in that it includes: a first acquisition component for acquiring from the bitstream information indicating whether each sub-picture forming the image includes only one rectangle; a second acquisition component for acquiring from the bitstream the maximum number of the sub-pictures; and a third acquisition component for acquiring from the bitstream a basic number of pixels, which is the vertical / horizontal size of the grid forming the sub-picture, wherein the sub-picture is a set of grids allocated the same number, and the number and the basic number of pixels are acquired only when the maximum number determined by the second acquisition component is not less than 2.
[0018] Effects of the Invention
[0019] According to the present invention, the syntax regarding the sub-pictures forming an image can be effectively encoded and the encoding can be efficiently improved.
[0020] From the following description in conjunction with the drawings, other features and advantages of the present invention will become apparent. Note that throughout the drawings, the same reference numerals denote the same or similar components. Description of the Drawings
[0021] The drawings included in the specification and constituting a part of the specification illustrate embodiments of the present invention and, together with the specification, are used to explain the principles of the present invention.
[0022] Figure 1 is a block diagram showing the configuration of the image encoding device;
[0023] Figure 2 is a block diagram showing the configuration of an image decoding device;
[0024] Figure 3 is a flowchart showing the image encoding process in an image encoding device;
[0025] Figure 4 is a flowchart showing the image decoding process in an image decoding device;
[0026] Figure 5 is a block diagram showing an example of the hardware configuration;
[0027] Figure 6 is a diagram showing the bitstream configuration;
[0028] Figure 7 is a diagram showing an example of image segmentation;
[0029] Figure 8 is a diagram showing an example of image segmentation;
[0030] Figure 9 is a diagram showing the bitstream configuration;
[0031] Figure 10 is a diagram showing an example of image segmentation; and
[0032] Figure 11 is a diagram showing the bitstream configuration.
[0033] List of Reference Numerals
[0034] 101, 112, 201, 208: Terminals
[0035] 102: Image segmentation unit
[0036] 103: Block segmentation unit
[0037] 104: Prediction unit
[0038] 105: Transform / quantization unit
[0039] 106, 204: Inverse quantization / inverse transform unit
[0040] 107, 205: Image reproduction unit
[0041] 108, 206: Frame memory
[0042] 109, 207: In-loop filter unit
[0043] 110: Encoding unit
[0044] 111: Integrated encoding unit
[0045] 202: Separation decoding unit
[0046] 203: Decoding unit Detailed implementation manners
[0047] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments are not intended to limit the scope of the claimed invention. In the embodiments, multiple features are described, but the invention does not require all of these features, and multiple of these features can be appropriately combined. In addition, in the accompanying drawings, the same reference numerals are given to the same or similar configurations, and redundant descriptions thereof are omitted.
[0048] (First embodiment)
[0049] Now, embodiments of the present invention will be described with reference to the accompanying drawings. Figure 1 is a block diagram showing an image encoding device according to this embodiment. Refer to Figure 1 , reference numeral 101 denotes a terminal configured to input image data.
[0050] Reference numeral 102 denotes an image segmentation unit that divides the input image into one or more slice rows or one or more slice columns. Each slice is a set of consecutive basic blocks that cover a rectangular area in the image. The image segmentation unit 102 also divides a slice into one or more bricks. Each brick is a rectangle formed by one or more rows of basic blocks, and these rows of basic blocks are rows of basic blocks in the slice. The image segmentation unit 102 divides the image into strips, and each strip is formed by one or more slices in the image or one or more bricks in a slice. A strip is a basic unit of encoding, and header information such as information indicating the strip type is added to each strip.
[0051] In this embodiment, as Figure 8 shown, a picture is divided into two sub - pictures. In addition, for simplicity of description, the sub - pictures, strips, slices, and bricks have the same vertical / horizontal size and the same position. However, the present invention is not limited thereto.
[0052] Reference numeral 103 denotes a block segmentation unit that divides the row image of basic blocks output from the image segmentation unit 102 into multiple basic blocks and outputs the image of each basic block to the subsequent stage.
[0053] Reference numeral 104 denotes a prediction unit that determines sub-block division of the image data of each basic block and performs intra prediction as intra-frame prediction or inter prediction as inter-frame prediction for each sub-block, thereby generating predicted image data. Intra prediction across blocks or motion vector prediction is not performed. Further, the prediction unit 104 calculates a prediction error based on the input image data and the predicted image data, and outputs the prediction error. Additionally, the prediction unit 104 outputs information required for prediction, such as information such as sub-block division, prediction mode, and motion vector, and the prediction error. The information required for prediction will hereinafter be referred to as prediction information.
[0054] Reference numeral 105 denotes a transform / quantization unit that orthogonally transforms the prediction error based on sub-blocks to obtain transform coefficients, and further quantizes the transform coefficients to obtain quantization coefficients. Reference numeral 106 denotes an inverse quantization / inverse transform unit that inverse quantizes the quantization coefficients output from the transform / quantization unit 105 to reproduce the transform coefficients, and also inverse orthogonally transforms the transform coefficients to reproduce the prediction error. Reference numeral 108 denotes a frame memory that stores the reproduced image data.
[0055] Reference numeral 107 denotes an image reproduction unit. The image reproduction unit 107 generates predicted image data by appropriately referring to the frame memory 108 based on the prediction information output from the prediction unit 104, generates reproduced image data based on the predicted image data and the input prediction error, and outputs the reproduced image data. Reference numeral 109 denotes an in-loop filter unit. The in-loop filter unit 109 performs in-loop filtering processing such as deblocking filtering processing or sample adaptive offset on the reproduced image, and outputs the image that has undergone the filtering processing.
[0056] Reference numeral 110 denotes an encoding unit. The encoding unit 110 generates code data by encoding the quantization coefficients output from the transform / quantization unit 105 and the prediction information output from the prediction unit 104, and outputs the code data. Reference numeral 111 denotes an integrated encoding unit. The integrated encoding unit 111 receives segmentation information from the image segmentation unit 102 and generates header code data. The integrated encoding unit 111 also forms a bitstream by combining the header code data with the code data output from the encoding unit 110, and outputs the bitstream. Reference numeral 112 denotes a terminal that outputs the bitstream generated by the integrated encoding unit 111 to the outside.
[0057] The image encoding operation of the image encoding device according to the first embodiment will be described below. In this embodiment, frame input moving image data is used. However, still image data corresponding to one frame may also be input. Further, in this embodiment, for the sake of easy description, only the intra prediction encoding process will be described. However, the present invention is not limited thereto, and it may also be applied to the inter prediction encoding process. Further, in this embodiment, for the sake of description, it is assumed that the block splitting unit 103 splits the image into basic blocks each having a size of 64×64 pixels for description. However, the present invention is not limited thereto.
[0058] In this embodiment, first, the image splitting unit 102 splits one frame of image data input from the terminal 101 into sub-pictures, as Figure 8 shown. In this embodiment, the case where one image data is split into two sub-pictures and formed by 18 grids will be described. The vertical / horizontal size of one grid is 256×256 pixels, and the vertical / horizontal size of the sub-picture is 768×768 pixels. For each grid of the left sub-picture, 0 is set as the index value. For each grid of the right sub-picture, 1 is set as the index value. As described above, in this embodiment, one image data is divided into a plurality of grids and split into a plurality of sub-pictures based on the plurality of grids. Further, each grid is assigned an index value as a number, and the grids forming one sub-picture are assigned the same index value.
[0059] In this embodiment, each sub-picture is formed by one strip and one slice / one block. Information about the size of the sub-picture or the strip is sent to the integrated encoding unit 111 as splitting information. Further, each block is split into basic block row images, and sent to the block splitting unit 103, where each basic block row image is the image data of each basic block row.
[0060] The block splitting unit 103 splits the input basic block row image into a plurality of basic blocks, and outputs the image of each basic block to the prediction unit 104. In this embodiment, the image of each basic block with a size of 64×64 pixels is output.
[0061] The prediction unit 104 performs prediction processing on the image data of each basic block input from the block splitting unit 103. More specifically, the sub-block splitting for splitting the basic block into finer sub-blocks is determined, and the intra prediction mode such as horizontal prediction or vertical prediction is determined based on the sub-blocks.
[0062] Prediction image data is generated based on the determined intra prediction mode and the encoded pixels. Further, a prediction error is generated based on the input image data and the prediction image data, and output to the transform / quantization unit 105. In addition, information about the sub-block splitting or the intra prediction mode is output to the encoding unit 110 and the image reproduction unit 107 as prediction information.
[0063] The transform / quantization unit 105 performs orthogonal transform / quantization on the input prediction error to generate quantization coefficients. First, an orthogonal transform process corresponding to the size of the sub-block is performed to generate orthogonal transform coefficients. Next, the orthogonal transform coefficients are quantized to generate quantization coefficients. The generated quantization coefficients are output to the coding unit 110 and the inverse quantization / inverse transform unit 106.
[0064] The inverse quantization / inverse transform unit 106 inverse quantizes the input quantization coefficients to reproduce the transform coefficients, and further inverse orthogonally transforms the reproduced transform coefficients to reproduce the prediction error. The reproduced prediction error is output to the image reproduction unit 107.
[0065] The image reproduction unit 107 reproduces the prediction image by appropriately referring to the frame memory 108 based on the prediction information input from the prediction unit 104. Then, the image data is reproduced according to the reproduced prediction image and the reproduced prediction error input from the inverse quantization / inverse transform unit 106. The image data is input and stored in the frame memory 108.
[0066] The in-loop filter unit 109 reads out the reproduced image from the frame memory 108 and performs in-loop filtering processes such as deblocking filtering. The image that has undergone the filtering process is input to the frame memory 108 again and stored again.
[0067] The coding unit 110 performs entropy coding on the quantization coefficients generated by the transform / quantization unit 105 and the prediction information input from the prediction unit 104 based on blocks, thereby generating code data. The method of entropy coding is not particularly specified, and Columbus coding, arithmetic coding, Huffman coding, etc. can be used. The generated code data is output to the integrated coding unit 111.
[0068] The integrated coding unit 111 receives segmentation information from the image segmentation unit 102, generates the coded data of the header, and multiplexes the code data input from the coding unit 110, etc., thereby forming a bitstream. Finally, the bitstream is output from the terminal 112 to the outside.
[0069] The format of the coded data using VVC encoded by the image coding device according to this embodiment is as Figure 6 shown. In Figure 6 the coded data shown, first, there is a sequence parameter set as header information including information about sequence coding. Subsequently, there are a picture parameter set as header information including information about picture coding, a strip header as header information including information about the coding of each strip, and the coded data of each block.
[0070] In the sequence parameter set, pic_width_in_luma_samples and pic_height_in_luma_samples exist as image size information. These represent the number of pixels in the horizontal direction and the number of pixels in the vertical direction with respect to the image luminance. In this embodiment, since Figure 8 the image shown in is encoded, pic_width_in_luma_samples is 1536, and pic_height_in_luma_samples is 768. Additionally, as basic block data segmentation information, there is log2_ctu_size_minus2 which represents the size of the basic block. The number of vertical and horizontal pixels of the basic block is represented by 1<<(log2_ctu_size_minus2 + 2). In this embodiment, since the basic block has 64×64 pixels, the value of log2_ctu_size_minus2 is 4.
[0071] Furthermore, as sub-picture information, there is subpics_present_flag which is information indicating whether sub-picture segmentation exists. If subpics_present_flag is 1, this indicates that the image is segmented into one or more sub-pictures. If subpics_present_flag is 0, this indicates that the image is not segmented.
[0072] If subpics_present_flag is 1, the following information is also encoded. That is, the information encoded at this time includes, for example, max_subpics_minus2 which represents the maximum number of sub-pictures. Additionally, the information encoded at this time includes, for example, information such as subpic_grid_col_width_minus1 and subpic_grid_row_height_minus1 which represent the basic number of pixels of the grid. Furthermore, the information encoded at this time includes, for example, information such as subpic_grid_idx[i][j] which is an index value. Additionally, the information encoded at this time includes, for example, information about filtering such as subpic_treated_as_pic_flag[i] and loop_filter_across_subpic_enabled_flag[i].
[0073] Traditionally, the maximum number of sub - pictures is expressed as "maximum number - 1". In this case, the maximum number of sub - pictures is 1, that is, there is one sub - picture, and the design is such that the code amount is minimized in a state where the image is not actually divided. However, in this embodiment, the maximum number of sub - pictures is expressed as "maximum number - 2". If the image is divided into two sub - pictures, the value of max_subpics_minus2 is 0, and it can be represented with the minimum code amount.
[0074] The picture parameter set includes slice data information, block data information, stripe data information, and basic block data information. The stripe header includes stripe data information, etc., and is followed by the encoded data of each block.
[0075] Figure 3 is a flowchart showing the encoding process in the image encoding device according to the first embodiment. First, in step S301, the image segmentation unit 102 divides the image into slices, blocks, and stripes as described above, and sends the segmentation information to the integrated encoding unit 111. In step S309, the integrated encoding unit 111 converts the segmentation information into header information and encodes it into a bitstream. The image segmentation unit 102 also divides the image into sub - pictures and sends these sub - pictures to the block segmentation unit 103.
[0076] In step S302, the block segmentation unit 103 divides the basic block row image into basic blocks. In step S303, the prediction unit 104 performs prediction processing on the image data of each basic block generated in step S302, thereby generating sub - block segmentation information, prediction information such as intra - prediction mode, and prediction image data. The prediction unit 104 also calculates the prediction error based on the input image data and the prediction image data.
[0077] In step S304, the transform / quantization unit 105 orthogonally transforms the prediction error calculated in step S303 to generate transform coefficients, and performs quantization to generate quantization coefficients.
[0078] In step S305, the inverse quantization / inverse transform unit 106 performs inverse quantization on the quantization coefficients generated in step S304, thereby reproducing the transform coefficients. The inverse quantization / inverse transform unit 106 also performs inverse orthogonal transformation on the transform coefficients, thereby reproducing the prediction error.
[0079] In step S306, the image reproduction unit 107 reproduces the prediction image based on the prediction information generated in step S303. The image reproduction unit 107 also reproduces the image data based on the reproduced prediction image and the prediction error generated in step S305.
[0080] In step S307, the encoding unit 110 encodes the prediction information generated in step S303 and the quantization coefficients generated in step S304 to generate code data.
[0081] In step S308, the image encoding device determines whether the encoding of all basic blocks in the strip has ended. If the encoding has ended, the process proceeds to step S309. Otherwise, the process returns to step S302 to process the next basic block.
[0082] In step S309, the integration encoding unit 111 generates header information based on the segmentation information sent from the image segmentation unit 102 and performs encoding.
[0083] Based on Figure 8 The following describes a detailed example in this embodiment based on the image segmentation shown. In this embodiment, since the image is segmented into two sub - pictures, subpics_present_flag is 1. Next, for max_subpics_minus2 representing the maximum number of sub - pictures, the value of "maximum number - 2" is set, that is, "0" in this embodiment. If this information is encoded by Golomb coding, the 1 - bit data "0" representing "0" is encoded. Compared with the case of encoding the 3 - bit data "010" representing the value "1", this can improve the encoding efficiency.
[0084] subpic_grid_col_width_minus1 is 63, and subpic_grid_row_height_minus1 is also 63. These two values are defined in units of four pixels, and more specifically, these are the values obtained by subtracting 1 from the value obtained by dividing the basic number of pixels in the grid by 4.
[0085] There is subpic_grid_idx[i][j] which is the index value of the sub - picture. Here, [i][j] represents the position of the grid, where i represents the row direction and j represents the column direction. In this embodiment, the range of the row direction is from 0 to 2, and the range of the column direction is from 0 to 5, and [0][0], [0][1], [0][2], [1][0], [1][1], [1][2], [2][0], [2][1], and [2][2] are 0. Additionally, [0][3], [0][4], [0][5], [1][3], [1][4], [1][5], [2][3], [2][4], and [2][5] are 1. The code data of the strip forming the sub - picture is combined based on these sub - picture information to create the code data of the sub - picture.
[0086] In step S310, the image encoding device determines whether the encoding of all the basic blocks in the frame has ended. If the encoding has ended, the process proceeds to step S311. Otherwise, the process returns to step S302 to process the next basic block.
[0087] In step S311, the in-loop filter unit 109 performs in-loop filtering on the image data reproduced in step S306 to generate a filtered image, and the process ends.
[0088] With the above configuration and operations, particularly in step S309, if an image is segmented into sub-pictures, the information representing the maximum number of sub-pictures is encoded by "maximum number - 2", thus effectively encoding the syntax regarding sub-pictures.
[0089] Note that in this embodiment, a sub-picture is formed by one stripe, one slice, or one block. However, the present invention is not limited thereto. Figure 7 An example in which a stripe is segmented into multiple slices is shown. More specifically, the stripes of index 0 and index 3 are each formed by four slices, the stripes of index 1 and index 2 are each formed by two slices, and the stripe of index 4 is formed by six slices. This embodiment can also be applied to this sub-picture configuration. Note that in this case, a slice and a block have the same basic pixel size.
[0090] Note that in this embodiment, the value obtained by subtracting 2 from the maximum number of sub-pictures is encoded. However, the present invention is not limited thereto. Instead, max_subpics_minus1 obtained by subtracting 1 from the maximum number of sub-pictures can be encoded, and when the value is 1 or greater, that is, only when the maximum number of sub-pictures is 2 or greater, the following parameters can be encoded. That is, parameters such as subpic_grid_col_width_minus1, subpic_grid_row_height_minus1, and subpic_grid_idx[i][j] can be encoded. That is, if the maximum number is 1, the parameters can be omitted so that these parameters are not encoded.
[0091] The format of the encoded data at this time is as Figure 9 shown. In this case, when the maximum number of sub-pictures is 1, the amount of information of subpic_grid_col_width_minus1, subpic_grid_row_height_minus1, and subpic_grid_idx[0][0] can be reduced. This enables effective encoding.
[0092] Note that in the above description, when the value of max_subpics_minus1 is 1, i.e., only when the maximum number of sub-pictures is 2 or greater, the following parameters are encoded. That is, when the maximum number of sub-pictures is 2 or greater, the parameters subpic_grid_col_width_minus1, subpic_grid_row_height_minus1, and subpic_grid_idx[i][j] are encoded. In addition, only when the maximum number of sub-pictures is 2 or greater, the following two values can be encoded. That is, only when the value of max_subpics_minus1 is 1 or greater, the subpic_treated_as_flag and loop_filter_across_subpic_enabled_flag can be encoded. This enables further reduction of the amount of information and efficient encoding.
[0093] In addition, in this embodiment, an example of dividing an image into two sub-pictures has been described. Figure 10 Examples of dividing an image into, for example, three sub-pictures are shown. The value that can be maximally represented by Ceil(log2(max_subpics_minus1 + 1)) bits can be set as subpic_grid_idx. When using a value of max_subpics_minus1 or less, effective encoding is possible. More specifically, since the image is divided into three parts, the maximum number of sub-pictures is 3. Therefore, max_subpics_minus1 is 2 obtained by subtracting 1 from 3. The number of bits at this time is 2 as a result of the calculation of Ceil(log2(2 + 1)), and a value that can be represented by 2 bits can be set. More specifically, four types of values 0, 1, 2, and 3 can be set. When the maximum value that can be set is limited to max_subpics_minus1, as Figure 10 shown, these values are limited to 0, 1, and 2, and the syntax regarding sub-pictures can be correctly encoded.
[0094] In addition, in this embodiment, a grid is defined based on four pixels. However, a grid can be defined based on CTU rows / columns. If the basic number of pixels of a CTU is 64×64, the value of subpic_grid_col_width_minus1 is 3, and the value of subpic_grid_row_height_minus1 is also 3. The information of the encoding object is small, and the amount of code can be reduced.
[0095] Note that, in this embodiment, subpic_grid_idx, which is an ID representing the position of the rectangular area of the sub-picture, is encoded in the sequence parameter set. However, the present invention is not limited thereto. Figure 11 The configuration of the bitstream is shown, in which subpic_idx, which is the ID of the corresponding sub-picture, is instead encoded in the slice header that is the header of each slice. When such a configuration is adopted, the amount of information representing the position of the rectangular area of the sub-picture can be reduced from the number of grids to the number of slices. In Figure 8 the example shown, the number of grids is 18, and the number of sub-pictures = the number of slices. Therefore, since the number of slices is 2, the amount of information of the encoding object can be reduced from 18 to 2.
[0096] In addition, subpic_idx is arranged in the slice header. However, the present invention is not limited thereto, and subpic_idx can be arranged in the block code data. This can reduce the amount of code, as in the case where subpic_idx is stored in the slice header.
[0097] (Second Embodiment)
[0098] Figure 2 is a block diagram showing the configuration of an image decoding apparatus according to a second embodiment of the present invention. The image decoding apparatus according to this embodiment can decode a bitstream including code data generated by encoding each of a plurality of sub-pictures. In this embodiment, decoding the encoded data generated in the first embodiment will be described below as an example.
[0099] Reference numeral 201 denotes a terminal for inputting an encoded bitstream. Reference numeral 202 denotes a separation decoding unit that separates code data related to information and coefficients regarding decoding processing from the bitstream and sends these to the decoding unit 203. In addition, the separation decoding unit 202 decodes the code data present in the header portion of the bitstream. In this embodiment, the segmentation information is generated by decoding header information regarding image segmentation (such as the sizes of slices, blocks, strips, and basic blocks) and is output to the image reproduction unit 205. The separation decoding unit 202 performs an operation Figure 1 opposite to the operation of the integration encoding unit 111 shown.
[0100] Reference numeral 203 denotes a decoding unit that decodes the coded data output from the separation decoding unit 202 to reproduce quantization coefficients and prediction information. Reference numeral 204 denotes an inverse quantization / inverse transformation unit that inverse-quantizes the quantization coefficients to obtain transform coefficients and performs an inverse orthogonal transformation on the transform coefficients to reproduce a prediction error. Reference numeral 206 denotes a frame memory 206 that stores the image data of the reproduced picture. Reference numeral 205 denotes an image reproduction unit. The image reproduction unit 205 generates prediction image data by appropriately referring to the frame memory 206 based on the input prediction information. Then, the image reproduction unit 205 generates reproduced image data based on the prediction image data and the prediction error reproduced by the inverse quantization / inverse transformation unit 204. The positions of slices, blocks, and stripes in a frame are specified based on the segmentation information input from the separation decoding unit 202, and the generated reproduced image data is output.
[0101] Reference numeral 207 denotes an in-loop filter unit. Similar to the in-loop filter unit 109 shown in Figure 1 the above, the in-loop filter unit 207 performs in-loop filtering processing such as deblocking filtering processing on the reproduced image and outputs the image that has undergone the in-loop filtering processing. Reference numeral 208 denotes a terminal that outputs the reproduced image data to the outside.
[0102] The image decoding operation of the image decoding device according to this embodiment will be described below. In this embodiment, a bitstream generated in the first embodiment based on a frame input is used. However, a still image bitstream corresponding to one frame can be input. Further, in this embodiment, for ease of description, only the intra prediction decoding process will be described. However, the present invention is not limited thereto, and it can also be applied to the inter prediction decoding process.
[0103] In Figure 2 this, a bitstream corresponding to one frame input from the terminal 201 is input to the separation decoding unit 202. The separation decoding unit 202 separates the coded data related to information and coefficients regarding the decoding process from the bitstream and decodes the coded data present in the header portion of the bitstream. More specifically, segmentation information is generated by decoding the basic block data segmentation information, slice data segmentation information, block data segmentation information, stripe data segmentation information, basic block row data synchronization information, and basic block row data position information shown in Figure 6 and is sent to the image reproduction unit 205. Next, the coded data of each basic block of the picture data is reproduced and output to the decoding unit 203.
[0104] The decoding unit 203 decodes the coded data to reproduce quantization coefficients and prediction information. The reproduced quantization coefficients are output to the inverse quantization / inverse transformation unit 204, and the reproduced prediction information is output to the image reproduction unit 205.
[0105] The inverse quantization / inverse transform unit 204 inverse quantizes the input quantized coefficients to generate orthogonal transform coefficients, and performs an inverse orthogonal transform to reproduce the prediction error. The reproduced prediction error is output to the image reproduction unit 205.
[0106] The image reproduction unit 205 reproduces a prediction image by appropriately referring to the frame memory 206 based on the prediction information input from the decoding unit 203. Image data is reproduced based on the prediction image and the prediction error input from the inverse quantization / inverse transform unit 204. For example, as Figure 8 shown, the shape and position of slices, strips, and blocks in a frame are specified based on the segmentation information input from the separation decoding unit 202, and the generated reproduced image data is input and stored in the frame memory 206. The stored image data is used as a reference in prediction.
[0107] Similar to Figure 1 the in-loop filter unit 109 shown, the in-loop filter unit 207 reads out the reproduced image from the frame memory 206, and performs in-loop filtering processes such as deblocking filtering. The image that has undergone the filtering process is input to the frame memory 206 again. The reproduced image stored in the frame memory 206 is finally output from the terminal 208 to the outside.
[0108] Figure 4 is a flowchart showing the image decoding process in the image decoding device according to the second embodiment. First, in step S401, the separation decoding unit 202 separates the code data related to the information and coefficients regarding the decoding process from the bitstream, and decodes the code data in the header part. The separation decoding unit 202 decodes Figure 6 the slice data information, block data information, strip data information, etc. shown, generates the information for decoding, and sends it to the image reproduction unit 205. The segmentation of the image stored in the bitstream in this embodiment is as Figure 8 shown.
[0109] From the values of pic_width_in_luma_samples and pic_height_in_luma_samples of the image size information, it is obtained that the image has 1536×768 pixels.
[0110] Next, since the value of log2_ctu_size_minus2 of the basic block data segmentation information is 4, the size of the basic block is obtained as 64×64 pixels from 1<<log2_ctu_size_minus2+2.
[0111] Next, obtain the sub - picture segmentation information. First, since the subpics_present_flag is 1, it can be determined that the image is segmented into sub - pictures. Next, obtain the information on the maximum number of sub - pictures represented by max_subpics_minus2. In this embodiment, when the value 0 is obtained and 2 is added to this value, the maximum number of sub - pictures is obtained as 2. After that, obtain the basic pixel information of the grid forming the sub - pictures represented by subpic_grid_col_width_minus1 and subpic_grid_row_height_minus1. Then, by a predetermined calculation method, derive the number of vertical and horizontal grids forming the image represented by NumSubPicGridRows and NumSubPicGridCols. In this embodiment, subpic_grid_col_width_minus1 can be obtained as 63, and subpic_grid_row_height_minus1 can be obtained as 63.
[0112] Calculate NumSubPicGridRows and NumSubPicGridCols through the following equations.
[0113] NumSubPicGridCols =
[0114] (pic_width_in_luma_samples +
[0115] subpic_grid_col_width_minus1 * 4+3) /
[0116] (subpic_grid_col_width_minus1 * 4 + 4)
[0117] NumSubPicGridRows =
[0118] (pic_height_in_luma_samples +
[0119] subpic_grid_row_height_minus1 * 4+3) /
[0120] (subpic_grid_row_height_minus1 * 4 + 4)
[0121] In this embodiment, NumSubPicGridCols is derived as 6, and NumSubPicGridRows is derived as 3.
[0122] Thereafter, the index information indicating which sub-picture each grid represented by subpic_grid_idx[i][j] belongs to is obtained as many as the necessary number represented by NumSubpicGridRows and NumSubpicGridCols. In this embodiment, this number is 18. As for the values, 0 is obtained for [0][0], [0][1], [0][2], [1][0], [1][1], [1][2], [2][0], [2][1], and [2][2]. Additionally, 1 is obtained for [0][3], [0][4], [0][5], [1][3], [1][4], [1][5], [2][3], [2][4], and [2][5]. According to this information, it can be restored that the image is segmented into two sub-pictures, and the two sub-pictures are segmented as shown in Figure 8 Thus, the bitstream that has been effectively encoded with respect to the syntax of the sub-picture generated in the first embodiment can be decoded.
[0123] After obtaining the information, information such as slice data segmentation information and block data segmentation information required for decoding is obtained. The segmentation information derived from the separation decoding unit 202 is sent to the image reproduction unit 205 and is used to specify the position of the data in the image that will be processed in step S404.
[0124] In step S402, the decoding unit 203 decodes the code data separated in step S401 and reproduces the quantization coefficients and prediction information. In step S403, the inverse quantization / inverse transformation unit 204 inverse quantizes the quantization coefficients to obtain transformation coefficients, and also performs an inverse orthogonal transformation to reproduce the prediction error.
[0125] In step S404, the image reproduction unit 205 reproduces the prediction information and the predicted image generated in step S403. The image reproduction unit 205 also reproduces the image data based on the reproduced predicted image and the prediction error generated in step S404. The reproduced image data is synthesized into the appropriate position in the image based on the segmentation information generated in step S401.
[0126] In step S405, the image decoding device determines whether the decoding of all basic blocks in the frame has ended. If the decoding has ended, the process proceeds to step S406. Otherwise, the process returns to step S402 to process the next basic block.
[0127] In step S406, the in-loop filter unit 207 performs an in-loop filtering process on the image data reproduced in step S404 to generate an image that has been filtered, and the process ends.
[0128] Using the above configuration and operations, the bitstream generated by effectively encoding the syntax of the sub-pictures generated in the first embodiment can be decoded.
[0129] Note that in this embodiment, a value obtained by subtracting 2 from the maximum number of sub-pictures is obtained. However, the present invention is not limited thereto. Instead, max_subpics_minus1 obtained by subtracting 1 from the maximum number of sub-pictures can be obtained, and when this value is 1 or greater, that is, only when the maximum number of sub-pictures is 2 or greater, the following information can be obtained. That is, subpic_grid_col_width_minus1, subpic_grid_row_height_minus1, and subpic_grid_idx[i][j] can be obtained. That is, if the maximum number is 1, these information may not be obtained. This enables decoding of a bitstream with reduced redundant sub-picture information.
[0130] Note that in the above description, when the value of max_subpics_minus1 is 1 or greater, that is, only when the maximum number of sub-pictures is 2 or greater, the following information is obtained. That is, when the maximum number of sub-pictures is 2 or greater, the information of subpic_grid_col_width_minus1, subpic_grid_row_height_minus1, and subpic_grid_idx[i][j] is obtained. When the maximum number of sub-pictures is 2 or greater, the following two information can also be obtained. That is, only when the value of max_subpics_minus1 is 1 or greater, the two values of subpic_treated_as_flag and loop_filter_across_subpic_enabled_flag can also be obtained. This enables decoding of a bitstream with further reduced redundant sub-picture information.
[0131] In addition, in this embodiment, an example of dividing an image into two sub-pictures has been described. However, the present invention is not limited thereto. For example, in Figure 10 the case where the shown image is divided into three sub-pictures, the bitstream created by an encoding method for setting the maximum value of subpic_grid_idx to max_subpics_minus1 or less can also be decoded.
[0132] In addition, in this embodiment, a four-pixel defined grid is used. However, a CTU row / column defined grid can be used. In this case, when restoring the values of subpic_grid_col_width_minus1 and subpic_grid_row_height_minus1, the bitstream with reduced code amount can also be decoded by calculating based on the number of basic pixels of the CTU.
[0133] Note that, in this embodiment, subpic_grid_idx, which is an ID representing the position of the rectangular area of the subpicture, is obtained from the sequence parameter set. However, the present invention is not limited to this. Figure 11 The configuration of the bitstream is shown, where subpic_idx, which is the ID of the corresponding subpicture, is changed to be encoded in the slice header that is the header of each slice. Figure 11 The configuration of the bitstream is shown, where subpic_idx, which is the ID of the corresponding subpicture, is changed to be encoded in the slice header that is the header of each slice. When such a configuration is adopted, the number of pieces of information representing the position of the rectangular area of the subpicture can be reduced from the number of grids to the number of slices. In Figure 8 the example shown, the number of grids is 18, and the number of subpictures = the number of slices. Therefore, since the number of slices is 2, decoding can be performed by obtaining two subpic_idx and calculating the same information as 18 subpic_grid_idx.
[0134] In addition, an embodiment in which subpic_idx is arranged in the slice header has been described. However, the present invention is not limited to this, and even if subpic_idx is arranged in the block code data, the same information as subpic_grid_idx can be decoded.
[0135] (Third Embodiment)
[0136] In the above embodiment, it is described assuming that Figure 1 the processing unit shown in 1 or 2 is formed by hardware. However, the processing performed by the processing unit shown in these drawings can be configured by a computer program.
[0137] Figure 5 is a block diagram showing an example of the configuration of the hardware of a computer applicable to the image processing apparatus according to each embodiment.
[0138] The CPU 501 uses computer programs and data stored in the RAM 502 or the ROM 503 to control the entire computer, and also executes the above-described various processes as the processes to be performed by the image processing apparatus according to the above embodiment. That is, the CPU 501 serves as Figure 1 or Figure 2The processing unit shown.
[0139] RAM 502 includes areas configured to temporarily store computer programs and data loaded from an external storage device 506 and data acquired from the outside via an I / F (interface) 507. In addition, RAM 502 includes a work area used by the CPU 501 to perform various processes. That is, for example, RAM 502 can be allocated as a frame memory, or various other types of areas can be appropriately provided.
[0140] ROM 503 stores setting data of the computer, a boot program, etc. The operation unit 504 is formed by a keyboard, a mouse, etc. When a user of the computer operates the operation unit 504, various instructions can be input into the CPU 501. The display unit 505 displays the processing results of the CPU 501. In addition, the display unit 505 is formed by, for example, a liquid crystal display.
[0141] The external storage device 506 is a mass information storage device represented by a hard disk drive. The external storage device 506 stores an OS (operating system) and a computer program configured to enable the CPU 501 to implement Figure 1 or Figure 2 the functions of the unit shown. The external storage device 506 can also store each image data to be processed.
[0142] Under the control of the CPU 501, the computer programs and data stored in the external storage device 506 are appropriately loaded into the RAM 502 and processed by the CPU 501. The I / F 507 can connect to a network such as a LAN or the Internet or other devices such as a projection device or a display device. The computer can acquire or send various information via the I / F 507. Reference numeral 508 denotes a bus connecting the above units.
[0143] For the operation formed by the above configuration, the CPU 501 plays a main role in controlling the operation described with reference to the above flowchart.
[0144] Other embodiments
[0145] The present invention can be implemented by the following processing: supplying a program for implementing one or more functions of the above embodiment to a system or device via a network or a storage medium, and causing one or more processors in the computer of the system or device to read and execute the program. The present invention can also be implemented by a circuit (for example, an ASIC) for implementing one or more functions.
[0146] The present invention is not limited to the above embodiments, and various changes and modifications can be made within the spirit and scope of the present invention. Therefore, in order to inform the public of the scope of the present invention, the appended claims are proposed.
[0147] This application claims the priority of Japanese Patent Application No. 2019-165580 filed on September 11, 2019, which is hereby incorporated by reference.
Claims
1. An image encoding device for dividing an image into one or more sub - pictures each including basic blocks and encoding the one or more sub - pictures, the image encoding device comprising: A first encoding unit configured to encode first information corresponding to the block size of each basic block into a sequence parameter set of a bitstream, wherein the block size of each basic block is derived by performing an arithmetic left - shift of 1 on the sum of the value of the first information and a predetermined integer, and wherein a basic block can be split into multiple blocks, and information on a given prediction mode for the blocks is encoded into the bitstream; A second encoding unit configured to encode a flag related to the presence of information of a sub - picture into the sequence parameter set of the bitstream; A third encoding unit configured to, when the flag has a value indicating the presence of information of the sub - picture, encode a syntax element corresponding to the number of sub - pictures included in the image and second information of an ID related to the sub - pictures in the image into the sequence parameter set of the bitstream; and A fourth encoding unit configured to encode third information for identifying the arrangement of sub - pictures in the image into the sequence parameter set of the bitstream, wherein when the block size derived based on the first information is 64, the value of the third information is equal to the result of subtracting 1 from the value identified in units of 64; wherein, when the value represented by the syntax element is one or more, the third information for identifying the arrangement of sub - pictures in the image is encoded into the sequence parameter set of the bitstream; wherein, when the value represented by the syntax element is not one or more, the third information for identifying the arrangement of sub - pictures in the image is not encoded into the sequence parameter set of the bitstream; wherein the sub - picture includes stripes, and wherein, in the bitstream, the sequence parameter set is located before a slice header including information related to the stripes.
2. The apparatus according to claim 1, wherein, The value of the first information is an integer.
3. The device according to claim 1, wherein The sub - picture is a rectangular area.
4. The device according to claim 1, wherein, The loop_filter_across_subpic_enabled_flag related to filtering can be encoded into the bitstream.
5. The device according to claim 4, wherein, When the value represented by the syntax element is not one or more, the loop_filter_across_subpic_enabled_flag is not encoded into the bitstream.
6. An image decoding device capable of decoding a bitstream obtained by encoding an image including one or more sub - pictures each including basic blocks, the image decoding device comprising: A first decoding unit configured to decode first information corresponding to the block size of each basic block from the sequence parameter set of the bitstream, wherein the block size of each basic block is derived by performing an arithmetic left - shift of 1 on the sum of the value of the first information and a predetermined integer, and wherein a basic block can be split into multiple blocks, and decode information on a given prediction mode for the blocks from the bitstream; A second decoding unit configured to decode, from the sequence parameter set of the bitstream, a flag related to the presence of information of a sub-picture; A third decoding unit configured to, when the flag has a value indicating the presence of information of the sub-picture, decode, from the sequence parameter set of the bitstream, a syntax element corresponding to the number of sub-pictures included in the picture and second information of an ID related to the sub-pictures in the picture; and A fourth decoding unit configured to decode, from the sequence parameter set of the bitstream, third information for identifying the arrangement of sub-pictures in the picture, wherein, when the block size derived based on the first information is 64, the value of the third information is equal to the result of subtracting 1 from the value identified in units of 64; wherein, when the numerical value represented by the syntax element is one or more than one, the third information for identifying the arrangement of sub-pictures in the picture is decoded from the sequence parameter set of the bitstream; wherein, when the numerical value represented by the syntax element is not one or more than one, the third information for identifying the arrangement of sub-pictures in the picture is not decoded from the sequence parameter set of the bitstream; wherein the sub-picture includes stripes, and wherein, in the bitstream, the sequence parameter set is located before a slice header including information related to the stripes.
7. The apparatus according to claim 6, wherein, The value of the first information is an integer.
8. The device according to claim 6, wherein, The sub-picture is a rectangular area.
9. The device according to claim 6, wherein The loop_filter_across_subpic_enabled_flag related to filtering can be decoded from the bitstream.
10. The device according to claim 9, wherein, When the numerical value represented by the syntax element is not one or more than one, the loop_filter_across_subpic_enabled_flag is not decoded from the bitstream.
11. An image encoding method for dividing an image into one or more than one sub-pictures each including basic blocks and encoding the one or more than one sub-pictures, the image encoding method comprising: Encoding first information corresponding to the block size of each basic block into a sequence parameter set of a bitstream, wherein the block size of each basic block is derived by performing an arithmetic left shift of 1 on the sum of the value of the first information and a predetermined integer, and wherein a basic block can be split into multiple blocks, and information on a given prediction mode for the blocks is encoded into the bitstream; Encoding a flag related to the presence of information of a sub-picture into the sequence parameter set of the bitstream; When the flag has a value indicating the presence of information of the sub-picture, encoding a syntax element corresponding to the number of sub-pictures included in the picture and second information of an ID related to the sub-pictures in the picture into the sequence parameter set of the bitstream; and Encoding third information for identifying the arrangement of sub-pictures in the picture into the sequence parameter set of the bitstream, wherein, when the block size derived based on the first information is 64, the value of the third information is equal to the result of subtracting 1 from the value identified in units of 64; Among them, when the value represented by the syntax element is one or more than one, the third information for identifying the arrangement of sub-pictures in the image is encoded into the sequence parameter set of the bitstream. Among them, when the value represented by the syntax element is not one or more than one, the third information for identifying the arrangement of sub-pictures in the image is not encoded into the sequence parameter set of the bitstream. Among them, the sub-picture includes stripes, and Among them, in the bitstream, the sequence parameter set is located before the slice header including information related to the stripes.
12. An image decoding method capable of decoding a bitstream obtained by encoding an image including one or more than one sub-picture each including a basic block, the image decoding method includes: Decoding first information corresponding to the block size of each basic block from the sequence parameter set of the bitstream, where the block size of each basic block is derived by performing an arithmetic left shift of 1 on the sum of the value of the first information and a predetermined integer, and where a basic block can be split into multiple blocks, and decoding information on a given prediction mode for the blocks from the bitstream; Decoding a flag related to the existence of information on sub-pictures from the sequence parameter set of the bitstream; When the flag has a value indicating the existence of information on the sub-picture, decoding a syntax element corresponding to the number of sub-pictures included in the image and second information on the ID related to the sub-pictures in the image from the sequence parameter set of the bitstream; and Decoding third information for identifying the arrangement of sub-pictures in the image from the sequence parameter set of the bitstream, where when the block size derived based on the first information is 64, the value of the third information is equal to the result of subtracting 1 from the value identified in units of 64. Among them, when the value represented by the syntax element is one or more than one, decoding the third information for identifying the arrangement of sub-pictures in the image from the sequence parameter set of the bitstream. Among them, when the value represented by the syntax element is not one or more than one, not decoding the third information for identifying the arrangement of sub-pictures in the image from the sequence parameter set of the bitstream. Among them, the sub-picture includes stripes, and Among them, in the bitstream, the sequence parameter set is located before the slice header including information related to the stripes.
Citation Information
Patent Citations
Image encoder, image encoding method and program, image decoder, and image decoding method and program
JP2014011638A
Power system, controller, power management method, program, and power management server
JP2019165580A